Commit Graph
37 Commits
Author SHA1 Message Date
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
Brandon Squizzato cc780d4e17 feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS

Add grpcRoute.backendTLSPolicy values to optionally create a
BackendTLSPolicy resource that enables end-to-end TLS between the
Gateway proxy and the OpenShell gateway pod. The Gateway proxy
terminates client-facing TLS and re-encrypts when connecting to the
backend, validating the pod's certificate against a user-supplied CA
ConfigMap.

This removes the requirement to set server.disableTls=true when using
HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other
platforms with BackendTLSPolicy support in the Gateway API
implementation.

Update OpenShift and ingress documentation with e2e TLS instructions.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm,server): auto-create backend CA ConfigMap in certgen hook

Extend the generate-certs command with --backend-ca-configmap-name and
--backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the
certgen pre-install hook creates the CA ConfigMap automatically:

- pkiInitJob mode (default): uses the CA from the generated PKI bundle.
  Fully automatic on first install.
- cert-manager mode: reads ca.crt from the server TLS Secret. On first
  install the Secret does not exist yet (cert-manager reconciles after
  templates are applied), so the ConfigMap is created on the first helm
  upgrade. Logs a warning on the initial skip.

The caCertificateConfigMapName value now defaults to <fullname>-backend-ca
when empty, so users only need to set backendTLSPolicy.enabled=true.

Update certgen RBAC to include configmaps get/create when the feature is
enabled. Add CLI arg parsing tests for the new flags.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* refactor(helm): add server.tls.enableMtls flag for mTLS control

Replace automatic mTLS disabling based on BackendTLSPolicy with an
explicit server.tls.enableMtls flag that defaults to true. The user is
now responsible for setting this to false when using BackendTLSPolicy,
as ingress proxies cannot present client certificates to backends.

Updated:
- values.yaml: Added server.tls.enableMtls (default true)
- gateway-config.yaml: Check enableMtls instead of backendTLSPolicy
- _gateway-workload.tpl: Check enableMtls for client CA mount
- Tests: Updated to use enableMtls flag
- Docs: Added enableMtls=false to BackendTLSPolicy examples
- README: Document new flag and BackendTLSPolicy requirement

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(helm): clarify cert-manager backend CA ConfigMap workflow

Update documentation to explain the two-step install process required
when using cert-manager with BackendTLSPolicy:

1. helm install - cert-manager issues the server certificate, but the
   certgen hook can't create the backend CA ConfigMap yet (cert-manager
   reconciles after templates are applied)
2. helm upgrade - certgen hook reads the CA from the cert-manager-issued
   certificate and creates the ConfigMap

Previously, the docs said "created on first upgrade" without explaining
why or that the feature won't work until then. The updated docs now:
- Explain the timing issue (cert-manager reconciles after chart install)
- Provide clear steps for the cert-manager workflow
- Note that pkiInitJob (default) creates it immediately on install
- Clarify that users must wait for the Certificate to be Ready before
  running the second upgrade

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy

BackendTLSPolicy validates the backend certificate against the service FQDN
(e.g., openshell.openshell.svc.cluster.local), not the external hostname.
The external hostname only needs to be on the Gateway listener certificate
for client-facing TLS.

The default certManager.serverDnsNames already includes all required service
FQDN variants, so no configuration is needed for BackendTLSPolicy to work.

Fixed incorrect documentation that claimed:
- "The server certificate SAN list must include the external hostname"
- Users need to "configure certManager.serverDnsNames with the external hostname"

Removed the unnecessary pkiInitJob.serverDnsNames override from the example
and clarified that:
- Gateway listener certificate needs the external hostname (for clients)
- Backend certificate needs the service FQDN (for Gateway proxy)
- The service FQDN is already in the defaults

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs: clarify ACME with LetsEncrypt reference

Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer"
to help users understand that LetsEncrypt is the most common ACME
provider and what ACME means in practice.

Updated:
- docs/kubernetes/managing-certificates.mdx
- docs/kubernetes/openshift.mdx
- deploy/helm/openshell/values.yaml
- deploy/helm/openshell/README.md
- deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname

Reorganize the OpenShift production deployment documentation:

1. Changed main section from "Production Deployments" to "Options for
   end-to-end TLS" for better clarity

2. Renamed subsections for consistency and clarity:
   - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)"
   - "End-to-end TLS using pass-through Route (all OpenShift versions)"

3. Clarified that the Gateway hostname is typically a wildcard:
   "typically a wildcard like *.openshell-ingress-gw.example.com"

4. Removed the recommendation to copy the cluster's wildcard certificate
   from openshift-ingress namespace, as this is not a recommended
   security best practice

These changes make it clearer that users have two end-to-end TLS options
and help them understand the typical naming pattern for Gateway hostnames.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager

When using BackendTLSPolicy with cert-manager, the certgen hook now polls
for up to 90 seconds waiting for cert-manager to issue the TLS certificate
before creating the backend CA ConfigMap. This eliminates the need for a
second `helm upgrade` in most cases.

The hook polls every 2 seconds with progress logging every 10 seconds.
If cert-manager takes longer than 90 seconds, the hook times out gracefully
and logs a warning, preserving the fallback to manual ConfigMap creation
or a second upgrade.

The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s
margin for ConfigMap creation and hook completion.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable timeout for certgen hook

Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long
the certgen hook Job can run. When using cert-manager with BackendTLSPolicy,
the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap
creation and cleanup.

This allows users to increase the timeout for environments where cert-manager
takes longer than 90 seconds to issue certificates, without requiring code
changes.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180  # Hook polls for 150 seconds
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): document configurable certgen timeout

Update documentation to mention the pkiInitJob.timeoutSeconds value and
how it affects the cert-manager polling behavior when using BackendTLSPolicy.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable failure behavior for certgen timeout

Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether
the certgen hook fails or succeeds when cert-manager does not issue a
certificate within the polling timeout.

When false (default), the hook succeeds with a warning and users can run
`helm upgrade` after cert-manager issues the certificate to create the
backend CA ConfigMap. This provides backwards-compatible behavior.

When true, the hook fails immediately if the timeout is reached, providing
clear feedback that BackendTLSPolicy is non-functional. This is useful for
strict validation requirements where incomplete installs should fail fast.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180
  failOnTimeout: true  # Fail install if cert-manager takes >150s
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): change failOnTimeout default to true and add troubleshooting docs

Change `pkiInitJob.failOnTimeout` default from false to true to provide
immediate feedback when cert-manager does not issue certificates within
the polling timeout. This prevents silent failures where BackendTLSPolicy
is non-functional but the install appears to succeed.

Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx
documenting the specific error "TLS error: Secret is not supplied by SDS"
that occurs when the backend CA ConfigMap is missing, with step-by-step
resolution instructions.

Updated comments in values.yaml to clearly document the default behavior
and explain when administrators might see connectivity errors if they
override the default to failOnTimeout=false.

BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm
installs will fail if cert-manager takes longer than (timeoutSeconds - 30)
seconds to issue certificates. To restore the old behavior of allowing
installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make cert-manager resources pre-install hooks to fix ordering

Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks
with weight -30, before the certgen hook (weight -20). This fixes the
chicken-and-egg problem where the certgen hook was waiting for Secrets
created by Certificates that hadn't been created yet.

**Hook ordering:**
1. Certificate and Issuer resources created (weight -30)
2. cert-manager issues certificates and creates Secrets
3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap
4. Main resources (StatefulSet, Service, etc.) created

Previously, the certgen pre-install hook would run before any resources
were created, poll for a non-existent Secret, timeout, and fail. The
Certificate resources would never get created because Helm waits for
all pre-install hooks to succeed before creating main resources.

This fix allows single-stage installs to work reliably as long as
cert-manager can issue certificates within the polling timeout.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add validation to prevent enableMtls with BackendTLSPolicy

Add Helm chart validation that fails the install if both
server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true
are set, since this is an invalid configuration.

BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy
cannot present client certificates to the backend. This validation provides
immediate, clear feedback at install time rather than allowing the
misconfiguration to be discovered through runtime errors.

Example error message:
```
Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because
the Gateway proxy cannot present client certificates to the backend;
set server.tls.enableMtls=false
```

Also updated documentation to mention this validation check.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior

Improve documentation to clearly explain that pkiInitJob.timeoutSeconds
controls the Job deadline, but the actual polling timeout is
(timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation
and cleanup.

Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling"
to make the relationship explicit and avoid confusion where users might
expect the hook to poll for the full timeout value.

Updated both values.yaml inline comments and ingress.mdx documentation
for consistency.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(openshift): remove outdated two-stage install instructions

Update OpenShift documentation to reflect that single-stage installs now
work with cert-manager and BackendTLSPolicy. The Certificate resources
run as pre-install hooks (weight -30) before certgen (weight -20),
allowing the hook to poll for and find the issued certificates.

Removed the outdated two-step process:
1. helm install (cert-manager issues cert, hook logs warning)
2. helm upgrade (hook creates ConfigMap)

Replaced with current single-stage behavior:
- Certificate resources created as pre-install hooks
- certgen hook polls for up to 90 seconds (configurable)
- Single helm install succeeds in most cases
- Fails fast by default if timeout reached

This brings openshift.mdx in line with the already-updated ingress.mdx
documentation.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration

The timeout value now represents the actual polling time that users
experience when waiting for cert-manager to issue certificates.
The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to
allow buffer time for ConfigMap creation and cleanup.

Previously, the hook polled for (timeoutSeconds - 30) seconds, which
was confusing when users set timeoutSeconds=180 and only got 150
seconds of actual polling.

Updated documentation in values.yaml, ingress.mdx, and openshift.mdx
to reflect the clearer behavior.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(helm): update values.yaml and README with correct polling duration

Updated the caCertificateConfigMapName description to reflect that the
hook polls for exactly pkiInitJob.timeoutSeconds seconds, not
(timeoutSeconds - 30) seconds.

Regenerated README.md with helm-docs.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): add OIDC configuration to helm install and CLI examples

Updated all helm install and openshell gateway add examples in ingress.mdx
and openshift.mdx to include OIDC issuer and audience configuration.

Examples now use concrete placeholder values:
- OIDC issuer: https://keycloak.example.com/realms/openshell
- OIDC audience: openshell-cli
- Hostname: gateway.example.com
- ClusterIssuer: letsencrypt-prod

This makes it clearer how to configure OIDC authentication, which is
required when using BackendTLSPolicy or HTTPS termination since the
Gateway proxy cannot present client certificates to the backend.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): explicitly list OIDC client ID in gateway add examples

Added --oidc-client-id openshell-cli to all openshell gateway add
commands in ingress.mdx and openshift.mdx, making the default client
ID explicit in the examples even though it's the CLI default.

This improves clarity and helps users understand the complete OIDC
configuration needed for gateway registration.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* fix(helm): address PR review feedback for BackendTLSPolicy

- Read backend CA from the authoritative server Secret instead of the
  in-memory PKI bundle so enabling BackendTLSPolicy on an existing
  release uses the CA that actually signed the server certificate.
- Reconcile the backend CA ConfigMap on every hook run (compare and
  update) instead of skipping when it already exists, so CA rotations
  propagate automatically.
- Remove hook annotations from cert-manager Issuer/Certificate resources
  so they remain regular release objects managed by Helm lifecycle. Split
  the cert-manager backend CA ConfigMap creation into a separate
  post-install/post-upgrade hook Job that polls after cert-manager
  Certificate resources are applied.
- Update architecture/gateway.md, docs/reference/gateway-config.mdx,
  debug-openshell-cluster skill, and helm-dev-environment skill with
  BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout
  documentation.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): resolve markdown lint errors in helm README and kubernetes docs

Escape inline HTML angle brackets in README.md template placeholders,
remove trailing spaces, and add blank lines around fenced code blocks
in numbered lists.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Update docs

* fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile

Wrap `<fullname>` and `<namespace>` template placeholders in backticks
so markdownlint does not flag them as inline HTML (MD033). Regenerate
mise.lock to match current mise.toml after rebase onto main.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Run 'mise lock'

* fix: align mise.lock with CI mise version output

The lockfile was regenerated locally with mise 2026.8.10 which resolves
uv Linux artifacts to gnu variants and adds provenance_verified fields,
but CI uses v2026.4.25 which produces musl variants without those fields.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): correct cert-manager hook ordering and clientCaSecretName comment

Update ingress.mdx and openshift.mdx to describe Certificate resources
as regular release objects with a post-install/post-upgrade Job, matching
the current implementation and architecture/gateway.md.

Fix values.yaml clientCaSecretName comment to state that "" disables
client certificate verification, matching the helper and access-control
docs.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

---------

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
2026-09-14 17:39:15 +00:00
Yuedong Wu 7cc9551677 feat(server): support EC and EdDSA keys in OIDC JWKS validation (#2593)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-04 00:07:29 +00:00
Dhiraj Bokde cc4ded2088 feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(helm): preserve split chart upgrade compatibility

Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology.

* fix(ci): preserve VM runtime for E2E

The Rust cache restores target/ after VM runtime artifacts are staged,
overwriting target/vm-runtime-compressed before openshell-driver-vm is built.
Stage the compressed runtime outside target and pass that location through
OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor.

Also locate the Helm split-ownership test repository root from the script
path rather than git rev-parse. The test runs in a container where the
GitHub checkout can be owned by a different UID and rejected as dubious
ownership.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

* fix(ci): install yq for Helm ownership test

The split-chart ownership regression uses yq to inspect rendered YAML,
but the Helm CI container installs only tools declared in mise.
Declare and lock yq so mise install --locked provides the test dependency.

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>

---------

Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com>
2026-09-02 00:23:39 +00:00
Yuedong Wu e508c169ea fix(helm): honor empty clientCaSecretName for HTTPS-only mode (#2235)
* fix(helm): honor empty clientCaSecretName for HTTPS-only mode

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

* docs(skills): document HTTPS-only clientCaSecretName in debug-openshell-cluster

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>

---------

Signed-off-by: Yuedong Wu <dwcn22@outlook.com>
2026-09-01 18:10:40 +00:00
Drew Newberry 9b6d904e88 feat(compute): delegate sandbox authentication to drivers (#2968)
* feat(compute): delegate sandbox authentication to drivers

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(kubernetes): align sandbox identity annotation

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* test(auth): restore sandbox bootstrap coverage

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* fix(kubernetes): satisfy ownership test lint

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
2026-08-31 17:41:56 +00:00
grs 40d1b48666 feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(proxy): add further tests for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(provider): add runnable example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(e2e): cover Podman token exchange grants

Signed-off-by: Gordon Sim <gsim@redhat.com>

* refactor(oauth): extract duplicated functionality from server and supervisor

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): evict nearest-to-expiry entry from intermediate token cache

Signed-off-by: Gordon Sim <gsim@redhat.com>

* doc(supervisor): add podman example for token exchange

Signed-off-by: Gordon Sim <gsim@redhat.com>

* fix(provider): withhold token-exchange subject credentials

Signed-off-by: Gordon Sim <gsim@redhat.com>

---------

Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-08-24 05:42:30 +00:00
Evan Lezar 3be2cd8a29 fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(kubernetes): share Agent Sandbox setup

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(e2e): wait for Agent Sandbox CRD status

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(canary): sparse-checkout sandbox helper

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-08-21 14:32:19 +00:00
Drew Newberry 998db04780 feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-08-19 21:20:14 +00:00
Jesse Jaggars 2eb0880a00 feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login

Closes #2793

Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow.

The device code flow:
- Requests a device code and user code from the IdP's device authorization endpoint
- Displays a verification URL and user code to the user
- Polls the token endpoint until the user completes authorization or the code expires
- Supports slow_down responses per RFC 8628 by increasing the polling interval

This implementation:
- Extends OidcDiscovery to optionally capture device_authorization_endpoint
- Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes
- Updates gateway add and gateway login to dispatch to device flow when browser is suppressed
- Adds comprehensive unit tests for device flow structs and response parsing
- Updates gateway authentication documentation to describe the device code fallback

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): validate OIDC device token responses

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): add PKCE to OIDC device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(cli): document PKCE device flow

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
2026-08-19 17:44:46 +00:00
LR90 d0c6dc3fd8 feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create

Gateway validation accepted proxy_auth_secret_key values that Kubernetes
rejects when creating the Secret (keys longer than 253 bytes, or the
reserved "."/".." names), turning an invalid deployment setting into
repeated sandbox Pod-provisioning failures instead of a startup error.
Reject them in validate_upstream_proxy_config so they fail closed at
gateway startup.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

* docs(skill): add corporate upstream proxy checks to debug-openshell-cluster

Add a Kubernetes corporate upstream proxy troubleshooting section covering
rendered [openshell.drivers.kubernetes] configuration, credential Secret
volume events, supervisor arguments and mounts confined to the network-
supervising container, and proxy reachability.

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>

---------

Signed-off-by: loveRhythm1990 <qiuweimin@126.com>
2026-08-14 16:16:43 +00:00
Jesse Jaggars c4b500a7de feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA

The supervisor gRPC client only trusted the CA configured via
OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root
store unless with_native_roots()/with_webpki_roots() is also enabled.
Deployments where the gateway server certificate is issued by a public CA
(e.g. cert-manager against an ACME issuer) caused every supervisor
connection to fail the TLS handshake with "UnknownCA", since the sandbox
mTLS CA and the server cert issuer were no longer the same.

Enable both native and webpki roots in addition to the configured CA.
tonic root store is a union of all configured sources, so this does not
weaken verification for existing self-signed deployments. webpki-roots
(compiled in) is enabled alongside native-roots since the supervisor
binary may run in minimal sandbox images without a populated system CA
bundle.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(helm): support external cert-manager issuers and OpenShift Route passthrough

Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS
client certificates can be issued by a real Issuer/ClusterIssuer (e.g.
ACME) instead of only the chart built-in self-signed CA.

Add openshiftRoute template for exposing the gateway via a TLS
passthrough Route so the gateway keeps terminating its own TLS/mTLS.

The server Certificate excludes internal-only SANs (cluster-local,
localhost, loopback) when an external issuer is configured, since ACME
issuers reject those per CA/Browser Forum baseline requirements. A
template-time fail guard catches the misconfiguration at helm install
time rather than asynchronously at cert-manager issuance time.

Includes Helm unittest coverage for both issuerRef overrides and Route
rendering, plus a CI values overlay for lint coverage.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs: document cert-manager external issuer and OpenShift Route

Update managing-certificates.mdx with the serverIssuerRef workflow and
install-time validation behavior. Add a production section to the
OpenShift guide covering passthrough Route with a real certificate.
Regenerate Helm README for new certManager and openshiftRoute values.
Sync debug-openshell-cluster skill with new troubleshooting steps for
ACME issuance failures and supervisor UnknownCA from mismatched CAs.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,core): address PR review feedback on cert-manager external issuer

Addresses all five blocking review items from #2468:

1. Remove .with_native_roots() from supervisor gRPC client -- the
   supervisor runs inside the user-selected sandbox image, so the
   image CA bundle is not operator-controlled. Keep .with_webpki_roots()
   (compiled-in, not user-controlled) alongside the configured CA.

2. Fail at render time when serverIssuerRef.name is set but
   clientCaFromServerTlsSecret is still true. Add negative Helm test.

3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both
   directions because trust bundles are not modeled separately. Change
   serverIssuerRef.kind default from ClusterIssuer to Issuer.

4. Add server.oidc.issuer and server.oidc.audience to the documented
   OpenShift production Helm command. Add Access Control prerequisite.

5. Fail at render time when openshiftRoute.enabled and disableTls are
   both true. Add negative Helm test.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(helm,drivers): guard default clientCaSecretName and add env-strip tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* feat(tls): SNI-based dual certificate for internal and external server TLS

Split the gateway server certificate into two: an internal cert issued by
the chart's own CA (for supervisor connections via cluster-local SANs) and
an external cert issued by an operator-configured Issuer such as ACME/Let's
Encrypt (for CLI and Route access via public SANs).

The gateway uses SNI-based certificate selection: connections whose SNI
hostname matches external_server_names receive the external cert; all
others (including those with no SNI) receive the internal cert.

Security improvement: remove .with_webpki_roots() from the supervisor
gRPC client so supervisors trust only the chart CA, closing a MITM vector
via publicly-trusted certificates in user-supplied container images.

Key changes:
- Add DualCertResolver with SNI-based cert selection and full test coverage
- Add external_cert_path, external_key_path, external_server_names to TlsConfig
- Validate partial external cert config (error on cert-without-key or vice versa)
- Validate empty external_server_names when external cert is configured
- Split cert-manager templates into internal + external Certificate resources
- Add Helm guards for misconfigured external issuer (empty serverDnsNames,
  internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret)
- Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs
- Update debug-openshell-cluster skill for dual-cert troubleshooting

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments

Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM
compute driver that Docker, Podman, and Kubernetes drivers already
perform. Without this, a sandbox user on the VM driver could override
the TLS server name the supervisor verifies.

Fix stale comments in Docker and Podman drivers that referenced
'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted
after the tls-webpki-roots removal.

Use tls-ring instead of bare channel for tonic in openshell-core so the
TLS API (ClientTlsConfig, Endpoint::tls_config) is available without
pulling in any root certificate store.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(tls,helm): wildcard SNI matching and Route host validation

Add RFC 6125 single-level wildcard matching to DualCertResolver so
external_server_names entries like *.example.com correctly match SNI
hostnames like gw.example.com. Previously only exact matches worked,
silently falling back to the internal cert for wildcard configurations.

Add a Helm fail guard in route.yaml that rejects openshiftRoute.host
values not listed in certManager.serverDnsNames when an external issuer
is configured — catches cert/route hostname mismatches at install time
instead of at TLS connect time.

Quote the host field in route.yaml for robustness.

Signed-off-by: Pi Agent <agent@openshell.local>

* fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate

1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true.
   The internal server certificate is always signed by the chart CA (the same
   CA that signs the client cert), so clientCaFromServerTlsSecret=true is
   correct — its filtered ca.crt is exactly the right trust anchor.  The old
   workaround (mounting openshell-ca-tls directly) unnecessarily exposed the
   CA private key to the gateway container.  Remove the client-CA overrides
   from docs, CI overlay, and production examples.

2. Route host validation now supports wildcard certificates per RFC 6125:
   single-level wildcards like *.example.com match gateway.example.com but
   not deep.sub.example.com.  Require an explicit openshiftRoute.host when
   an external issuer is configured — without one, OpenShift generates a
   hostname absent from serverDnsNames.

3. Reject serverIssuerRef.name when certManager.enabled is false — the
   external certificate, its Secret mount, and the gateway TLS config all
   require cert-manager to be enabled.

Validated on ROSA (dev.dyee.p3) with branch-built images:
- Fresh install with letsencrypt-prod ClusterIssuer
- SNI dual-cert: external hostname served Let's Encrypt cert
- Supervisor mTLS via internal cert path: ConnectSupervisor accepted
- Client CA volume: filtered ca.crt from internal server secret (no key)
- CLI connected via Route + OIDC

Helm tests: 81 pass across 7 suites.

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Pi Agent <agent@openshell.local>
2026-08-14 00:12:11 +00:00
Matthew Grossman 170961997f chore(ci): disable telemetry in internal test runs (#2648)
* chore(ci): disable telemetry in internal test runs

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(ci): remove brittle telemetry wiring test

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* docs: trim CI telemetry guidance

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* test(e2e): share telemetry default with OpenShift

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
2026-08-10 21:07:29 +00:00
Taylor MutchandSeth Jennings 8eacb4779f feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology

Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar process id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar process id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): avoid similar proxy id names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): clarify sidecar topology limits

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): keep sidecar process leaf capless

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): refresh sidecar provider env snapshots

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* test(supervisor): align hot-swap identity regression

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): stage sidecar mtls files before proxy chown

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): simplify sidecar supervisor topology

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* chore(helm): reuse sidecar skaffold values

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid similar iptables helper names

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(e2e): harden kube gateway wrapper setup

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(supervisor): avoid nft batch rollback on OCP

Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): preserve process identity in sidecar topology

Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table.

Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): replace sidecar snapshots with control socket

Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files.

Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* feat(kubernetes): support relaxed sidecar network identity

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): satisfy sidecar clippy lint

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): standardize topology naming

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(sandbox): satisfy linux clippy timeout import

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): support kata sidecar on ipv4 pods

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): satisfy linux clippy for sidecar fallback

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* chore(kubernetes): remove stale supervisor topology references

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): enable sidecar binary policy inspection

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): harden sidecar control boundary

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): couple sidecar supervisor lifecycles

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Seth Jennings <sjenning@redhat.com>
Co-authored-by: Seth Jennings <sjenning@redhat.com>
2026-07-10 13:01:39 -07:00
Christian Zaccaria 8c0ecac8cf docs(openshift): simplify install steps and add Helm README entries for OpenShift overrides (#2125)
Signed-off-by: ChristianZaccaria <christian.zaccaria.cz@gmail.com>
2026-07-10 08:45:56 -07:00
Mesut Oezdil 290297ffa3 docs(kubernetes): bump cert-manager to v1.20.3 (#2129) 2026-07-06 17:16:33 +00:00
Huabing (Robin) Zhao abcd15d1f2 feat(helm): add TLS termination for Envoy Gateway ingress (#2015)
The chart's optional Gateway API ingress only rendered a plaintext HTTP
listener, so the gateway could not be exposed over TLS. Add an HTTPS
listener option that terminates TLS at the Envoy Gateway and forwards
plaintext gRPC to the gateway pod.

- gateway.yaml renders an HTTPS listener with `tls.mode: Terminate` and
  `certificateRefs` when `grpcRoute.gateway.listener.protocol=HTTPS`,
  keeping the default HTTP listener unchanged. Guards fail the render when
  `certificateRefs` is empty or `server.disableTls` is not true (the chart
  does not render a BackendTLSPolicy for re-encryption).
- values.yaml adds `grpcRoute.gateway.listener.tls.certificateRefs`.
- ci/values-gateway-tls.yaml exercises the HTTPS branch in lint/render.
- docs/kubernetes/ingress.mdx documents HTTPS setup and clarifies that
  Envoy Gateway only terminates TLS (no OIDC SecurityPolicy); client
  identity uses OIDC bearer tokens, with the client-credentials grant for
  headless agents.
- debug-openshell-cluster skill gains HTTPS-ingress troubleshooting rows.
- Regenerated the chart README values table.

Signed-off-by: Huabing (Robin) Zhao <zhaohuabing@gmail.com>
2026-07-01 08:49:35 -07:00
Taylor Mutch 914da339b4 feat(kubernetes): add combined topology config surface (#2074)
* feat(kubernetes): add combined topology config surface

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): clarify topology defaults

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-30 15:07:55 -07:00
Yuedong Wu 8cb16de9ea chore(deploy): use OCI registry for cert-manager Helm chart (#2041) 2026-06-29 06:44:07 -07:00
Taylor Mutch ba21bb32a2 feat(kubernetes): support agent-sandbox v1beta1 (#2009)
* feat(kubernetes): support agent-sandbox v1beta1

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* ci(kubernetes): test agent-sandbox api versions

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): retry agent-sandbox raw 404s

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): harden agent-sandbox api setup

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(kubernetes): cache agent-sandbox api version

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* docs(kubernetes): document agent-sandbox upgrade behavior

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* refactor(kubernetes): collapse agent-sandbox api selection

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-26 11:44:52 -07:00
Huabing (Robin) Zhao 62b03f0058 fix(docs): add step for creating the GatewayClass (#1984)
Signed-off-by: Huabing (Robin) Zhao <zhaohuabing@gmail.com>
2026-06-24 10:11:01 -07:00
Taylor Mutch 7dab612feb feat(helm): support Deployment kind in HA gateway workloads (#1867)
* feat(helm): support Deployment kind in HA gateway workloads

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(helm): handle null workload values

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-10 23:52:07 +00:00
Taylor Mutch 702cbc4f63 feat(providers): support SPIFFE-backed token grants (#1784)
* feat(providers): support SPIFFE-backed token grants

Add provider profile token_grant metadata and expand endpoint-specific
dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs,
exchange them with an OAuth-style token endpoint, cache returned access
tokens, and inject bearer tokens into matching HTTP requests.

Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload
API socket into sandbox pods for token grant exchange.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Gordon Sim <gsim@redhat.com>

* test(examples): add SPIFFE token grant demo

Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer
and protected services, imports a token-grant provider profile, creates a
sandbox, and verifies endpoint-specific bearer tokens.

The script leaves Kubernetes workloads in place, deletes sandboxes through
openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof
of life.

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(providers): harden SPIFFE token grants

* fix(providers): harden dynamic token grants

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>

* fix(providers): harden token grant handling

---------

Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
Signed-off-by: Gordon Sim <gsim@redhat.com>
2026-06-10 10:54:39 -07:00
Taylor Mutch c4ca283c1a refactor(helm): require external postgres for ha (#1844) 2026-06-09 16:39:24 -07:00
Taylor Mutch e26a1b1ffc fix(kubernetes): configure sandbox apparmor profile (#1767)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com>
2026-06-04 17:24:59 -07:00
Roshni Malani 586c385bdf chore(k8s): use upstream agent-sandbox manifest in CI/e2e (#1657)
Drop the vendored deploy/kube/manifests/agent-sandbox.yaml. CI and e2e
scripts now apply manifest.yaml directly from
github.com/kubernetes-sigs/agent-sandbox releases, pinned via
AGENT_SANDBOX_VERSION env (default v0.4.6, overridable) so internal runs
stay reproducible.

Public docs and the helm chart README continue to reference
/releases/latest/download/ — no OpenShell/agent-sandbox support matrix
exists yet (see #1649).

Adds an air-gap note in docs/kubernetes/setup.mdx enumerating the
manifest and image operators need to mirror to an internal registry.

Signed-off-by: Roshni Malani <rmalani@nvidia.com>
2026-06-04 11:35:06 -07:00
Taylor Mutch 5f58cb018f fix(helm): create sandbox JWT secret when cert-manager is enabled (#1700)
* fix(helm): create sandbox JWT secret under cert-manager

The cert-manager install path (certManager.enabled=true,
pkiInitJob.enabled=false) left the gateway StatefulSet unable to start
because nothing created the openshell-jwt-keys Secret: cert-manager owns
TLS Secrets but does not mint the sandbox JWT signing key, and the
certgen hook only rendered when pkiInitJob.enabled was true.

Separate JWT signing-key provisioning from TLS PKI provisioning:

- certgen: add a --jwt-only mode that creates only the Opaque JWT
  signing Secret, for use when another controller owns TLS Secrets.
- certgen.yaml: render the hook when pkiInitJob.enabled OR
  certManager.enabled is true. cert-manager takes precedence and runs
  the hook with --jwt-only even if pkiInitJob.enabled remains true.
  Remove the mutual-exclusion failure between the two values.
- _helpers.tpl: add openshell.sandboxJwtSecretName, shared by the hook
  and the StatefulSet mount.
- Update values, README, docs, architecture, and the
  debug-openshell-cluster skill to reflect the new precedence; the
  documented cert-manager install no longer needs pkiInitJob.enabled=false.

Closes #1691

* fix(helm): honor cert-manager precedence for client CA volume

The client CA volume logic treated pkiInitJob.enabled as proof that
built-in PKI owns the client CA. With cert-manager precedence now
allowing certManager.enabled=true alongside the default
pkiInitJob.enabled=true, that assumption mounts the server TLS cert
secret as the client CA and ignores
certManager.clientCaFromServerTlsSecret=false, which can break mTLS or
trust the wrong CA.

Gate the pkiInitJob.enabled term with (not certManager.enabled) in all
three client CA conditions (volume mount, volume definition, and secret
selection) so cert-manager owns TLS when enabled. Add a Helm test suite
covering built-in PKI, cert-manager shared CA, the regression config
(cert-manager + clientCaFromServerTlsSecret=false + default pkiInitJob),
and the no-client-CA case.
2026-06-03 12:26:49 -07:00
Taylor Mutch d9908222f2 feat(kubernetes): support sandbox image pull secrets (#1671) 2026-06-01 20:28:50 -07:00
alangou 2bdc968ed9 fix(gateway): make readiness health checks dependency-aware (#1328)
* feat(gateway): add readiness probe metrics and test-only store close

Emit Prometheus readiness metrics for database probes (healthy gauge and
outcome-labeled latency histogram) with coverage in health HTTP tests.
Restrict Store::close behind test support cfg to prevent accidental runtime
pool shutdown under live traffic.

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* test(e2e): add simple e2e test with kubernetes to test /readyz

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-05-27 11:00:06 -07:00
Mesut Oezdil fafde3e1b3 docs(kubernetes): add RBAC section to setup page (#1540)
Documents the ServiceAccount, Role, and ClusterRole created by the Helm
chart inline on the setup page, per reviewer feedback on #1250. Reflects
the current chart templates including pods/get for sandbox identity and
tokenreviews/create for projected token validation.

Closes #1018
2026-05-27 08:53:25 -07:00
Taylor Mutch a3b16c18ab feat(auth): per-sandbox authentication to gateway (#1404) 2026-05-21 17:58:55 -07:00
Taylor MutchandDrew Newberry b61a98dbad feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)

Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.

Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.

The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.

Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.

* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples

Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.

* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003

* docs(gateway): add per-driver TOML example configurations

Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.

A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.

* refactor(gateway): drop image_pull_policy from shared inheritance

Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:

  - Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
    verbatim to the K8s API).
  - Podman: `always | missing | never | newer` (strict lowercase enum
    deserialised into `ImagePullPolicy`).

No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.

Make `image_pull_policy` driver-local:

  - Remove the field from `GatewayFileSection` and from
    `inheritable_keys()` for both Kubernetes and Podman.
  - Drop the file→`RunArgs` merge for the gateway-scope key.
  - Stop unconditionally clobbering the driver value with
    `config.sandbox_image_pull_policy` in the runtime wiring — only apply
    the CLI/env override when it was set (and, for Podman, only when it
    parses into the lowercase enum).
  - Move the key under `[openshell.drivers.kubernetes]` and
    `[openshell.drivers.podman]` in every example, the RFC, the
    architecture doc, and the Helm-rendered gateway ConfigMap.

The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.

* fix(gateway): address review feedback on TOML configuration

Resolves the P1 and P2 issues raised in PR #1317:

- Helm gateway ConfigMap moves `grpc_endpoint` under
  `[openshell.drivers.kubernetes]` so the default install no longer fails
  the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
  the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
  value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
  vocabulary "missing"), so default deployments let the Kubernetes API
  apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
  through the TOML ConfigMap instead of relying on env vars dropped from
  the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
  annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
  `health_bind_address` / `metrics_bind_address`, so a loopback-pinned
  health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
  was previously accepted by the loader but never read).

New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.

* docs(gateway): consolidate gateway TOML examples into docs reference

Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.

The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).

* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver

The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.

* refactor(core): move Podman bridge default into Podman driver

DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.

* docs(auth): scrub remaining SSH handshake secret references

Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.

* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers

The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.

* refactor(gateway): move driver options into config (#1394)

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses

The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.

Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.

Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.

The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).

* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host

CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).

Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).

* fix(e2e): repair podman harness on macOS

Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".

- Discover the macOS socket via `podman machine inspect` instead of the
  hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
  attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
  driver picks up the discovered socket; the driver reads TOML only
  after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
  was no longer enough.

* fix(server): use clone_from for TLS client CA assignment

clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-05-15 12:43:48 -07:00
Taylor Mutch 7a0c444445 refactor!(auth): drop SSH handshake secret (#1274)
* refactor!(auth): drop SSH handshake secret in favor of mTLS

The OPENSHELL_SSH_HANDSHAKE_SECRET / x-sandbox-secret mechanism was
misnamed: it does not authenticate SSH (which flows over the
RelayStream gRPC RPC and is gated by mTLS plus supervisor Unix-socket
permissions). It only gated a small set of sandbox-to-gateway
control-plane RPCs, and production deployments already enforce mTLS
on that channel — so the shared secret was redundant.

Replace the secret check with an mTLS-presence marker. Sandbox-class
methods (ReportPolicyStatus, PushSandboxLogs,
GetSandboxProviderEnvironment, SubmitPolicyAnalysis, GetSandboxConfig,
GetInferenceBundle) accept callers without a Bearer token; the gRPC
mTLS handshake is the trust boundary. Dual-auth methods treat
Bearer-present as full-scope CLI access and Bearer-absent as
sandbox-restricted scope via validate_sandbox_caller_update.

Drops the secret from all drivers (K8s, Podman, VM), the sandbox gRPC
interceptor, the Helm chart (values + pre-install hook + StatefulSet
env), the RPM bootstrap script, the man pages, and the
debug-openshell-cluster skill. Also removes the never-read
ssh_handshake_skew_secs flag and config field.

BREAKING CHANGE: --ssh-handshake-secret / OPENSHELL_SSH_HANDSHAKE_SECRET
and --ssh-handshake-skew-secs / OPENSHELL_SSH_HANDSHAKE_SKEW_SECS are
removed from the gateway, sandbox, and all driver binaries. The
openshell-ssh-handshake K8s Secret is no longer managed by the chart;
operators may delete the orphan. Deployments using
--disable-gateway-auth must enforce caller authentication at the
fronting proxy, since the gateway no longer validates a per-request
secret on sandbox-class methods.

Refs OS-174.

* docs(auth): scrub residual SSH handshake secret references

Sweep across docs, e2e scripts, the Podman driver README/NETWORKING
notes, the RPM/Helm/setup guides, the gateway man page, and RFC 0003
to remove instructions and examples that still referenced
OPENSHELL_SSH_HANDSHAKE_SECRET / --ssh-handshake-secret /
ssh_handshake_skew_secs. The mechanism is gone; nothing should still
suggest setting it. Negative-assertion regression tests are kept so
the env var cannot silently be re-introduced.

* test(drivers): drop SSH handshake secret negative-assertion tests

The supporting code, env vars, CLI flags, and config plumbing are
gone — these tests asserted absence of strings that no longer have
any path to being set. Remove the guards from the Docker, Podman,
Kubernetes, and VM driver test modules.
2026-05-14 13:14:30 -07:00
Piotr Mlocek 0797fefa44 feat(gateway): add local-domain service routing (#1101) 2026-05-12 17:44:55 -07:00
Miyoung Choi 8322e4fd00 docs: style fixes (#1341)
* docs: style fixes

* docs: drop observability section overview page and rename a section title

* docs: title updates
2026-05-12 16:36:15 -07:00
Mesut Oezdil 957daa0a0f docs(helm): document supervisor.sideloadMethod and sandboxNamespace default (#1309)
Add supervisor.sideloadMethod to the Kubernetes setup chart values table
and the compute drivers reference table. The value was added in the
ImageVolumeSource sideload PR and controls whether the supervisor binary
is delivered via an OCI image volume mount or an init container, with
auto-detection based on cluster version when left empty.

Update the server.sandboxNamespace description in both pages to reflect
that the Helm chart now derives it from the release namespace by default
when the value is left empty.
2026-05-11 13:21:56 -07:00
Taylor Mutch d8f614fa31 docs(kubernetes): add initial reference docs (#1243)
* docs(kubernetes): add initial reference docs

Add a Kubernetes reference section under docs/reference/kubernetes/ covering
gateway setup, OIDC and reverse-proxy authentication, cert-manager PKI, and
Gateway API external access. Update installation.mdx to link to the new setup
page and nest the new folder under the Reference section in index.yml.

* docs(kubernetes): promote Kubernetes to top-level nav and retitle pages

Move docs/reference/kubernetes/ to docs/kubernetes/ so the section appears
as a top-level nav entry rather than nested under Reference. Retitle the
pages to broader, action-oriented names (Managing Certificates, Ingress,
Access Control) and rename the files to match. Update cross-references in
setup.mdx and the installation page.

* docs(nav): revert Reference to folder mode

The Reference section was switched to explicit section/contents to host
the nested Kubernetes subsection. With Kubernetes promoted to a top-level
folder, restore folder-mode auto-detection. Side effect: sandbox-compute-drivers
is picked up again, which was silently absent from the explicit list.

* docs(kubernetes): add OpenShift install page

Mirror the OpenShift install instructions from PR #1240 into the
Kubernetes docs section. Documents the SCC binding for sandbox pods
and the chart overrides required by OpenShift's Security Context
Constraints (disabled PKI init Job, plaintext TLS, cleared fsGroup
and runAsUser).
2026-05-07 12:55:56 -07:00