mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 16:39:35 +08:00
* feat(helm): add optional BackendTLSPolicy for e2e TLS Add grpcRoute.backendTLSPolicy values to optionally create a BackendTLSPolicy resource that enables end-to-end TLS between the Gateway proxy and the OpenShell gateway pod. The Gateway proxy terminates client-facing TLS and re-encrypts when connecting to the backend, validating the pod's certificate against a user-supplied CA ConfigMap. This removes the requirement to set server.disableTls=true when using HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other platforms with BackendTLSPolicy support in the Gateway API implementation. Update OpenShift and ingress documentation with e2e TLS instructions. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm,server): auto-create backend CA ConfigMap in certgen hook Extend the generate-certs command with --backend-ca-configmap-name and --backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the certgen pre-install hook creates the CA ConfigMap automatically: - pkiInitJob mode (default): uses the CA from the generated PKI bundle. Fully automatic on first install. - cert-manager mode: reads ca.crt from the server TLS Secret. On first install the Secret does not exist yet (cert-manager reconciles after templates are applied), so the ConfigMap is created on the first helm upgrade. Logs a warning on the initial skip. The caCertificateConfigMapName value now defaults to <fullname>-backend-ca when empty, so users only need to set backendTLSPolicy.enabled=true. Update certgen RBAC to include configmaps get/create when the feature is enabled. Add CLI arg parsing tests for the new flags. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * refactor(helm): add server.tls.enableMtls flag for mTLS control Replace automatic mTLS disabling based on BackendTLSPolicy with an explicit server.tls.enableMtls flag that defaults to true. The user is now responsible for setting this to false when using BackendTLSPolicy, as ingress proxies cannot present client certificates to backends. Updated: - values.yaml: Added server.tls.enableMtls (default true) - gateway-config.yaml: Check enableMtls instead of backendTLSPolicy - _gateway-workload.tpl: Check enableMtls for client CA mount - Tests: Updated to use enableMtls flag - Docs: Added enableMtls=false to BackendTLSPolicy examples - README: Document new flag and BackendTLSPolicy requirement Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(helm): clarify cert-manager backend CA ConfigMap workflow Update documentation to explain the two-step install process required when using cert-manager with BackendTLSPolicy: 1. helm install - cert-manager issues the server certificate, but the certgen hook can't create the backend CA ConfigMap yet (cert-manager reconciles after templates are applied) 2. helm upgrade - certgen hook reads the CA from the cert-manager-issued certificate and creates the ConfigMap Previously, the docs said "created on first upgrade" without explaining why or that the feature won't work until then. The updated docs now: - Explain the timing issue (cert-manager reconciles after chart install) - Provide clear steps for the cert-manager workflow - Note that pkiInitJob (default) creates it immediately on install - Clarify that users must wait for the Certificate to be Ready before running the second upgrade Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy BackendTLSPolicy validates the backend certificate against the service FQDN (e.g., openshell.openshell.svc.cluster.local), not the external hostname. The external hostname only needs to be on the Gateway listener certificate for client-facing TLS. The default certManager.serverDnsNames already includes all required service FQDN variants, so no configuration is needed for BackendTLSPolicy to work. Fixed incorrect documentation that claimed: - "The server certificate SAN list must include the external hostname" - Users need to "configure certManager.serverDnsNames with the external hostname" Removed the unnecessary pkiInitJob.serverDnsNames override from the example and clarified that: - Gateway listener certificate needs the external hostname (for clients) - Backend certificate needs the service FQDN (for Gateway proxy) - The service FQDN is already in the defaults Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs: clarify ACME with LetsEncrypt reference Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer" to help users understand that LetsEncrypt is the most common ACME provider and what ACME means in practice. Updated: - docs/kubernetes/managing-certificates.mdx - docs/kubernetes/openshift.mdx - deploy/helm/openshell/values.yaml - deploy/helm/openshell/README.md - deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname Reorganize the OpenShift production deployment documentation: 1. Changed main section from "Production Deployments" to "Options for end-to-end TLS" for better clarity 2. Renamed subsections for consistency and clarity: - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)" - "End-to-end TLS using pass-through Route (all OpenShift versions)" 3. Clarified that the Gateway hostname is typically a wildcard: "typically a wildcard like *.openshell-ingress-gw.example.com" 4. Removed the recommendation to copy the cluster's wildcard certificate from openshift-ingress namespace, as this is not a recommended security best practice These changes make it clearer that users have two end-to-end TLS options and help them understand the typical naming pattern for Gateway hostnames. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager When using BackendTLSPolicy with cert-manager, the certgen hook now polls for up to 90 seconds waiting for cert-manager to issue the TLS certificate before creating the backend CA ConfigMap. This eliminates the need for a second `helm upgrade` in most cases. The hook polls every 2 seconds with progress logging every 10 seconds. If cert-manager takes longer than 90 seconds, the hook times out gracefully and logs a warning, preserving the fallback to manual ConfigMap creation or a second upgrade. The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s margin for ConfigMap creation and hook completion. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable timeout for certgen hook Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long the certgen hook Job can run. When using cert-manager with BackendTLSPolicy, the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap creation and cleanup. This allows users to increase the timeout for environments where cert-manager takes longer than 90 seconds to issue certificates, without requiring code changes. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 # Hook polls for 150 seconds ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): document configurable certgen timeout Update documentation to mention the pkiInitJob.timeoutSeconds value and how it affects the cert-manager polling behavior when using BackendTLSPolicy. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable failure behavior for certgen timeout Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether the certgen hook fails or succeeds when cert-manager does not issue a certificate within the polling timeout. When false (default), the hook succeeds with a warning and users can run `helm upgrade` after cert-manager issues the certificate to create the backend CA ConfigMap. This provides backwards-compatible behavior. When true, the hook fails immediately if the timeout is reached, providing clear feedback that BackendTLSPolicy is non-functional. This is useful for strict validation requirements where incomplete installs should fail fast. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 failOnTimeout: true # Fail install if cert-manager takes >150s ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): change failOnTimeout default to true and add troubleshooting docs Change `pkiInitJob.failOnTimeout` default from false to true to provide immediate feedback when cert-manager does not issue certificates within the polling timeout. This prevents silent failures where BackendTLSPolicy is non-functional but the install appears to succeed. Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx documenting the specific error "TLS error: Secret is not supplied by SDS" that occurs when the backend CA ConfigMap is missing, with step-by-step resolution instructions. Updated comments in values.yaml to clearly document the default behavior and explain when administrators might see connectivity errors if they override the default to failOnTimeout=false. BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm installs will fail if cert-manager takes longer than (timeoutSeconds - 30) seconds to issue certificates. To restore the old behavior of allowing installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make cert-manager resources pre-install hooks to fix ordering Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks with weight -30, before the certgen hook (weight -20). This fixes the chicken-and-egg problem where the certgen hook was waiting for Secrets created by Certificates that hadn't been created yet. **Hook ordering:** 1. Certificate and Issuer resources created (weight -30) 2. cert-manager issues certificates and creates Secrets 3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap 4. Main resources (StatefulSet, Service, etc.) created Previously, the certgen pre-install hook would run before any resources were created, poll for a non-existent Secret, timeout, and fail. The Certificate resources would never get created because Helm waits for all pre-install hooks to succeed before creating main resources. This fix allows single-stage installs to work reliably as long as cert-manager can issue certificates within the polling timeout. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add validation to prevent enableMtls with BackendTLSPolicy Add Helm chart validation that fails the install if both server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true are set, since this is an invalid configuration. BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend. This validation provides immediate, clear feedback at install time rather than allowing the misconfiguration to be discovered through runtime errors. Example error message: ``` Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend; set server.tls.enableMtls=false ``` Also updated documentation to mention this validation check. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior Improve documentation to clearly explain that pkiInitJob.timeoutSeconds controls the Job deadline, but the actual polling timeout is (timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation and cleanup. Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling" to make the relationship explicit and avoid confusion where users might expect the hook to poll for the full timeout value. Updated both values.yaml inline comments and ingress.mdx documentation for consistency. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(openshift): remove outdated two-stage install instructions Update OpenShift documentation to reflect that single-stage installs now work with cert-manager and BackendTLSPolicy. The Certificate resources run as pre-install hooks (weight -30) before certgen (weight -20), allowing the hook to poll for and find the issued certificates. Removed the outdated two-step process: 1. helm install (cert-manager issues cert, hook logs warning) 2. helm upgrade (hook creates ConfigMap) Replaced with current single-stage behavior: - Certificate resources created as pre-install hooks - certgen hook polls for up to 90 seconds (configurable) - Single helm install succeeds in most cases - Fails fast by default if timeout reached This brings openshift.mdx in line with the already-updated ingress.mdx documentation. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration The timeout value now represents the actual polling time that users experience when waiting for cert-manager to issue certificates. The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to allow buffer time for ConfigMap creation and cleanup. Previously, the hook polled for (timeoutSeconds - 30) seconds, which was confusing when users set timeoutSeconds=180 and only got 150 seconds of actual polling. Updated documentation in values.yaml, ingress.mdx, and openshift.mdx to reflect the clearer behavior. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(helm): update values.yaml and README with correct polling duration Updated the caCertificateConfigMapName description to reflect that the hook polls for exactly pkiInitJob.timeoutSeconds seconds, not (timeoutSeconds - 30) seconds. Regenerated README.md with helm-docs. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): add OIDC configuration to helm install and CLI examples Updated all helm install and openshell gateway add examples in ingress.mdx and openshift.mdx to include OIDC issuer and audience configuration. Examples now use concrete placeholder values: - OIDC issuer: https://keycloak.example.com/realms/openshell - OIDC audience: openshell-cli - Hostname: gateway.example.com - ClusterIssuer: letsencrypt-prod This makes it clearer how to configure OIDC authentication, which is required when using BackendTLSPolicy or HTTPS termination since the Gateway proxy cannot present client certificates to the backend. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): explicitly list OIDC client ID in gateway add examples Added --oidc-client-id openshell-cli to all openshell gateway add commands in ingress.mdx and openshift.mdx, making the default client ID explicit in the examples even though it's the CLI default. This improves clarity and helps users understand the complete OIDC configuration needed for gateway registration. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * fix(helm): address PR review feedback for BackendTLSPolicy - Read backend CA from the authoritative server Secret instead of the in-memory PKI bundle so enabling BackendTLSPolicy on an existing release uses the CA that actually signed the server certificate. - Reconcile the backend CA ConfigMap on every hook run (compare and update) instead of skipping when it already exists, so CA rotations propagate automatically. - Remove hook annotations from cert-manager Issuer/Certificate resources so they remain regular release objects managed by Helm lifecycle. Split the cert-manager backend CA ConfigMap creation into a separate post-install/post-upgrade hook Job that polls after cert-manager Certificate resources are applied. - Update architecture/gateway.md, docs/reference/gateway-config.mdx, debug-openshell-cluster skill, and helm-dev-environment skill with BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout documentation. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): resolve markdown lint errors in helm README and kubernetes docs Escape inline HTML angle brackets in README.md template placeholders, remove trailing spaces, and add blank lines around fenced code blocks in numbered lists. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Update docs * fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile Wrap `<fullname>` and `<namespace>` template placeholders in backticks so markdownlint does not flag them as inline HTML (MD033). Regenerate mise.lock to match current mise.toml after rebase onto main. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Run 'mise lock' * fix: align mise.lock with CI mise version output The lockfile was regenerated locally with mise 2026.8.10 which resolves uv Linux artifacts to gnu variants and adds provenance_verified fields, but CI uses v2026.4.25 which produces musl variants without those fields. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): correct cert-manager hook ordering and clientCaSecretName comment Update ingress.mdx and openshift.mdx to describe Certificate resources as regular release objects with a post-install/post-upgrade Job, matching the current implementation and architecture/gateway.md. Fix values.yaml clientCaSecretName comment to state that "" disables client certificate verification, matching the helper and access-control docs. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> --------- Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
254 lines
12 KiB
Plaintext
254 lines
12 KiB
Plaintext
---
|
|
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
title: "OpenShift"
|
|
sidebar-title: "OpenShift"
|
|
description: "Install the OpenShell Helm chart on OpenShift, including the SCC binding and chart overrides required by OpenShift's Security Context Constraints."
|
|
keywords: "Generative AI, Cybersecurity, Kubernetes, OpenShift, SCC, Security Context Constraints, Helm, Gateway, Installation"
|
|
position: 6
|
|
---
|
|
|
|
<Warning>
|
|
The OpenShift install path is experimental. It currently requires running sandbox pods under the `privileged` SCC and installing the gateway with TLS disabled. Use only for evaluation on a private network.
|
|
</Warning>
|
|
|
|
OpenShift's [Security Context Constraints](https://docs.openshift.com/container-platform/latest/authentication/managing-security-context-constraints.html) reject the chart's default pod security settings. Installing on OpenShift requires precreating the namespace, granting the `privileged` SCC to the sandbox service account, and overriding a few chart values so the cluster admission controller can assign UIDs and FS groups itself.
|
|
|
|
OpenShell installs sandbox nftables rules as individual commands. On OpenShift
|
|
nodes where optional conntrack or packet log expressions are unavailable, those
|
|
optional rules can fail without rolling back the required proxy bypass reject
|
|
rules.
|
|
|
|
## Prerequisites
|
|
|
|
- OpenShift 4.x cluster with `oc` configured
|
|
- Helm 3.x
|
|
- [Agent Sandbox](/kubernetes/setup#install-agent-sandbox) controller and CRDs installed
|
|
|
|
## Install
|
|
|
|
<Steps>
|
|
|
|
## Create the namespace
|
|
|
|
Pre-create the namespace so the SCC binding can be applied before the chart installs:
|
|
|
|
```shell
|
|
oc create ns openshell
|
|
```
|
|
|
|
## Grant the privileged SCC to sandbox pods
|
|
|
|
Sandbox pods run under the `openshell-sandbox` service account in the `openshell` namespace and require the `privileged` SCC:
|
|
|
|
```shell
|
|
oc adm policy add-scc-to-user privileged -z openshell-sandbox -n openshell
|
|
```
|
|
|
|
## Install the chart with OpenShift overrides
|
|
|
|
```shell
|
|
helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart \
|
|
--version <version> \
|
|
--namespace openshell \
|
|
--set server.disableTls=true \
|
|
--set podSecurityContext.fsGroup=null \
|
|
--set securityContext.runAsUser=null
|
|
```
|
|
|
|
| Override | Reason |
|
|
|---|---|
|
|
| `server.disableTls=true` | Runs the gateway over plaintext HTTP for simpler evaluation. |
|
|
| `podSecurityContext.fsGroup=null` / `securityContext.runAsUser=null` | Clear the chart's hardcoded UID and fsGroup so OpenShift's SCC admission can assign them. |
|
|
|
|
## Wait for the gateway to be ready
|
|
|
|
```shell
|
|
oc -n openshell rollout status statefulset/openshell
|
|
```
|
|
|
|
If you set `workload.kind=deployment`, use
|
|
`oc -n openshell rollout status deployment/openshell` instead.
|
|
|
|
</Steps>
|
|
|
|
## Connect to the gateway
|
|
|
|
The gateway is now running over plaintext HTTP. Connect with `oc port-forward`:
|
|
|
|
```shell
|
|
oc -n openshell port-forward svc/openshell 8080:8080
|
|
```
|
|
|
|
Register the gateway with the CLI:
|
|
|
|
```shell
|
|
openshell gateway add http://127.0.0.1:8080 --local --name openshift
|
|
openshell status
|
|
```
|
|
|
|
## Options for end-to-end TLS
|
|
|
|
The steps above run the gateway over plaintext HTTP for quick evaluation. For production deployments, choose one of the approaches below based on your OpenShift version and preferences.
|
|
|
|
### End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)
|
|
|
|
OpenShift 4.22 and later support `BackendTLSPolicy` in the Gateway API, enabling end-to-end TLS between the OpenShift router and the OpenShell gateway pod. The traffic flow is:
|
|
|
|
```text
|
|
client → HTTPS → OpenShift Gateway (terminate TLS) → TLS (re-encrypt) → openshell gateway pod
|
|
```
|
|
|
|
This removes the requirement to run the gateway with `server.disableTls=true`. The OpenShift router terminates client-facing TLS at the listener and re-encrypts when connecting to the backend service, validating the backend's certificate against a CA you provide.
|
|
|
|
#### Prerequisites
|
|
|
|
- OpenShift 4.22+ cluster with the Gateway API enabled
|
|
- cert-manager installed (recommended) or the built-in pkiInitJob for server certificates
|
|
- A `GatewayClass` registered for the OpenShift gateway controller
|
|
|
|
#### Create the GatewayClass
|
|
|
|
If your cluster does not already have an OpenShift GatewayClass, create one:
|
|
|
|
```shell
|
|
oc apply -f - <<'EOF'
|
|
apiVersion: gateway.networking.k8s.io/v1
|
|
kind: GatewayClass
|
|
metadata:
|
|
name: openshift-default
|
|
spec:
|
|
controllerName: openshift.io/gateway-controller/v1
|
|
EOF
|
|
```
|
|
|
|
#### Create the Gateway
|
|
|
|
Create a Gateway resource in the `openshift-ingress` namespace. Replace `<external-hostname>` with your cluster's route hostname (typically a wildcard like `*.openshell-ingress-gw.example.com`):
|
|
|
|
```shell
|
|
oc apply -f - <<'EOF'
|
|
apiVersion: gateway.networking.k8s.io/v1
|
|
kind: Gateway
|
|
metadata:
|
|
name: openshell-gateway
|
|
namespace: openshift-ingress
|
|
spec:
|
|
gatewayClassName: openshift-default
|
|
listeners:
|
|
- name: grpc
|
|
hostname: "<external-hostname>"
|
|
port: 443
|
|
protocol: HTTPS
|
|
tls:
|
|
mode: Terminate
|
|
certificateRefs:
|
|
- name: <listener-tls-secret>
|
|
kind: Secret
|
|
allowedRoutes:
|
|
namespaces:
|
|
from: Selector
|
|
selector:
|
|
matchLabels:
|
|
kubernetes.io/metadata.name: openshell
|
|
EOF
|
|
```
|
|
|
|
The listener TLS Secret should contain the certificate for the external hostname.
|
|
|
|
#### Install with e2e TLS
|
|
|
|
Install the chart with the GRPCRoute and BackendTLSPolicy enabled. The certgen hook automatically creates the backend CA ConfigMap from the generated PKI bundle:
|
|
|
|
```shell
|
|
helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart \
|
|
--version <version> \
|
|
--namespace openshell \
|
|
--set podSecurityContext.fsGroup=null \
|
|
--set securityContext.runAsUser=null \
|
|
--set server.tls.enableMtls=false \
|
|
--set grpcRoute.enabled=true \
|
|
--set grpcRoute.gateway.name=openshell-gateway \
|
|
--set grpcRoute.gateway.namespace=openshift-ingress \
|
|
--set 'grpcRoute.hostnames[0]=gateway.example.com' \
|
|
--set grpcRoute.backendTLSPolicy.enabled=true \
|
|
--set server.oidc.issuer=https://keycloak.example.com/realms/openshell \
|
|
--set server.oidc.audience=openshell-cli
|
|
```
|
|
|
|
| Override | Reason |
|
|
|---|---|
|
|
| `podSecurityContext.fsGroup=null` / `securityContext.runAsUser=null` | Let OpenShift's SCC admission assign UIDs. |
|
|
| `server.tls.enableMtls=false` | Disable mTLS client certificate authentication. BackendTLSPolicy only validates the server certificate; the ingress proxy cannot present a client certificate to the backend. Use OIDC for authentication instead. |
|
|
| `grpcRoute.enabled=true` | Create a GRPCRoute pointing at the external Gateway. |
|
|
| `grpcRoute.gateway.name` / `namespace` | Reference the Gateway created above in `openshift-ingress`. |
|
|
| `grpcRoute.backendTLSPolicy.enabled=true` | Create a BackendTLSPolicy for TLS re-encryption to the gateway pod. The certgen hook auto-creates the backend CA ConfigMap. The Gateway proxy validates the backend certificate against the service FQDN, which is already in the default server certificate SANs. |
|
|
| `grpcRoute.hostnames` | External hostname for the GRPCRoute. This goes on the Gateway listener certificate, not the backend certificate. |
|
|
|
|
Note that `server.disableTls` is **not** set — the gateway pod serves TLS over HTTPS without requiring client certificates. Use OIDC for authentication (see [Access Control](/kubernetes/access-control)).
|
|
|
|
**Using cert-manager instead of pkiInitJob:** Add `--set certManager.enabled=true` to the install command. The default `certManager.serverDnsNames` already includes the service FQDN needed for BackendTLSPolicy validation. The Certificate resources are regular release objects, and a separate post-install/post-upgrade Job (`<release>-certgen-backend-ca`) polls for up to 120 seconds waiting for cert-manager to issue the server certificate, then creates the backend CA ConfigMap. A single `helm install` is sufficient in most cases.
|
|
|
|
If cert-manager takes longer than 120 seconds to issue certificates, increase the polling timeout with `--set pkiInitJob.timeoutSeconds=<seconds>`. The hook polls for exactly this many seconds. For example, `timeoutSeconds=180` polls for 180 seconds. By default (`pkiInitJob.failOnTimeout=true`), the install fails if the timeout is reached, providing clear feedback that the BackendTLSPolicy is non-functional.
|
|
|
|
#### Register over HTTPS
|
|
|
|
```shell
|
|
openshell gateway add https://gateway.example.com \
|
|
--name openshift \
|
|
--oidc-issuer https://keycloak.example.com/realms/openshell \
|
|
--oidc-client-id openshell-cli
|
|
openshell status
|
|
```
|
|
|
|
### End-to-end TLS using pass-through Route (all OpenShift versions)
|
|
|
|
For OpenShift versions prior to 4.22, or when you prefer Route-based ingress, cert-manager can issue the gateway's server certificate from a real Issuer or ClusterIssuer (for example, a LetsEncrypt/ACME issuer), and an OpenShift Route with TLS passthrough exposes it externally while the gateway keeps terminating its own TLS and mTLS.
|
|
|
|
Install cert-manager and configure a working `ClusterIssuer` first — see
|
|
[Managing Certificates](/kubernetes/managing-certificates) for the
|
|
`certManager.serverIssuerRef` details. Configure an OIDC provider as described
|
|
in [Access Control](/kubernetes/access-control) — remote gateways authenticate
|
|
CLI users via OIDC, not mTLS, so the gateway must know the OIDC issuer URL.
|
|
Install the chart with:
|
|
|
|
```shell
|
|
helm install openshell oci://ghcr.io/nvidia/openshell/helm-chart \
|
|
--version <version> \
|
|
--namespace openshell \
|
|
--set podSecurityContext.fsGroup=null \
|
|
--set securityContext.runAsUser=null \
|
|
--set server.disableTls=false \
|
|
--set certManager.enabled=true \
|
|
--set certManager.serverIssuerRef.name=letsencrypt-prod \
|
|
--set certManager.serverIssuerRef.kind=ClusterIssuer \
|
|
--set certManager.serverDnsNames[0]=gateway.example.com \
|
|
--set openshiftRoute.enabled=true \
|
|
--set openshiftRoute.host=gateway.example.com \
|
|
--set server.oidc.issuer=https://keycloak.example.com/realms/openshell \
|
|
--set server.oidc.audience=openshell-cli
|
|
```
|
|
|
|
| Override | Reason |
|
|
|---|---|
|
|
| `certManager.serverIssuerRef` | Creates a second server certificate from your Issuer or ClusterIssuer for external clients. The gateway uses SNI to present this cert for the external hostname while continuing to present the internal (chart CA) cert to supervisors. The internal certificate's `ca.crt` is the chart CA that also signed the client cert, so the default `clientCaFromServerTlsSecret=true` is correct. |
|
|
| `openshiftRoute.enabled` / `openshiftRoute.host` | Creates an OpenShift Route with TLS passthrough — the router forwards the encrypted connection by SNI without decrypting, so the gateway uses the SNI hostname to select the external certificate. |
|
|
| `server.oidc.issuer` / `server.oidc.audience` | Configures server-side OIDC validation. Without these, the gateway expects mTLS client certificates and rejects OIDC-only CLI connections. See [Access Control](/kubernetes/access-control). |
|
|
|
|
Register the gateway with the CLI over OIDC. Remote gateways authenticate CLI
|
|
users via OIDC, not mTLS — see [Access Control](/kubernetes/access-control):
|
|
|
|
```shell
|
|
openshell gateway add https://gateway.example.com \
|
|
--name openshift \
|
|
--oidc-issuer https://keycloak.example.com/realms/openshell \
|
|
--oidc-client-id openshell-cli
|
|
openshell gateway login openshift
|
|
```
|
|
|
|
## Next Steps
|
|
|
|
- For more on certificate provisioning modes, refer to [Managing Certificates](/kubernetes/managing-certificates).
|
|
- To expose the gateway externally through the Kubernetes Gateway API instead of a Route, refer to [Ingress](/kubernetes/ingress).
|
|
- To configure OIDC authentication, refer to [Access Control](/kubernetes/access-control).
|