mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
54
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
021400be8a |
refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(auth): clarify gateway mTLS behavior Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(auth): include workspace scope in TLS authorization checks Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): bound service auth sandbox names for large PIDs Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
2fe5a0e19c |
perf(kubernetes): use a TCP readiness probe for the supervisor (#3700)
- Kubernetes now checks supervisor readiness by connecting to TCP port 5501 - Stop starting a supervisor process in every sandbox each second - The supervisor opens the port only while its gateway session is up - Accept IPv4 and IPv6 probes, even when net.ipv6.bindv6only is set - Keep the health socket for Docker, Podman, and debugging - Add tests and update the docs Signed-off-by: divesh <dgude@nvidia.com> |
||
|
|
d009f30121 |
feat(helm): make cluster-scoped RBAC optional (#3459)
The gateway chart always rendered the ClusterRole and ClusterRoleBinding, so
every install and upgrade required cluster-admin even when only namespaced
objects were needed. Installers that are namespace-admin GitOps or platform
controllers could not run the release at all, and clusters where cluster-scoped
RBAC is owned by a separate team had no supported way to split the install.
Add an rbac values block so a cluster-admin can apply the cluster-scoped objects
once and a namespace-admin can install and upgrade the release without
cluster-scoped permissions:
rbac:
create: true
clusterScoped:
create: true
clusterRoleName: ""
clusterRoleBindingName: ""
rbac.clusterScoped.create gates the ClusterRole and ClusterRoleBinding, and is
independent of the workspace mode. rbac.create additionally gates the namespaced
sandbox Role and RoleBinding, which matters because Kubernetes escalation
prevention stops an installer holding only the built-in admin role from creating
a Role that grants agents.x-k8s.io verbs it does not itself hold. The certgen
hook and credential driver RBAC keep their existing flags.
Both flags default to true, so current installs are unchanged. The helpers treat
a missing rbac block as enabled so upgrades with --reuse-values do not drop RBAC,
matching the existing workspaceResources pattern. The ClusterRoleBinding roleRef
follows clusterRoleName so a separately applied ClusterRole can carry a name the
cluster-admin chooses.
Document the migration for a release that already owns the cluster-scoped
objects: Helm deletes objects that leave the manifest, so annotate them with
helm.sh/resource-policy=keep before setting the flag, otherwise the gateway
loses TokenReview until a cluster-admin re-applies them.
Signed-off-by: ansjindal <ansjindal@nvidia.com>
|
||
|
|
d7f921190b |
docs(policy): refresh policy documentation and references (#3563)
* docs(policy): correct schema and default policy guidance Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): add network recipes and update command reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): organize lifecycle guidance and troubleshooting Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): split policy overview into concepts and management tasks Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): reorganize network recipes as a cookbook Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): restructure schema reference by field group and protocol Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): align troubleshooting, advisor, and reference pages Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): fix first policy tutorial and security guidance Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): keep overview high level and move network rules to their own page Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): focus policy management on CLI workflows and remove command reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): clarify policy views and sandbox deletion in management guide Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): streamline network rule concepts and examples Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct request path wildcard semantics Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): rewrite policy advisor guide for clarity Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): clarify policy advisor scope, setup, and review Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): rewrite policy prover guide for clarity Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): explain the two uses of the policy prover Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe policy prover uses, boundaries, and coverage Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): place prover before advisor and troubleshooting last Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): remove unsupported CI guidance from prover page Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): tighten policy prover introduction Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): move policy change behavior into management guide Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): name prover check types and note expanding coverage Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): prefix prover and advisor sidebar labels Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): streamline policy schema reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): place default policy before schema reference Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): fold troubleshooting into policy management guide Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct tutorial log samples and GitHub push policy steps The first policy tutorial said the 403 body begins with error, policy, and rule, but the proxy serializes the body with sorted keys. Its log samples also showed the wrong CONNECT deny reason for a sandbox without network rules, and the L7 deny sample omitted the :443 authority, the `l7` engine, and the reason tag that the shorthand formatter emits. The GitHub tutorial filtered denials with `--level warn`, which hides the INFO level OCSF policy events, and showed the retired key=value log format. Its hand-written policy also omitted /bin from the restrictive default, so `policy set` would reject the file for removing a filesystem path on a live sandbox. Start from `policy get --base` and add only the network rules. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): improve flow and terminology across policy pages Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct network rule matching and protocol details Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): align policy management steps with CLI behavior Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct policy advisor proposal and approval details Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct policy section, default, and schema details Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): correct prover installation and coverage limits Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): recommend tls skip for server-first protocols Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): fix stale baseline path and interpreter examples Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): move policy pages under how-it-works and fix links Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): align native TCP guidance in security best practices Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): restore policy.local and policy DNS details from main Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): state exact glob matching rules Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
9244868056 |
docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe updated security architecture neutrally Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: highlight new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandboxes): clarify how to disconnect Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align architecture and guides with current navigation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(extensibility): streamline extension authentication guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
52cb8ecee7 |
fix(kubernetes): remove NetworkPolicy acknowledgement (#3677)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
679b190677 |
feat(testing): support independent gateway and supervisor image overrides (#3341)
* feat(testing): normalize configurable test images Signed-off-by: Bobbins228 <mcampbel@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(helm): add global image overrides Signed-off-by: Bobbins228 <mcampbel@redhat.com> * feat(helm): support image registry overrides Signed-off-by: Bobbins228 <mcampbel@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> * refactor(helm): simplify image configuration Signed-off-by: Bobbins228 <mcampbel@redhat.com> * fix(e2e): avoid reloading reused kind sandbox image Signed-off-by: Bobbins228 <mcampbel@redhat.com> * fix(helm): default sandbox image to nvcr.io/nvidia/base/ubuntu:24.04 Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Bobbins228 <mcampbel@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> Co-authored-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
0b351c4a9b |
fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace The Kubernetes Secrets credential driver now stores every credential in its configured namespace in all workspace modes and rejects handles that reference any other namespace before contacting the Kubernetes API. The gateway reaches credential Secrets through the Role in that namespace; this allows removing the Secret rules from the ClusterRole. - Remove the workspace_mode, gateway_id, and allow_reference_namespace driver settings and stop rendering them from Helm. Configurations that set them fail at startup. Existing credential state is not migrated. - Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a dedicated credential namespace. The namespace is kept on uninstall, adopted by a reinstall of the same release, and left untouched when something else owns it. - Update the gateway config reference, Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(helm): reduce gateway Secret permissions Remove the gateway's Secret list permission in every workspace mode and grant source Secret reads through a Role in the sandbox namespace. Bootstrap Secret cleanup deletes Secrets by exact name instead of listing them. - Grant get on the copied client TLS and image-pull Secrets through a Role in the sandbox namespace. The ClusterRole keeps get and patch on those names for the ownership check and server-side apply into workspace namespaces. - Delete sandbox and supervisor bootstrap Secrets by exact name, derived from the runtime generation recorded on the Sandbox and, on restart, the target generation, tolerating 404. The generation annotation is cleared only after cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so garbage collection removes any generation the driver does not name. - Drop Secret list from the ClusterRole and the shared-mode sandbox Role. - Extend the managed e2e RBAC checks to Secret list. - Update the Kubernetes setup and sandbox runtime docs, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(driver-kubernetes): stage workspace Secrets per runtime generation Every Secret the Kubernetes driver writes into a workspace namespace is now scoped to one sandbox runtime generation, immutable, and created with create only, so the gateway never reads, patches, or adopts an existing Secret there. This removes the gateway's cluster-wide get and patch on the copied client TLS and image-pull Secret names. - In managed mode, create an immutable copy of each configured image-pull Secret per generation, named os-pull-<id>-<generation>-<n> and owned by the generation's workload and supervisor Pods. Pods and the restarted Sandbox template reference those names, and generation cleanup deletes them by name. A Secret already holding a generation name fails the create. - Outside shared mode, stage the gateway client TLS material into the supervisor bootstrap Secret instead of copying the client TLS Secret into the workspace namespace. - Remove the fixed-name TLS and image-pull copies, the target ownership read, and the ClusterRole get and patch rule on the copied names. Source reads stay in the sandbox-namespace Role. - Update the managed e2e to expect generation image-pull Secrets and client TLS material in the supervisor bootstrap Secret, and to check that the gateway cannot read the copied names in workspace namespaces. - Update the gateway config and compute driver references, Kubernetes setup and sandbox runtime docs, compute-runtime architecture, driver README, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(helm)!: grant operator-mode Secret permissions through the workspace chart The operator-mode gateway ClusterRole grants no Secret permissions. The openshell-workspace chart Role, installed in each operator-managed namespace, grants the gateway create and delete on Secrets for sandbox runtime generations. Operator-managed namespaces require the workspace chart. - Fail the chart tests on any ClusterRole rule that includes Secrets in operator and shared modes. - Install the workspace chart when the operator e2e provisions a namespace, and assert that the gateway has no Secret permissions in a namespace without it. - Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
a649aa42fc |
docs: add 0.1.0 upgrade guide outline (#3540)
* docs: add 0.1.0 upgrade guide outline Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: add upgrade change provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: update 0.1.0 upgrade guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: move upgrade navigation below security Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: remove release notes page Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: announce OpenShell 0.1.0 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: reorganize guides and require user-owned images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: reorganize navigation around core concepts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: simplify navigation and tutorial catalog Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: refine 0.1.0 release highlights Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align navigation and 0.1.0 guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
a408f5dd08 |
chore(kubernetes): update Agent Sandbox to v1.0.3 (#3578)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
84960e70a3 |
fix(kubernetes): scope resource admission RBAC (#3571)
* fix(kubernetes): scope resource admission RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(helm): gate PVC admission reads Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1e34e8c576 |
fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): address resource admission review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): reserve driver-owned admission labels Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): clarify workspace admission label Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): clarify resource admission failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): configure resource admission fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): retry forbidden admission lookups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): preserve external driver admission defaults Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
293fab75d4 |
fix(sandbox): support kernels < 5.19 via seccomp WAIT_KILLABLE_RECV fallback (#3420)
* fix(sandbox): fall back to plain seccomp listener when WAIT_KILLABLE_RECV is unavailable SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV was added in Linux 5.19. On older kernels (for example RHEL 9.x / 5.14 nodes such as RHCOS on OpenShift) the flag is rejected with EINVAL, which made the capability-free sandbox fail to start during the notification probe with "notification launcher disappeared". Install the notification listener with WAIT_KILLABLE_RECV when the kernel supports it and fall back to a plain NEW_LISTENER on EINVAL. The fallback listener records wait_killable_recv = false: its notification receive is uninterruptible, but the sandbox is otherwise fully functional. Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): address review — legacy read-only listener mode + cancellation invariant Follow-up to the PR review (GATOR-de00bfcc-01 / mrunalp): make the < 5.19 fallback cancellation-safe instead of racing broker writes. - Record an explicit ListenerMode (Killable vs LegacyReadOnly); add writes_disabled()/mode() and emit the selected mode in qualification output (seccomp_listener_mode). - Centralize task-memory output writes behind NotificationListener:: write_task_output; in LegacyReadOnly mode getpeername, accept/accept4 with a non-null address, and sendmmsg length write-backs fail closed with EOPNOTSUPP. accept with a null address, socket/connect/bind/listen/sendto/sendmsg keep working (copied inputs, scalar responses, atomic ADDFD_SEND). - Enforce the launch invariant `cancellation || task_memory_writes_disabled` in SandboxConfirmEvidence::validate() rather than dropping cancellation unconditionally; add task_memory_writes_disabled to SeccompEvidence. - Add per-path fail-closed tests (write_task_output, write_socket_addr, a real plain listener installed on a modern kernel) and confirmation-invariant tests. - Correct the flag-semantics comments and document both modes plus the reduced legacy syscall compatibility in architecture/sandbox.md. Signed-off-by: Akram <akram.benaissi@gmail.com> * docs: document legacy read-only sandbox mode on kernels before 5.19 Kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (< 5.19, e.g. RHEL 9.x / RHCOS 5.14) run the sandbox in a legacy read-only cancellation mode where the broker fails closed with EOPNOTSUPP on the mediated operations that write results back into workload memory (getpeername, accept/accept4 with a non-null address, sendmmsg length write-backs). Document this observable behavior and its syscall limitations in the public Fern docs: the support-matrix kernel requirements and the OpenShift runtime guidance. Signed-off-by: Akram <akram.benaissi@gmail.com> * style(sandbox): satisfy rustfmt and clippy doc_markdown Match the pinned rustfmt (Rust 1.95.0) line-wrapping for the write_task_output call, and backtick `legacy_read_only` in the qualification-report doc comment so clippy::doc_markdown (-D warnings) passes. Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): migrate task_memory_writes_disabled into backend protocol Add the `task_memory_writes_disabled` field to `SeccompEvidence` in the backend protocol contract and relax the validation from requiring `cancellation` to accepting `cancellation || task_memory_writes_disabled`. This completes the rebase migration missed by the isolation-interface refactor: the sandbox reports this field but the backend struct lacked it, and the validator rejected every pre-5.19 legacy listener. Signed-off-by: Akram <akram.benaissi@gmail.com> --------- Signed-off-by: Akram <akram.benaissi@gmail.com> |
||
|
|
c8a4ff5f19 |
fix(kubernetes): bind bootstrap to runtime identity (#3531)
* fix(kubernetes): bind bootstrap to runtime identity Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(compute): compensate runtime binding failures Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(compute): clean up backend on store failure Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(compute): merge runtime binding after start Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(auth): bind restarted sandbox sessions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(compute): fold runtime binding into authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
0a770d9173 |
feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com> |
||
|
|
07d4ac5474 |
fix(helm): grant secret cleanup permissions (#3363)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
cc780d4e17 |
feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS Add grpcRoute.backendTLSPolicy values to optionally create a BackendTLSPolicy resource that enables end-to-end TLS between the Gateway proxy and the OpenShell gateway pod. The Gateway proxy terminates client-facing TLS and re-encrypts when connecting to the backend, validating the pod's certificate against a user-supplied CA ConfigMap. This removes the requirement to set server.disableTls=true when using HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other platforms with BackendTLSPolicy support in the Gateway API implementation. Update OpenShift and ingress documentation with e2e TLS instructions. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm,server): auto-create backend CA ConfigMap in certgen hook Extend the generate-certs command with --backend-ca-configmap-name and --backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the certgen pre-install hook creates the CA ConfigMap automatically: - pkiInitJob mode (default): uses the CA from the generated PKI bundle. Fully automatic on first install. - cert-manager mode: reads ca.crt from the server TLS Secret. On first install the Secret does not exist yet (cert-manager reconciles after templates are applied), so the ConfigMap is created on the first helm upgrade. Logs a warning on the initial skip. The caCertificateConfigMapName value now defaults to <fullname>-backend-ca when empty, so users only need to set backendTLSPolicy.enabled=true. Update certgen RBAC to include configmaps get/create when the feature is enabled. Add CLI arg parsing tests for the new flags. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * refactor(helm): add server.tls.enableMtls flag for mTLS control Replace automatic mTLS disabling based on BackendTLSPolicy with an explicit server.tls.enableMtls flag that defaults to true. The user is now responsible for setting this to false when using BackendTLSPolicy, as ingress proxies cannot present client certificates to backends. Updated: - values.yaml: Added server.tls.enableMtls (default true) - gateway-config.yaml: Check enableMtls instead of backendTLSPolicy - _gateway-workload.tpl: Check enableMtls for client CA mount - Tests: Updated to use enableMtls flag - Docs: Added enableMtls=false to BackendTLSPolicy examples - README: Document new flag and BackendTLSPolicy requirement Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(helm): clarify cert-manager backend CA ConfigMap workflow Update documentation to explain the two-step install process required when using cert-manager with BackendTLSPolicy: 1. helm install - cert-manager issues the server certificate, but the certgen hook can't create the backend CA ConfigMap yet (cert-manager reconciles after templates are applied) 2. helm upgrade - certgen hook reads the CA from the cert-manager-issued certificate and creates the ConfigMap Previously, the docs said "created on first upgrade" without explaining why or that the feature won't work until then. The updated docs now: - Explain the timing issue (cert-manager reconciles after chart install) - Provide clear steps for the cert-manager workflow - Note that pkiInitJob (default) creates it immediately on install - Clarify that users must wait for the Certificate to be Ready before running the second upgrade Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy BackendTLSPolicy validates the backend certificate against the service FQDN (e.g., openshell.openshell.svc.cluster.local), not the external hostname. The external hostname only needs to be on the Gateway listener certificate for client-facing TLS. The default certManager.serverDnsNames already includes all required service FQDN variants, so no configuration is needed for BackendTLSPolicy to work. Fixed incorrect documentation that claimed: - "The server certificate SAN list must include the external hostname" - Users need to "configure certManager.serverDnsNames with the external hostname" Removed the unnecessary pkiInitJob.serverDnsNames override from the example and clarified that: - Gateway listener certificate needs the external hostname (for clients) - Backend certificate needs the service FQDN (for Gateway proxy) - The service FQDN is already in the defaults Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs: clarify ACME with LetsEncrypt reference Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer" to help users understand that LetsEncrypt is the most common ACME provider and what ACME means in practice. Updated: - docs/kubernetes/managing-certificates.mdx - docs/kubernetes/openshift.mdx - deploy/helm/openshell/values.yaml - deploy/helm/openshell/README.md - deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname Reorganize the OpenShift production deployment documentation: 1. Changed main section from "Production Deployments" to "Options for end-to-end TLS" for better clarity 2. Renamed subsections for consistency and clarity: - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)" - "End-to-end TLS using pass-through Route (all OpenShift versions)" 3. Clarified that the Gateway hostname is typically a wildcard: "typically a wildcard like *.openshell-ingress-gw.example.com" 4. Removed the recommendation to copy the cluster's wildcard certificate from openshift-ingress namespace, as this is not a recommended security best practice These changes make it clearer that users have two end-to-end TLS options and help them understand the typical naming pattern for Gateway hostnames. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager When using BackendTLSPolicy with cert-manager, the certgen hook now polls for up to 90 seconds waiting for cert-manager to issue the TLS certificate before creating the backend CA ConfigMap. This eliminates the need for a second `helm upgrade` in most cases. The hook polls every 2 seconds with progress logging every 10 seconds. If cert-manager takes longer than 90 seconds, the hook times out gracefully and logs a warning, preserving the fallback to manual ConfigMap creation or a second upgrade. The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s margin for ConfigMap creation and hook completion. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable timeout for certgen hook Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long the certgen hook Job can run. When using cert-manager with BackendTLSPolicy, the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap creation and cleanup. This allows users to increase the timeout for environments where cert-manager takes longer than 90 seconds to issue certificates, without requiring code changes. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 # Hook polls for 150 seconds ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): document configurable certgen timeout Update documentation to mention the pkiInitJob.timeoutSeconds value and how it affects the cert-manager polling behavior when using BackendTLSPolicy. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable failure behavior for certgen timeout Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether the certgen hook fails or succeeds when cert-manager does not issue a certificate within the polling timeout. When false (default), the hook succeeds with a warning and users can run `helm upgrade` after cert-manager issues the certificate to create the backend CA ConfigMap. This provides backwards-compatible behavior. When true, the hook fails immediately if the timeout is reached, providing clear feedback that BackendTLSPolicy is non-functional. This is useful for strict validation requirements where incomplete installs should fail fast. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 failOnTimeout: true # Fail install if cert-manager takes >150s ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): change failOnTimeout default to true and add troubleshooting docs Change `pkiInitJob.failOnTimeout` default from false to true to provide immediate feedback when cert-manager does not issue certificates within the polling timeout. This prevents silent failures where BackendTLSPolicy is non-functional but the install appears to succeed. Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx documenting the specific error "TLS error: Secret is not supplied by SDS" that occurs when the backend CA ConfigMap is missing, with step-by-step resolution instructions. Updated comments in values.yaml to clearly document the default behavior and explain when administrators might see connectivity errors if they override the default to failOnTimeout=false. BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm installs will fail if cert-manager takes longer than (timeoutSeconds - 30) seconds to issue certificates. To restore the old behavior of allowing installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make cert-manager resources pre-install hooks to fix ordering Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks with weight -30, before the certgen hook (weight -20). This fixes the chicken-and-egg problem where the certgen hook was waiting for Secrets created by Certificates that hadn't been created yet. **Hook ordering:** 1. Certificate and Issuer resources created (weight -30) 2. cert-manager issues certificates and creates Secrets 3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap 4. Main resources (StatefulSet, Service, etc.) created Previously, the certgen pre-install hook would run before any resources were created, poll for a non-existent Secret, timeout, and fail. The Certificate resources would never get created because Helm waits for all pre-install hooks to succeed before creating main resources. This fix allows single-stage installs to work reliably as long as cert-manager can issue certificates within the polling timeout. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add validation to prevent enableMtls with BackendTLSPolicy Add Helm chart validation that fails the install if both server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true are set, since this is an invalid configuration. BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend. This validation provides immediate, clear feedback at install time rather than allowing the misconfiguration to be discovered through runtime errors. Example error message: ``` Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend; set server.tls.enableMtls=false ``` Also updated documentation to mention this validation check. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior Improve documentation to clearly explain that pkiInitJob.timeoutSeconds controls the Job deadline, but the actual polling timeout is (timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation and cleanup. Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling" to make the relationship explicit and avoid confusion where users might expect the hook to poll for the full timeout value. Updated both values.yaml inline comments and ingress.mdx documentation for consistency. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(openshift): remove outdated two-stage install instructions Update OpenShift documentation to reflect that single-stage installs now work with cert-manager and BackendTLSPolicy. The Certificate resources run as pre-install hooks (weight -30) before certgen (weight -20), allowing the hook to poll for and find the issued certificates. Removed the outdated two-step process: 1. helm install (cert-manager issues cert, hook logs warning) 2. helm upgrade (hook creates ConfigMap) Replaced with current single-stage behavior: - Certificate resources created as pre-install hooks - certgen hook polls for up to 90 seconds (configurable) - Single helm install succeeds in most cases - Fails fast by default if timeout reached This brings openshift.mdx in line with the already-updated ingress.mdx documentation. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration The timeout value now represents the actual polling time that users experience when waiting for cert-manager to issue certificates. The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to allow buffer time for ConfigMap creation and cleanup. Previously, the hook polled for (timeoutSeconds - 30) seconds, which was confusing when users set timeoutSeconds=180 and only got 150 seconds of actual polling. Updated documentation in values.yaml, ingress.mdx, and openshift.mdx to reflect the clearer behavior. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(helm): update values.yaml and README with correct polling duration Updated the caCertificateConfigMapName description to reflect that the hook polls for exactly pkiInitJob.timeoutSeconds seconds, not (timeoutSeconds - 30) seconds. Regenerated README.md with helm-docs. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): add OIDC configuration to helm install and CLI examples Updated all helm install and openshell gateway add examples in ingress.mdx and openshift.mdx to include OIDC issuer and audience configuration. Examples now use concrete placeholder values: - OIDC issuer: https://keycloak.example.com/realms/openshell - OIDC audience: openshell-cli - Hostname: gateway.example.com - ClusterIssuer: letsencrypt-prod This makes it clearer how to configure OIDC authentication, which is required when using BackendTLSPolicy or HTTPS termination since the Gateway proxy cannot present client certificates to the backend. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): explicitly list OIDC client ID in gateway add examples Added --oidc-client-id openshell-cli to all openshell gateway add commands in ingress.mdx and openshift.mdx, making the default client ID explicit in the examples even though it's the CLI default. This improves clarity and helps users understand the complete OIDC configuration needed for gateway registration. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * fix(helm): address PR review feedback for BackendTLSPolicy - Read backend CA from the authoritative server Secret instead of the in-memory PKI bundle so enabling BackendTLSPolicy on an existing release uses the CA that actually signed the server certificate. - Reconcile the backend CA ConfigMap on every hook run (compare and update) instead of skipping when it already exists, so CA rotations propagate automatically. - Remove hook annotations from cert-manager Issuer/Certificate resources so they remain regular release objects managed by Helm lifecycle. Split the cert-manager backend CA ConfigMap creation into a separate post-install/post-upgrade hook Job that polls after cert-manager Certificate resources are applied. - Update architecture/gateway.md, docs/reference/gateway-config.mdx, debug-openshell-cluster skill, and helm-dev-environment skill with BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout documentation. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): resolve markdown lint errors in helm README and kubernetes docs Escape inline HTML angle brackets in README.md template placeholders, remove trailing spaces, and add blank lines around fenced code blocks in numbered lists. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Update docs * fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile Wrap `<fullname>` and `<namespace>` template placeholders in backticks so markdownlint does not flag them as inline HTML (MD033). Regenerate mise.lock to match current mise.toml after rebase onto main. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Run 'mise lock' * fix: align mise.lock with CI mise version output The lockfile was regenerated locally with mise 2026.8.10 which resolves uv Linux artifacts to gnu variants and adds provenance_verified fields, but CI uses v2026.4.25 which produces musl variants without those fields. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): correct cert-manager hook ordering and clientCaSecretName comment Update ingress.mdx and openshift.mdx to describe Certificate resources as regular release objects with a post-install/post-upgrade Job, matching the current implementation and architecture/gateway.md. Fix values.yaml clientCaSecretName comment to state that "" disables client certificate verification, matching the helper and access-control docs. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> --------- Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> |
||
|
|
7cc9551677 |
feat(server): support EC and EdDSA keys in OIDC JWKS validation (#2593)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
cc4ded2088 |
feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(helm): preserve split chart upgrade compatibility Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology. * fix(ci): preserve VM runtime for E2E The Rust cache restores target/ after VM runtime artifacts are staged, overwriting target/vm-runtime-compressed before openshell-driver-vm is built. Stage the compressed runtime outside target and pass that location through OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor. Also locate the Helm split-ownership test repository root from the script path rather than git rev-parse. The test runs in a container where the GitHub checkout can be owned by a different UID and rejected as dubious ownership. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(ci): install yq for Helm ownership test The split-chart ownership regression uses yq to inspect rendered YAML, but the Helm CI container installs only tools declared in mise. Declare and lock yq so mise install --locked provides the test dependency. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> --------- Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> |
||
|
|
e508c169ea |
fix(helm): honor empty clientCaSecretName for HTTPS-only mode (#2235)
* fix(helm): honor empty clientCaSecretName for HTTPS-only mode Signed-off-by: Yuedong Wu <dwcn22@outlook.com> * docs(skills): document HTTPS-only clientCaSecretName in debug-openshell-cluster Signed-off-by: Yuedong Wu <dwcn22@outlook.com> --------- Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
9b6d904e88 |
feat(compute): delegate sandbox authentication to drivers (#2968)
* feat(compute): delegate sandbox authentication to drivers Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(kubernetes): align sandbox identity annotation Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(auth): restore sandbox bootstrap coverage Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(kubernetes): satisfy ownership test lint Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
40d1b48666 |
feat(provider): support for SPIFFE backed token exchange (#1970)
* feat(provider): add ability to request token exchange instead of client credentials as OAuth grant_type Signed-off-by: Gordon Sim <gsim@redhat.com> * test(proxy): add further tests for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(provider): add runnable example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover Podman token exchange grants Signed-off-by: Gordon Sim <gsim@redhat.com> * refactor(oauth): extract duplicated functionality from server and supervisor Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): evict nearest-to-expiry entry from intermediate token cache Signed-off-by: Gordon Sim <gsim@redhat.com> * doc(supervisor): add podman example for token exchange Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(provider): withhold token-exchange subject credentials Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
3be2cd8a29 |
fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(kubernetes): share Agent Sandbox setup Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(e2e): wait for Agent Sandbox CRD status Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(canary): sparse-checkout sandbox helper Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
998db04780 |
feat(policy): allow non-root sandbox identities (#2785)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
2eb0880a00 |
feat(cli): support OIDC device authorization grant for headless login (#2795)
* feat(cli): support OIDC device authorization grant for headless login Closes #2793 Add OAuth 2.0 Device Authorization Grant (RFC 8628) support to the OpenShell CLI's OIDC login flow. When running in a headless environment (OPENSHELL_NO_BROWSER=1) without a client secret configured, the CLI now uses the device code flow instead of the browser-based PKCE flow. The device code flow: - Requests a device code and user code from the IdP's device authorization endpoint - Displays a verification URL and user code to the user - Polls the token endpoint until the user completes authorization or the code expires - Supports slow_down responses per RFC 8628 by increasing the polling interval This implementation: - Extends OidcDiscovery to optionally capture device_authorization_endpoint - Adds oidc_device_code_flow function with proper error handling for all RFC 8628 error codes - Updates gateway add and gateway login to dispatch to device flow when browser is suppressed - Adds comprehensive unit tests for device flow structs and response parsing - Updates gateway authentication documentation to describe the device code fallback Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): validate OIDC device token responses Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): add PKCE to OIDC device flow Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(cli): document PKCE device flow Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> |
||
|
|
d0c6dc3fd8 |
feat(kubernetes): support corporate upstream proxy (#2633)
* feat(kubernetes): support corporate upstream proxy Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * fix(kubernetes): reject proxy_auth_secret_key values Kubernetes cannot create Gateway validation accepted proxy_auth_secret_key values that Kubernetes rejects when creating the Secret (keys longer than 253 bytes, or the reserved "."/".." names), turning an invalid deployment setting into repeated sandbox Pod-provisioning failures instead of a startup error. Reject them in validate_upstream_proxy_config so they fail closed at gateway startup. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> * docs(skill): add corporate upstream proxy checks to debug-openshell-cluster Add a Kubernetes corporate upstream proxy troubleshooting section covering rendered [openshell.drivers.kubernetes] configuration, credential Secret volume events, supervisor arguments and mounts confined to the network- supervising container, and proxy reachability. Signed-off-by: loveRhythm1990 <qiuweimin@126.com> --------- Signed-off-by: loveRhythm1990 <qiuweimin@126.com> |
||
|
|
c4b500a7de |
feat(helm): cert-manager external issuer + OpenShift passthrough Route (#2468)
* fix(core): trust public root CAs alongside the sandbox mTLS CA The supervisor gRPC client only trusted the CA configured via OPENSHELL_TLS_CA, since tonic ClientTlsConfig starts with an empty root store unless with_native_roots()/with_webpki_roots() is also enabled. Deployments where the gateway server certificate is issued by a public CA (e.g. cert-manager against an ACME issuer) caused every supervisor connection to fail the TLS handshake with "UnknownCA", since the sandbox mTLS CA and the server cert issuer were no longer the same. Enable both native and webpki roots in addition to the configured CA. tonic root store is a union of all configured sources, so this does not weaken verification for existing self-signed deployments. webpki-roots (compiled in) is enabled alongside native-roots since the supervisor binary may run in minimal sandbox images without a populated system CA bundle. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(helm): support external cert-manager issuers and OpenShift Route passthrough Add certManager.serverIssuerRef/clientIssuerRef so the gateway and mTLS client certificates can be issued by a real Issuer/ClusterIssuer (e.g. ACME) instead of only the chart built-in self-signed CA. Add openshiftRoute template for exposing the gateway via a TLS passthrough Route so the gateway keeps terminating its own TLS/mTLS. The server Certificate excludes internal-only SANs (cluster-local, localhost, loopback) when an external issuer is configured, since ACME issuers reject those per CA/Browser Forum baseline requirements. A template-time fail guard catches the misconfiguration at helm install time rather than asynchronously at cert-manager issuance time. Includes Helm unittest coverage for both issuerRef overrides and Route rendering, plus a CI values overlay for lint coverage. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs: document cert-manager external issuer and OpenShift Route Update managing-certificates.mdx with the serverIssuerRef workflow and install-time validation behavior. Add a production section to the OpenShift guide covering passthrough Route with a real certificate. Regenerate Helm README for new certManager and openshiftRoute values. Sync debug-openshell-cluster skill with new troubleshooting steps for ACME issuance failures and supervisor UnknownCA from mismatched CAs. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,core): address PR review feedback on cert-manager external issuer Addresses all five blocking review items from #2468: 1. Remove .with_native_roots() from supervisor gRPC client -- the supervisor runs inside the user-selected sandbox image, so the image CA bundle is not operator-controlled. Keep .with_webpki_roots() (compiled-in, not user-controlled) alongside the configured CA. 2. Fail at render time when serverIssuerRef.name is set but clientCaFromServerTlsSecret is still true. Add negative Helm test. 3. Remove clientIssuerRef -- changing only clientIssuerRef breaks both directions because trust bundles are not modeled separately. Change serverIssuerRef.kind default from ClusterIssuer to Issuer. 4. Add server.oidc.issuer and server.oidc.audience to the documented OpenShift production Helm command. Add Access Control prerequisite. 5. Fail at render time when openshiftRoute.enabled and disableTls are both true. Add negative Helm test. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME from Docker and Podman env Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(helm,drivers): guard default clientCaSecretName and add env-strip tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * feat(tls): SNI-based dual certificate for internal and external server TLS Split the gateway server certificate into two: an internal cert issued by the chart's own CA (for supervisor connections via cluster-local SANs) and an external cert issued by an operator-configured Issuer such as ACME/Let's Encrypt (for CLI and Route access via public SANs). The gateway uses SNI-based certificate selection: connections whose SNI hostname matches external_server_names receive the external cert; all others (including those with no SNI) receive the internal cert. Security improvement: remove .with_webpki_roots() from the supervisor gRPC client so supervisors trust only the chart CA, closing a MITM vector via publicly-trusted certificates in user-supplied container images. Key changes: - Add DualCertResolver with SNI-based cert selection and full test coverage - Add external_cert_path, external_key_path, external_server_names to TlsConfig - Validate partial external cert config (error on cert-without-key or vice versa) - Validate empty external_server_names when external cert is configured - Split cert-manager templates into internal + external Certificate resources - Add Helm guards for misconfigured external issuer (empty serverDnsNames, internal-only SANs with external issuer, conflicting clientCaFromServerTlsSecret) - Update gateway-config.mdx, managing-certificates.mdx, openshift.mdx docs - Update debug-openshell-cluster skill for dual-cert troubleshooting Signed-off-by: Pi Agent <agent@openshell.local> * fix(drivers): strip GATEWAY_TLS_SERVER_NAME in VM driver and correct comments Add the same GATEWAY_TLS_SERVER_NAME environment stripping to the VM compute driver that Docker, Podman, and Kubernetes drivers already perform. Without this, a sandbox user on the VM driver could override the TLS server name the supervisor verifies. Fix stale comments in Docker and Podman drivers that referenced 'with WebPKI roots trusted' — WebPKI roots are explicitly not trusted after the tls-webpki-roots removal. Use tls-ring instead of bare channel for tonic in openshell-core so the TLS API (ClientTlsConfig, Endpoint::tls_config) is available without pulling in any root certificate store. Signed-off-by: Pi Agent <agent@openshell.local> * fix(tls,helm): wildcard SNI matching and Route host validation Add RFC 6125 single-level wildcard matching to DualCertResolver so external_server_names entries like *.example.com correctly match SNI hostnames like gw.example.com. Previously only exact matches worked, silently falling back to the internal cert for wildcard configurations. Add a Helm fail guard in route.yaml that rejects openshiftRoute.host values not listed in certManager.serverDnsNames when an external issuer is configured — catches cert/route hostname mismatches at install time instead of at TLS connect time. Quote the host field in route.yaml for robustness. Signed-off-by: Pi Agent <agent@openshell.local> * fix(helm): address blocking review items — client-CA guard, wildcard Route, serverIssuerRef gate 1. Remove the obsolete guard rejecting serverIssuerRef + clientCaFromServerTlsSecret=true. The internal server certificate is always signed by the chart CA (the same CA that signs the client cert), so clientCaFromServerTlsSecret=true is correct — its filtered ca.crt is exactly the right trust anchor. The old workaround (mounting openshell-ca-tls directly) unnecessarily exposed the CA private key to the gateway container. Remove the client-CA overrides from docs, CI overlay, and production examples. 2. Route host validation now supports wildcard certificates per RFC 6125: single-level wildcards like *.example.com match gateway.example.com but not deep.sub.example.com. Require an explicit openshiftRoute.host when an external issuer is configured — without one, OpenShift generates a hostname absent from serverDnsNames. 3. Reject serverIssuerRef.name when certManager.enabled is false — the external certificate, its Secret mount, and the gateway TLS config all require cert-manager to be enabled. Validated on ROSA (dev.dyee.p3) with branch-built images: - Fresh install with letsencrypt-prod ClusterIssuer - SNI dual-cert: external hostname served Let's Encrypt cert - Supervisor mTLS via internal cert path: ConnectSupervisor accepted - Client CA volume: filtered ca.crt from internal server secret (no key) - CLI connected via Route + OIDC Helm tests: 81 pass across 7 suites. Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Pi Agent <agent@openshell.local> |
||
|
|
170961997f |
chore(ci): disable telemetry in internal test runs (#2648)
* chore(ci): disable telemetry in internal test runs Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(ci): remove brittle telemetry wiring test Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * docs: trim CI telemetry guidance Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> * test(e2e): share telemetry default with OpenShift Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> --------- Signed-off-by: Matthew Grossman <mgrossman@nvidia.com> |
||
|
|
8eacb4779f |
feat(kubernetes): add sidecar supervisor topology (#2076)
* feat(kubernetes): add sidecar supervisor topology Add the Kubernetes sidecar supervisor topology, its Helm/Skaffold configuration, topology documentation, and sidecar e2e matrix coverage. Skip root-only sandbox identity rewriting when process enforcement is network-only so the low-permission sidecar process container can start successfully. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar process id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): avoid similar proxy id names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify sidecar topology limits Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): keep sidecar process leaf capless Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): refresh sidecar provider env snapshots Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * test(supervisor): align hot-swap identity regression Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): stage sidecar mtls files before proxy chown Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): simplify sidecar supervisor topology Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(helm): reuse sidecar skaffold values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid similar iptables helper names Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(e2e): harden kube gateway wrapper setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(supervisor): avoid nft batch rollback on OCP Run nftables setup as individual commands so optional conntrack and log expressions can fail without rolling back required table, chain, and reject rules. Signed-off-by: Seth Jennings <sjenning@redhat.com> Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): preserve process identity in sidecar topology Render sidecar pods with a shared process namespace, keep binary-aware network policy enabled, and move Kubernetes sidecar settings under the nested sidecar config table. Also apply unprivileged Landlock/seccomp setup in NetworkOnly supervisor mode so sidecar topology keeps sandbox child hardening without privileged process setup. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): replace sidecar snapshots with control socket Coordinate sidecar policy and provider bootstrap over a local Unix socket so the process leaf no longer reads policy/provider snapshot files. Report entrypoint startup through the control channel and keep gateway credentials confined to the network sidecar. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * feat(kubernetes): support relaxed sidecar network identity Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy sidecar clippy lint Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): standardize topology naming Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(sandbox): satisfy linux clippy timeout import Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): support kata sidecar on ipv4 pods Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): satisfy linux clippy for sidecar fallback Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * chore(kubernetes): remove stale supervisor topology references Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): enable sidecar binary policy inspection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden sidecar control boundary Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): couple sidecar supervisor lifecycles Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
8c0ecac8cf |
docs(openshift): simplify install steps and add Helm README entries for OpenShift overrides (#2125)
Signed-off-by: ChristianZaccaria <christian.zaccaria.cz@gmail.com> |
||
|
|
290297ffa3 | docs(kubernetes): bump cert-manager to v1.20.3 (#2129) | ||
|
|
abcd15d1f2 |
feat(helm): add TLS termination for Envoy Gateway ingress (#2015)
The chart's optional Gateway API ingress only rendered a plaintext HTTP listener, so the gateway could not be exposed over TLS. Add an HTTPS listener option that terminates TLS at the Envoy Gateway and forwards plaintext gRPC to the gateway pod. - gateway.yaml renders an HTTPS listener with `tls.mode: Terminate` and `certificateRefs` when `grpcRoute.gateway.listener.protocol=HTTPS`, keeping the default HTTP listener unchanged. Guards fail the render when `certificateRefs` is empty or `server.disableTls` is not true (the chart does not render a BackendTLSPolicy for re-encryption). - values.yaml adds `grpcRoute.gateway.listener.tls.certificateRefs`. - ci/values-gateway-tls.yaml exercises the HTTPS branch in lint/render. - docs/kubernetes/ingress.mdx documents HTTPS setup and clarifies that Envoy Gateway only terminates TLS (no OIDC SecurityPolicy); client identity uses OIDC bearer tokens, with the client-credentials grant for headless agents. - debug-openshell-cluster skill gains HTTPS-ingress troubleshooting rows. - Regenerated the chart README values table. Signed-off-by: Huabing (Robin) Zhao <zhaohuabing@gmail.com> |
||
|
|
914da339b4 |
feat(kubernetes): add combined topology config surface (#2074)
* feat(kubernetes): add combined topology config surface Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): clarify topology defaults Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
8cb16de9ea | chore(deploy): use OCI registry for cert-manager Helm chart (#2041) | ||
|
|
ba21bb32a2 |
feat(kubernetes): support agent-sandbox v1beta1 (#2009)
* feat(kubernetes): support agent-sandbox v1beta1 Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * ci(kubernetes): test agent-sandbox api versions Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): retry agent-sandbox raw 404s Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): harden agent-sandbox api setup Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(kubernetes): cache agent-sandbox api version Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * docs(kubernetes): document agent-sandbox upgrade behavior Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * refactor(kubernetes): collapse agent-sandbox api selection Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
62b03f0058 |
fix(docs): add step for creating the GatewayClass (#1984)
Signed-off-by: Huabing (Robin) Zhao <zhaohuabing@gmail.com> |
||
|
|
7dab612feb |
feat(helm): support Deployment kind in HA gateway workloads (#1867)
* feat(helm): support Deployment kind in HA gateway workloads Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(helm): handle null workload values Signed-off-by: Taylor Mutch <taylormutch@gmail.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
702cbc4f63 |
feat(providers): support SPIFFE-backed token grants (#1784)
* feat(providers): support SPIFFE-backed token grants Add provider profile token_grant metadata and expand endpoint-specific dynamic credentials so sandbox supervisors can request SPIFFE JWT-SVIDs, exchange them with an OAuth-style token endpoint, cache returned access tokens, and inject bearer tokens into matching HTTP requests. Wire Kubernetes and Helm deployments to mount the provider SPIFFE Workload API socket into sandbox pods for token grant exchange. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> * test(examples): add SPIFFE token grant demo Add a reusable alpha/beta demo that deploys a SPIFFE-verifying token issuer and protected services, imports a token-grant provider profile, creates a sandbox, and verifies endpoint-specific bearer tokens. The script leaves Kubernetes workloads in place, deletes sandboxes through openshell unless KEEP_SANDBOX=1, and prints protected service logs as proof of life. Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden SPIFFE token grants * fix(providers): harden dynamic token grants Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(providers): harden token grant handling --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
c4ca283c1a | refactor(helm): require external postgres for ha (#1844) | ||
|
|
e26a1b1ffc |
fix(kubernetes): configure sandbox apparmor profile (#1767)
Signed-off-by: Taylor Mutch <taylormutch@gmail.com> |
||
|
|
586c385bdf |
chore(k8s): use upstream agent-sandbox manifest in CI/e2e (#1657)
Drop the vendored deploy/kube/manifests/agent-sandbox.yaml. CI and e2e scripts now apply manifest.yaml directly from github.com/kubernetes-sigs/agent-sandbox releases, pinned via AGENT_SANDBOX_VERSION env (default v0.4.6, overridable) so internal runs stay reproducible. Public docs and the helm chart README continue to reference /releases/latest/download/ — no OpenShell/agent-sandbox support matrix exists yet (see #1649). Adds an air-gap note in docs/kubernetes/setup.mdx enumerating the manifest and image operators need to mirror to an internal registry. Signed-off-by: Roshni Malani <rmalani@nvidia.com> |
||
|
|
5f58cb018f |
fix(helm): create sandbox JWT secret when cert-manager is enabled (#1700)
* fix(helm): create sandbox JWT secret under cert-manager The cert-manager install path (certManager.enabled=true, pkiInitJob.enabled=false) left the gateway StatefulSet unable to start because nothing created the openshell-jwt-keys Secret: cert-manager owns TLS Secrets but does not mint the sandbox JWT signing key, and the certgen hook only rendered when pkiInitJob.enabled was true. Separate JWT signing-key provisioning from TLS PKI provisioning: - certgen: add a --jwt-only mode that creates only the Opaque JWT signing Secret, for use when another controller owns TLS Secrets. - certgen.yaml: render the hook when pkiInitJob.enabled OR certManager.enabled is true. cert-manager takes precedence and runs the hook with --jwt-only even if pkiInitJob.enabled remains true. Remove the mutual-exclusion failure between the two values. - _helpers.tpl: add openshell.sandboxJwtSecretName, shared by the hook and the StatefulSet mount. - Update values, README, docs, architecture, and the debug-openshell-cluster skill to reflect the new precedence; the documented cert-manager install no longer needs pkiInitJob.enabled=false. Closes #1691 * fix(helm): honor cert-manager precedence for client CA volume The client CA volume logic treated pkiInitJob.enabled as proof that built-in PKI owns the client CA. With cert-manager precedence now allowing certManager.enabled=true alongside the default pkiInitJob.enabled=true, that assumption mounts the server TLS cert secret as the client CA and ignores certManager.clientCaFromServerTlsSecret=false, which can break mTLS or trust the wrong CA. Gate the pkiInitJob.enabled term with (not certManager.enabled) in all three client CA conditions (volume mount, volume definition, and secret selection) so cert-manager owns TLS when enabled. Add a Helm test suite covering built-in PKI, cert-manager shared CA, the regression config (cert-manager + clientCaFromServerTlsSecret=false + default pkiInitJob), and the no-client-CA case. |
||
|
|
d9908222f2 | feat(kubernetes): support sandbox image pull secrets (#1671) | ||
|
|
2bdc968ed9 |
fix(gateway): make readiness health checks dependency-aware (#1328)
* feat(gateway): add readiness probe metrics and test-only store close Emit Prometheus readiness metrics for database probes (healthy gauge and outcome-labeled latency histogram) with coverage in health HTTP tests. Restrict Store::close behind test support cfg to prevent accidental runtime pool shutdown under live traffic. Signed-off-by: Adrien Langou <alangou@nvidia.com> * test(e2e): add simple e2e test with kubernetes to test /readyz Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
fafde3e1b3 |
docs(kubernetes): add RBAC section to setup page (#1540)
Documents the ServiceAccount, Role, and ClusterRole created by the Helm chart inline on the setup page, per reviewer feedback on #1250. Reflects the current chart templates including pods/get for sandbox identity and tokenreviews/create for projected token validation. Closes #1018 |
||
|
|
a3b16c18ab | feat(auth): per-sandbox authentication to gateway (#1404) | ||
|
|
b61a98dbad |
feat(gateway): add TOML configuration file (RFC 0003) (#1317)
* feat(gateway): add TOML configuration file (RFC 0003)
Introduces an opt-in --config / OPENSHELL_GATEWAY_CONFIG flag that loads a
TOML file with gateway-wide settings and per-driver tables. Source
precedence is CLI > env > file > built-in default, implemented via clap's
ValueSource so existing flags and env vars keep their priority.
Driver crates (kubernetes, docker, podman, vm) now derive Deserialize on
their config structs. SupervisorSideloadMethod gains Deserialize with
kebab-case rename. A per-driver inheritance allowlist on the loader side
overlays [openshell.gateway] shared defaults (default_image,
supervisor_image, image_pull_policy, guest_tls_*, ssh_handshake_skew_secs,
client_tls_secret_name, host_gateway_ip, enable_user_namespaces) onto
each [openshell.drivers.<name>] table before deserialization.
The Helm chart renders a new gateway-config ConfigMap and mounts it at
/etc/openshell/gateway.toml. The migrated OPENSHELL_* env entries are
dropped from the StatefulSet — only the Secret-backed
OPENSHELL_SSH_HANDSHAKE_SECRET remains. database_url stays on --db-url.
Adds examples/gateway/gateway.example.toml and updates architecture/gateway.md
with the source precedence and inheritance rules.
* docs(gateway): drop ssh_handshake_skew_secs and ssh_handshake_secret from examples
Both fields are scheduled for removal. Remove the example values and the
env-only note so the gateway.toml example and the architecture doc stop
recommending settings that will not exist much longer.
* docs(rfc): correct OPENSHELL_CONFIG to OPENSHELL_GATEWAY_CONFIG in RFC 0003
* docs(gateway): add per-driver TOML example configurations
Adds focused single-driver examples next to the comprehensive
gateway.example.toml: kubernetes, docker, podman, and microvm. Each one
demonstrates the realistic settings for that driver plus how shared
[openshell.gateway] defaults inherit into the driver table.
A new unit test (`checked_in_examples_parse`) loads every example through
the config_file loader so schema drift fails CI rather than silently
shipping a broken example.
* refactor(gateway): drop image_pull_policy from shared inheritance
Kubernetes and Podman use mutually-incompatible vocabularies for the same
TOML key:
- Kubernetes: `Always | IfNotPresent | Never` (free-form string passed
verbatim to the K8s API).
- Podman: `always | missing | never | newer` (strict lowercase enum
deserialised into `ImagePullPolicy`).
No value means the same thing in both drivers. Sharing the key at
`[openshell.gateway]` scope and inheriting it into every active driver's
table meant any value safe for one driver was either wrong or silently
dropped for the other (`IfNotPresent` → `ImagePullPolicy::Missing` after
`.unwrap_or_default()`). Operators run one driver per gateway, so the
"shared default" never pays for itself.
Make `image_pull_policy` driver-local:
- Remove the field from `GatewayFileSection` and from
`inheritable_keys()` for both Kubernetes and Podman.
- Drop the file→`RunArgs` merge for the gateway-scope key.
- Stop unconditionally clobbering the driver value with
`config.sandbox_image_pull_policy` in the runtime wiring — only apply
the CLI/env override when it was set (and, for Podman, only when it
parses into the lowercase enum).
- Move the key under `[openshell.drivers.kubernetes]` and
`[openshell.drivers.podman]` in every example, the RFC, the
architecture doc, and the Helm-rendered gateway ConfigMap.
The supervisor pull policy follows the same shape: it is K8s-only and
moves into `[openshell.drivers.kubernetes]` alongside `image_pull_policy`
in the Helm template.
* fix(gateway): address review feedback on TOML configuration
Resolves the P1 and P2 issues raised in PR #1317:
- Helm gateway ConfigMap moves `grpc_endpoint` under
`[openshell.drivers.kubernetes]` so the default install no longer fails
the gateway's `deny_unknown_fields` schema check.
- `kubernetes_config_from_file` and `podman_config_from_file` only let
the gateway-wide CLI/env `grpc_endpoint` overwrite the driver-table
value when it was actually supplied, preserving file-only configs.
- Kubernetes driver default `image_pull_policy` is now empty (was Podman
vocabulary "missing"), so default deployments let the Kubernetes API
apply its own policy instead of being rejected.
- New `disable_tls` gateway field plumbs `.Values.server.disableTls`
through the TOML ConfigMap instead of relying on env vars dropped from
the StatefulSet.
- StatefulSet pod template now carries a `checksum/gateway-config`
annotation so `helm upgrade` rolls pods when the ConfigMap changes.
- Auxiliary listener resolution preserves the full `SocketAddr` from
`health_bind_address` / `metrics_bind_address`, so a loopback-pinned
health port is not silently relocated onto the public bind address.
- `ssh_session_ttl_secs` from the file is now applied to `Config` (it
was previously accepted by the loader but never read).
New regression coverage: cli-level merge tests for the new fields plus
helm-unittest assertions for the ConfigMap shape, checksum annotation,
and `disable_tls` rendering.
* docs(gateway): consolidate gateway TOML examples into docs reference
Replaces the per-driver example files under examples/gateway/ with a
single published reference page at docs/reference/gateway-config.mdx
covering source precedence, layout, the full example, and the four
per-driver examples (Kubernetes, Docker, Podman, microVM). Drew flagged
during PR #1317 review that the examples belong with the user-facing
docs rather than in a sibling examples/ directory.
The cross-references in architecture/gateway.md and RFC 0003 are updated
to point at the new docs page; the round-trip test in config_file.rs is
removed (schema coverage stays on the inline parses_full_example test
and per-field merge tests — doc-snippet drift belongs in a separate
docs-lint, not in a cross-tree Rust unit test).
* refactor(core): move DEFAULT_K8S_NAMESPACE into K8s driver
The constant is Kubernetes-specific (used only by KubernetesComputeConfig's
Default impl) and does not belong in openshell-core. Relocate it to the
driver crate that owns the K8s vocabulary; openshell-core retains only
truly cross-cutting defaults.
* refactor(core): move Podman bridge default into Podman driver
DEFAULT_NETWORK_NAME is Podman vocabulary, consumed only by the Podman
driver. Also drops the unused DEFAULT_IMAGE_PULL_POLICY constant.
* docs(auth): scrub remaining SSH handshake secret references
Sweeps the trailing mentions left after the rebase: the gateway
config-file module doc, the Helm gateway-config ConfigMap header,
the gateway-config.mdx env-only note, and the RPM systemd unit
comment for init-gateway-env.sh.
* docs(gateway): clarify OPENSHELL_GRPC_ENDPOINT applies to all drivers
The previous comment implied the callback endpoint was Kubernetes-only,
but the value is propagated to every compute driver (Kubernetes, Docker,
Podman, VM) and must be reachable from wherever the sandbox runs.
* refactor(gateway): move driver options into config (#1394)
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): regenerate gateway config via TOML for docker + podman harnesses
The gateway CLI flags moved into TOML config tables in 560550d2 (#1394),
which made every existing e2e/with-{docker,podman}-gateway.sh invocation
fail with "unexpected argument '--sandbox-namespace'" (and a long tail
of similar driver-specific options) before the gateway could even bind.
Replace the obsolete CLI flags with a synthesized
`[openshell.drivers.<driver>]` table written to `${STATE_DIR}/gateway.toml`
and passed via the new `--config` flag. Only the gateway-wide flags
that survived 560550d2 (bind-address, port, drivers, db-url, tls-*,
disable-tls, log-level, health-port) stay on the command line.
Both scripts get a small `toml_string` helper to properly TOML-quote
the values (the previous `%q` printf format produced bash-escape, not
TOML-escape). The Docker harness also corrects two field names that
diverged from the driver schema: `docker_network_name` →
`network_name`, and the supervisor binary/image plumbing now reads
through to `supervisor_bin` / `supervisor_image` in the same table.
The Podman harness drops `--ssh-gateway-port` (deleted in 560550d2 —
gRPC + SSH are multiplexed on the same port now) and substitutes
`network_name` + `gateway_port` for the obsolete `--sandbox-namespace`
(which the Podman driver never had as a typed field).
* fix(core): swap bind-only 0.0.0.0 SSH gateway host for cluster URL host
CLI's resolve_ssh_gateway treated 0.0.0.0 as a loopback "keep as-is"
when the cluster URL was also loopback, so the SSH proxy connected to
0.0.0.0:port. The unspecified address is never a valid connect target
and is not present in any TLS cert SAN, which produced BadCertificate
TLS handshake failures during `openshell sandbox create -- ...` in
docker/podman e2e (e.g. bypass_detection).
Resolution: when the server returns 0.0.0.0 or :: as the gateway host
and both endpoints are loopback, fall back to the cluster URL's host
(which the CLI is already using to reach the gateway, so it must
resolve and match the cert).
* fix(e2e): repair podman harness on macOS
Podman 5.x with the applehv/libkrun provider no longer creates the legacy
~/.local/share/containers/podman/machine/podman.sock symlink, and
`podman system service` is a Linux-only subcommand — the macOS client
delegates the API service to the VM. Both assumptions in the harness
were stale, so the script tried to start a temporary service that podman
rejected with "unknown flag: --time".
- Discover the macOS socket via `podman machine inspect` instead of the
hardcoded path.
- On Darwin, fail fast with a "start podman machine" message rather than
attempting the Linux-only `podman system service` fallback.
- Write socket_path into [openshell.drivers.podman] so the in-process
driver picks up the discovered socket; the driver reads TOML only
after the config refactor (560550d2), so OPENSHELL_PODMAN_SOCKET alone
was no longer enough.
* fix(server): use clone_from for TLS client CA assignment
clippy 1.95.0 rejects assigning the result of `Clone::clone()` to an
existing variable under `-D warnings` (`assigning_clones`). Switch to
`clone_from(&...)` to satisfy the lint and avoid the redundant
allocation.
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
|
||
|
|
7a0c444445 |
refactor!(auth): drop SSH handshake secret (#1274)
* refactor!(auth): drop SSH handshake secret in favor of mTLS The OPENSHELL_SSH_HANDSHAKE_SECRET / x-sandbox-secret mechanism was misnamed: it does not authenticate SSH (which flows over the RelayStream gRPC RPC and is gated by mTLS plus supervisor Unix-socket permissions). It only gated a small set of sandbox-to-gateway control-plane RPCs, and production deployments already enforce mTLS on that channel — so the shared secret was redundant. Replace the secret check with an mTLS-presence marker. Sandbox-class methods (ReportPolicyStatus, PushSandboxLogs, GetSandboxProviderEnvironment, SubmitPolicyAnalysis, GetSandboxConfig, GetInferenceBundle) accept callers without a Bearer token; the gRPC mTLS handshake is the trust boundary. Dual-auth methods treat Bearer-present as full-scope CLI access and Bearer-absent as sandbox-restricted scope via validate_sandbox_caller_update. Drops the secret from all drivers (K8s, Podman, VM), the sandbox gRPC interceptor, the Helm chart (values + pre-install hook + StatefulSet env), the RPM bootstrap script, the man pages, and the debug-openshell-cluster skill. Also removes the never-read ssh_handshake_skew_secs flag and config field. BREAKING CHANGE: --ssh-handshake-secret / OPENSHELL_SSH_HANDSHAKE_SECRET and --ssh-handshake-skew-secs / OPENSHELL_SSH_HANDSHAKE_SKEW_SECS are removed from the gateway, sandbox, and all driver binaries. The openshell-ssh-handshake K8s Secret is no longer managed by the chart; operators may delete the orphan. Deployments using --disable-gateway-auth must enforce caller authentication at the fronting proxy, since the gateway no longer validates a per-request secret on sandbox-class methods. Refs OS-174. * docs(auth): scrub residual SSH handshake secret references Sweep across docs, e2e scripts, the Podman driver README/NETWORKING notes, the RPM/Helm/setup guides, the gateway man page, and RFC 0003 to remove instructions and examples that still referenced OPENSHELL_SSH_HANDSHAKE_SECRET / --ssh-handshake-secret / ssh_handshake_skew_secs. The mechanism is gone; nothing should still suggest setting it. Negative-assertion regression tests are kept so the env var cannot silently be re-introduced. * test(drivers): drop SSH handshake secret negative-assertion tests The supporting code, env vars, CLI flags, and config plumbing are gone — these tests asserted absence of strings that no longer have any path to being set. Remove the guards from the Docker, Podman, Kubernetes, and VM driver test modules. |