Commit Graph
22 Commits
Author SHA1 Message Date
pkhodade-NVandShailendra Singh 9e397e8a43 fix(mxc): reject cpu/memory limits instead of silently discarding them (#3548)
* fix(mxc): reject cpu/memory limits instead of silently discarding them

CreateSandbox accepted --cpu/--memory and reached Ready with no Job
Object enforcement and no diagnostic, leaving the SDD's T11
host-exhaustion mitigation silently unmet. MXC's schema does not
expose CPU rate control or memory limiting outside the WSLC backend,
so reject requests carrying cpu/memory limits synchronously at
CreateSandbox, matching the existing fail-closed GPU rejection.

NVBug 6782894

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>

* fix(server): validate sandbox resource quantities

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* docs(skill): clarify MXC resource limit behavior

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* chore(go): regenerate protobuf bindings

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

* fix(go): align generated protobuf comments

Signed-off-by: Shailendra Singh <shailendras@nvidia.com>

---------

Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Co-authored-by: Shailendra Singh <shailendras@nvidia.com>
2026-09-23 11:16:54 -07:00
Prekshi Vyas b031dc037a fix(mxc): reject unsupported live policy updates (#3480)
* fix(mxc): reject unsupported live policy updates (NVBug 6782891)

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): gate all live policy mutations

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(mxc): gate composed policy mutations

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>

* fix(ci): satisfy provider update lint

* fix(ci): order provider validation branches

* test(mxc): make policy synchronization deterministic

* fix(server): scope MXC policy synchronization

* fix(server): serialize provider-backed sandbox creation

* test(server): use valid provider create fixture name

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
2026-09-22 15:44:44 -07:00
Drew Newberry 49b4f0eb7f feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-17 12:06:06 -07:00
Seth Jennings dbe36eaf85 fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002.

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-09-15 19:56:09 +00:00
John T. Myers 607db99915 fix(deps): refresh gateway Debian runtime image (#3350)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-15 19:38:13 +00:00
John T. MyersandJohn Myers dfd5238d0d fix(gator): make supervised lifecycle sandbox-native (#3343)
* fix(gator): run supervisor as sandbox main process

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* feat(gator): persist supervised state history

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* refactor(gator): remove obsolete background launch mode

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-15 19:04:13 +00:00
Mrunal Patel b799fccb8b fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(e2e): pass OIDC HTTP acknowledgement value

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-15 17:00:45 +00:00
Shiju fd3fd9cf74 feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection
results together in sandbox status. Keep sandbox lifecycle readiness
separate so an external connection failure does not mark the sandbox unready.

Expose direct endpoint records through the CLI and SDKs, with plain-language
failure explanations and gateway acceptance times. Keep observation tracking,
runtime reporting, and gateway validation in dedicated endpoint status modules.

Preserve bounded reporting, request attribution, retry ordering, and
configuration and supervisor authority checks. Clear obsolete observations
while retaining the configured addresses, and document the distinction
between an observed HTTP response, current availability, and tool success.

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-15 13:46:38 +00:00
Jorge 42e9bcf2b0 feat(e2e): support the Vault credential-driver lane on OpenShift (#3312)
Implements https://github.com/NVIDIA/OpenShell/issues/3212

Running e2e:kubernetes with the Vault credential driver failed on
OpenShift in two ways: the OpenBao fixture pod was rejected by the
restricted-v2 SCC, and provider-creating tests hit HTTP 403 from
auth/kubernetes/login because OpenBao's Kubernetes auth was provisioned
only inside a single feature-gated test.

- Deploy OpenBao with the chart's OpenShift mode (global.openshift=true)
  when OpenShift is detected, so the pod inherits a namespace-assigned,
  SCC-compliant security context with no manual SCC grant. Hoist
  OpenShift detection ahead of the credential-driver fixtures so the
  flag is set before the fixture is deployed.
- Provision the OpenBao KV store, Kubernetes auth method, storage
  policy, and gateway login role in the harness (deploy_vault_fixture),
  making a Vault-backed gateway usable by the whole suite instead of
  only the credential_drivers test. Remove the now-redundant
  configure_vault_storage helper from the test.
- Harden openbao_exec so it tolerates only the idempotent "path is
  already in use" error on reruns and fails fast with output on any
  other error, instead of a blanket `|| true` that masked genuine
  failures (e.g. an unresponsive pod) until a later cryptic write.
- Document the OpenShift Vault credential-store SCC and Kubernetes-auth
  403 troubleshooting in the debug-openshell-cluster skill.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-14 21:55:43 +00:00
Brandon Squizzato cc780d4e17 feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS

Add grpcRoute.backendTLSPolicy values to optionally create a
BackendTLSPolicy resource that enables end-to-end TLS between the
Gateway proxy and the OpenShell gateway pod. The Gateway proxy
terminates client-facing TLS and re-encrypts when connecting to the
backend, validating the pod's certificate against a user-supplied CA
ConfigMap.

This removes the requirement to set server.disableTls=true when using
HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other
platforms with BackendTLSPolicy support in the Gateway API
implementation.

Update OpenShift and ingress documentation with e2e TLS instructions.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm,server): auto-create backend CA ConfigMap in certgen hook

Extend the generate-certs command with --backend-ca-configmap-name and
--backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the
certgen pre-install hook creates the CA ConfigMap automatically:

- pkiInitJob mode (default): uses the CA from the generated PKI bundle.
  Fully automatic on first install.
- cert-manager mode: reads ca.crt from the server TLS Secret. On first
  install the Secret does not exist yet (cert-manager reconciles after
  templates are applied), so the ConfigMap is created on the first helm
  upgrade. Logs a warning on the initial skip.

The caCertificateConfigMapName value now defaults to <fullname>-backend-ca
when empty, so users only need to set backendTLSPolicy.enabled=true.

Update certgen RBAC to include configmaps get/create when the feature is
enabled. Add CLI arg parsing tests for the new flags.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* refactor(helm): add server.tls.enableMtls flag for mTLS control

Replace automatic mTLS disabling based on BackendTLSPolicy with an
explicit server.tls.enableMtls flag that defaults to true. The user is
now responsible for setting this to false when using BackendTLSPolicy,
as ingress proxies cannot present client certificates to backends.

Updated:
- values.yaml: Added server.tls.enableMtls (default true)
- gateway-config.yaml: Check enableMtls instead of backendTLSPolicy
- _gateway-workload.tpl: Check enableMtls for client CA mount
- Tests: Updated to use enableMtls flag
- Docs: Added enableMtls=false to BackendTLSPolicy examples
- README: Document new flag and BackendTLSPolicy requirement

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(helm): clarify cert-manager backend CA ConfigMap workflow

Update documentation to explain the two-step install process required
when using cert-manager with BackendTLSPolicy:

1. helm install - cert-manager issues the server certificate, but the
   certgen hook can't create the backend CA ConfigMap yet (cert-manager
   reconciles after templates are applied)
2. helm upgrade - certgen hook reads the CA from the cert-manager-issued
   certificate and creates the ConfigMap

Previously, the docs said "created on first upgrade" without explaining
why or that the feature won't work until then. The updated docs now:
- Explain the timing issue (cert-manager reconciles after chart install)
- Provide clear steps for the cert-manager workflow
- Note that pkiInitJob (default) creates it immediately on install
- Clarify that users must wait for the Certificate to be Ready before
  running the second upgrade

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy

BackendTLSPolicy validates the backend certificate against the service FQDN
(e.g., openshell.openshell.svc.cluster.local), not the external hostname.
The external hostname only needs to be on the Gateway listener certificate
for client-facing TLS.

The default certManager.serverDnsNames already includes all required service
FQDN variants, so no configuration is needed for BackendTLSPolicy to work.

Fixed incorrect documentation that claimed:
- "The server certificate SAN list must include the external hostname"
- Users need to "configure certManager.serverDnsNames with the external hostname"

Removed the unnecessary pkiInitJob.serverDnsNames override from the example
and clarified that:
- Gateway listener certificate needs the external hostname (for clients)
- Backend certificate needs the service FQDN (for Gateway proxy)
- The service FQDN is already in the defaults

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs: clarify ACME with LetsEncrypt reference

Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer"
to help users understand that LetsEncrypt is the most common ACME
provider and what ACME means in practice.

Updated:
- docs/kubernetes/managing-certificates.mdx
- docs/kubernetes/openshift.mdx
- deploy/helm/openshell/values.yaml
- deploy/helm/openshell/README.md
- deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname

Reorganize the OpenShift production deployment documentation:

1. Changed main section from "Production Deployments" to "Options for
   end-to-end TLS" for better clarity

2. Renamed subsections for consistency and clarity:
   - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)"
   - "End-to-end TLS using pass-through Route (all OpenShift versions)"

3. Clarified that the Gateway hostname is typically a wildcard:
   "typically a wildcard like *.openshell-ingress-gw.example.com"

4. Removed the recommendation to copy the cluster's wildcard certificate
   from openshift-ingress namespace, as this is not a recommended
   security best practice

These changes make it clearer that users have two end-to-end TLS options
and help them understand the typical naming pattern for Gateway hostnames.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager

When using BackendTLSPolicy with cert-manager, the certgen hook now polls
for up to 90 seconds waiting for cert-manager to issue the TLS certificate
before creating the backend CA ConfigMap. This eliminates the need for a
second `helm upgrade` in most cases.

The hook polls every 2 seconds with progress logging every 10 seconds.
If cert-manager takes longer than 90 seconds, the hook times out gracefully
and logs a warning, preserving the fallback to manual ConfigMap creation
or a second upgrade.

The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s
margin for ConfigMap creation and hook completion.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable timeout for certgen hook

Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long
the certgen hook Job can run. When using cert-manager with BackendTLSPolicy,
the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap
creation and cleanup.

This allows users to increase the timeout for environments where cert-manager
takes longer than 90 seconds to issue certificates, without requiring code
changes.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180  # Hook polls for 150 seconds
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): document configurable certgen timeout

Update documentation to mention the pkiInitJob.timeoutSeconds value and
how it affects the cert-manager polling behavior when using BackendTLSPolicy.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add configurable failure behavior for certgen timeout

Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether
the certgen hook fails or succeeds when cert-manager does not issue a
certificate within the polling timeout.

When false (default), the hook succeeds with a warning and users can run
`helm upgrade` after cert-manager issues the certificate to create the
backend CA ConfigMap. This provides backwards-compatible behavior.

When true, the hook fails immediately if the timeout is reached, providing
clear feedback that BackendTLSPolicy is non-functional. This is useful for
strict validation requirements where incomplete installs should fail fast.

Example usage:
```yaml
pkiInitJob:
  timeoutSeconds: 180
  failOnTimeout: true  # Fail install if cert-manager takes >150s
```

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): change failOnTimeout default to true and add troubleshooting docs

Change `pkiInitJob.failOnTimeout` default from false to true to provide
immediate feedback when cert-manager does not issue certificates within
the polling timeout. This prevents silent failures where BackendTLSPolicy
is non-functional but the install appears to succeed.

Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx
documenting the specific error "TLS error: Secret is not supplied by SDS"
that occurs when the backend CA ConfigMap is missing, with step-by-step
resolution instructions.

Updated comments in values.yaml to clearly document the default behavior
and explain when administrators might see connectivity errors if they
override the default to failOnTimeout=false.

BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm
installs will fail if cert-manager takes longer than (timeoutSeconds - 30)
seconds to issue certificates. To restore the old behavior of allowing
installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make cert-manager resources pre-install hooks to fix ordering

Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks
with weight -30, before the certgen hook (weight -20). This fixes the
chicken-and-egg problem where the certgen hook was waiting for Secrets
created by Certificates that hadn't been created yet.

**Hook ordering:**
1. Certificate and Issuer resources created (weight -30)
2. cert-manager issues certificates and creates Secrets
3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap
4. Main resources (StatefulSet, Service, etc.) created

Previously, the certgen pre-install hook would run before any resources
were created, poll for a non-existent Secret, timeout, and fail. The
Certificate resources would never get created because Helm waits for
all pre-install hooks to succeed before creating main resources.

This fix allows single-stage installs to work reliably as long as
cert-manager can issue certificates within the polling timeout.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* feat(helm): add validation to prevent enableMtls with BackendTLSPolicy

Add Helm chart validation that fails the install if both
server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true
are set, since this is an invalid configuration.

BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy
cannot present client certificates to the backend. This validation provides
immediate, clear feedback at install time rather than allowing the
misconfiguration to be discovered through runtime errors.

Example error message:
```
Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because
the Gateway proxy cannot present client certificates to the backend;
set server.tls.enableMtls=false
```

Also updated documentation to mention this validation check.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior

Improve documentation to clearly explain that pkiInitJob.timeoutSeconds
controls the Job deadline, but the actual polling timeout is
(timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation
and cleanup.

Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling"
to make the relationship explicit and avoid confusion where users might
expect the hook to poll for the full timeout value.

Updated both values.yaml inline comments and ingress.mdx documentation
for consistency.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* docs(openshift): remove outdated two-stage install instructions

Update OpenShift documentation to reflect that single-stage installs now
work with cert-manager and BackendTLSPolicy. The Certificate resources
run as pre-install hooks (weight -30) before certgen (weight -20),
allowing the hook to poll for and find the issued certificates.

Removed the outdated two-step process:
1. helm install (cert-manager issues cert, hook logs warning)
2. helm upgrade (hook creates ConfigMap)

Replaced with current single-stage behavior:
- Certificate resources created as pre-install hooks
- certgen hook polls for up to 90 seconds (configurable)
- Single helm install succeeds in most cases
- Fails fast by default if timeout reached

This brings openshift.mdx in line with the already-updated ingress.mdx
documentation.

Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>

* fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration

The timeout value now represents the actual polling time that users
experience when waiting for cert-manager to issue certificates.
The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to
allow buffer time for ConfigMap creation and cleanup.

Previously, the hook polled for (timeoutSeconds - 30) seconds, which
was confusing when users set timeoutSeconds=180 and only got 150
seconds of actual polling.

Updated documentation in values.yaml, ingress.mdx, and openshift.mdx
to reflect the clearer behavior.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(helm): update values.yaml and README with correct polling duration

Updated the caCertificateConfigMapName description to reflect that the
hook polls for exactly pkiInitJob.timeoutSeconds seconds, not
(timeoutSeconds - 30) seconds.

Regenerated README.md with helm-docs.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): add OIDC configuration to helm install and CLI examples

Updated all helm install and openshell gateway add examples in ingress.mdx
and openshift.mdx to include OIDC issuer and audience configuration.

Examples now use concrete placeholder values:
- OIDC issuer: https://keycloak.example.com/realms/openshell
- OIDC audience: openshell-cli
- Hostname: gateway.example.com
- ClusterIssuer: letsencrypt-prod

This makes it clearer how to configure OIDC authentication, which is
required when using BackendTLSPolicy or HTTPS termination since the
Gateway proxy cannot present client certificates to the backend.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* docs(kubernetes): explicitly list OIDC client ID in gateway add examples

Added --oidc-client-id openshell-cli to all openshell gateway add
commands in ingress.mdx and openshift.mdx, making the default client
ID explicit in the examples even though it's the CLI default.

This improves clarity and helps users understand the complete OIDC
configuration needed for gateway registration.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>

* fix(helm): address PR review feedback for BackendTLSPolicy

- Read backend CA from the authoritative server Secret instead of the
  in-memory PKI bundle so enabling BackendTLSPolicy on an existing
  release uses the CA that actually signed the server certificate.
- Reconcile the backend CA ConfigMap on every hook run (compare and
  update) instead of skipping when it already exists, so CA rotations
  propagate automatically.
- Remove hook annotations from cert-manager Issuer/Certificate resources
  so they remain regular release objects managed by Helm lifecycle. Split
  the cert-manager backend CA ConfigMap creation into a separate
  post-install/post-upgrade hook Job that polls after cert-manager
  Certificate resources are applied.
- Update architecture/gateway.md, docs/reference/gateway-config.mdx,
  debug-openshell-cluster skill, and helm-dev-environment skill with
  BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout
  documentation.

Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): resolve markdown lint errors in helm README and kubernetes docs

Escape inline HTML angle brackets in README.md template placeholders,
remove trailing spaces, and add blank lines around fenced code blocks
in numbered lists.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Update docs

* fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile

Wrap `<fullname>` and `<namespace>` template placeholders in backticks
so markdownlint does not flag them as inline HTML (MD033). Regenerate
mise.lock to match current mise.toml after rebase onto main.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* Run 'mise lock'

* fix: align mise.lock with CI mise version output

The lockfile was regenerated locally with mise 2026.8.10 which resolves
uv Linux artifacts to gnu variants and adds provenance_verified fields,
but CI uses v2026.4.25 which produces musl variants without those fields.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

* fix(docs): correct cert-manager hook ordering and clientCaSecretName comment

Update ingress.mdx and openshift.mdx to describe Certificate resources
as regular release objects with a post-install/post-upgrade Job, matching
the current implementation and architecture/gateway.md.

Fix values.yaml clientCaSecretName comment to state that "" disables
client certificate verification, matching the helper and access-control
docs.

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>

---------

Signed-off-by: Brandon Squizzato <bsquizza@redhat.com>
Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com>
Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com>
2026-09-14 17:39:15 +00:00
alangou 9b4b63ec69 fix(deps): update DOMPurify and runtime image packages (#3276)
Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-09-11 10:45:29 +00:00
Jesse JaggarsandDrew Newberry 02b664bb0d refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): introduce canonical gateway fields

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* refactor(config): enforce gateway schema version 2

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve compute driver runtime guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address schema v2 review regressions

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): complete schema v2 migration safeguards

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): expand schema v2 regression coverage

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(config): add schema v2 parity manifest

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): correct parity manifest inventory

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): record schema v2 intentional changes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): disposition schema v2 parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add dual schema parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): establish compute lifecycle parity baseline

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve gateway option compatibility

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record gateway option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* docs(config): close gateway-wide parity gaps

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(podman): apply configured pids limit

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): validate Podman option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add Kubernetes option parity harness

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record Kubernetes option parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition VM parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): add external driver parity lane

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): preserve external driver pull policy

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity artifacts and launches

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): require clean parity build sources

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): use isolated supervisor tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): qualify parity image tags

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): serve parity supervisor locally

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): isolate parity podman services

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): harden parity evidence provenance

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): pin parity sandbox artifacts

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): attest parity runtime inputs

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): bind parity runtime evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): record compute boundary parity

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(e2e): disposition cross-cutting parity lanes

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight gateway config upgrades

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): preserve rebase integration guarantees

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(ci): isolate temporary git signing config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): update remaining schema v2 consumers

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(ci): provide e2fs tools to VM tests

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): align preflight with gateway startup

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(vm): preserve rootfs tar configuration

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* chore(config): adopt duration unit constructors

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(packaging): preflight RPM gateway config

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(config): address driver review findings

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(e2e): require fresh semantic parity evidence

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(docker): update tests for renamed sandbox label

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* test(gateway): preserve selective driver coverage after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 05:00:24 +00:00
Drew Newberry 38f2aef930 feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(gateway): support selective Windows MXC builds

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:36:09 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
John T. MyersandJohn Myers 226a83b323 fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies

Closes #2904

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(supervisor): preserve same-provider placeholders in request bodies

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-10 23:10:01 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Evie Howard 6e6b3c8905 refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com>
2026-09-09 13:28:17 +00:00
Jorge 519e5eb35f feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift

Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.

The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:

- Drives the gateway through a passthrough OpenShift Route secured with
  mandatory mTLS instead of port-forward, so the connect suites
  (live_policy_update, port_forward, sync, connect-based
  sandbox_lifecycle, settings_management) actually pass. Computes the
  Route host from the cluster ingress domain, extracts client mTLS
  material from the openshell-client-tls secret, waits for the Route to
  serve mTLS, asserts a certless caller is rejected at the TLS
  handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
  runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
  range.
- Grants the privileged SCC to openshell-sandbox before Helm install
  and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
  DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.

The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.

The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.

The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.

TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.

The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout

Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.

Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.

Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.

Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>

---------

Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
2026-09-08 20:37:19 +00:00
Jesse Jaggars 2ad86c1b2e fix(cli): fail closed when OIDC refresh fails (#2817)
* fix(cli): fail closed when OIDC refresh fails

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

* fix(cli): fall back to unexpired OIDC token on refresh failure

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>

---------

Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
2026-09-08 15:37:26 +00:00
John T. MyersandJohn Myers fc0929749c fix(policy): harden advisor transport proposals (#3136)
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-04 19:38:00 +00:00
Adel ZaaloukandJohn Myers 48a8a4bf09 feat(ocsf): configurable schema version for SIEM backward compatibility (#2717)
* feat(ocsf): configurable schema version for SIEM backward compatibility

Add a gateway-configurable OCSF schema version target that downgrades
JSONL output for SIEMs that only support older schema versions. AWS
Security Lake requires v1.1.0, Splunk CIM Add-On targets v1.1-v1.3.

The downgrade filter strips profile-gated fields (ai_model, container,
observation_point_id), removes unknown profiles from metadata.profiles,
rewrites metadata.version, and adds an unmapped.downgraded_from
breadcrumb so auditors can distinguish "no model involved" from "model
attribution stripped."

Supported target versions (1.1, 1.3) are enforced by an allow-list in
the settings registry. Invalid values are rejected with a clear error.

The setting flows to sandboxes via the settings bundle and takes effect
on the next poll cycle. The shorthand log output is unaffected.

Closes #2662

Signed-off-by: Adel Zaalouk <zanetworker@gmail.com>
Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>

* fix(ocsf): align downgrade with schema version

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Adel Zaalouk <zanetworker@gmail.com>
Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-04 18:43:58 +00:00
Johnny Greco 8bc7955263 feat(skills): separate public and contributor workflows (#2899)
* feat(skills): separate public and contributor workflows

Closes #2736

Publish the four user-facing OpenShell skills from the top-level skills directory, mark contributor workflows internal, and update portability guidance, validation, and documentation.

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(skills): clarify public skill audit scope

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(skills): use markdown documentation links

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

* docs(skills): align public and contributor guidance

Signed-off-by: Johnny Greco <jogreco@nvidia.com>

---------

Signed-off-by: Johnny Greco <jogreco@nvidia.com>
2026-09-02 16:02:17 +00:00