mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
428
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8751e35e28 |
fix(supervisor-network): reject malformed OPA policy containers (#3337)
Validate raw OPA object and array containers before normalization and access-preset expansion can skip malformed values. Return one fixed structural error without embedding authored policy data. Preserve versionless and runtime-only OPA data and existing semantic validation. Cover initial string/file loading, middleware callback order, valid deny-rule enforcement, and rejected reload state and generation. Document the loader contract and its engine-local rejection behavior. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
dbe36eaf85 |
fix(security): harden Vault credential transport (#3329)
Reject non-loopback plaintext Vault endpoints, disable redirects, and support private CA bundles without weakening hostname verification. Update Helm configuration, documentation, operator skills, and regression coverage for OSSR-002. Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
dfd5238d0d |
fix(gator): make supervised lifecycle sandbox-native (#3343)
* fix(gator): run supervisor as sandbox main process Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * feat(gator): persist supervised state history Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * refactor(gator): remove obsolete background launch mode Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
481ce566e1 |
fix(ocsf): correct HTTP activity context (#3316)
Previously, metadata events omitted both HTTP request and response objects, early proxy rejections used HTTP Activity without request context, and unsupported-scheme events did not expose enough safe HTTP context to satisfy the OCSF 1.8 schema. Now, metadata events include a method-only request and their actual HTTP response codes without recording the metadata URL. Unsupported-scheme events also include a method-only request plus the generated 400 response. Authority mismatches and credential-resolution denials use HTTP Activity with their generated 403 or 500 responses, and HTTP activity IDs are derived from the request method. Additionally, HttpActivityBuilder now enforces the OCSF request-or-response constraint at compile time. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
39cf4823f7 |
feat(api): add structured gateway errors and SDK decoding (#3313)
* feat(api): expose structured gateway errors across SDKs Refs #3051. Add standard validation, conflict, and retry details; preserve raw transport status in Rust, Go, TypeScript, and Python; document status and recovery guidance. This is the structured-error foundation only. Mutation result shapes, allow_missing, durable request deduplication, and exec retry semantics remain follow-up work. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(python): preserve wrapped RPC cleanup handling Inspect the original gRPC call when handling missing sandboxes during deletion waits and managed cleanup. Add intercepted cleanup regressions and clarify the error-wrapper migration contract. Addresses the cleanup review on #3313; part of #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
fd3fd9cf74 |
feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection results together in sandbox status. Keep sandbox lifecycle readiness separate so an external connection failure does not mark the sandbox unready. Expose direct endpoint records through the CLI and SDKs, with plain-language failure explanations and gateway acceptance times. Keep observation tracking, runtime reporting, and gateway validation in dedicated endpoint status modules. Preserve bounded reporting, request attribution, retry ordering, and configuration and supervisor authority checks. Clear obsolete observations while retaining the configured addresses, and document the distinction between an observed HTTP response, current availability, and tool success. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
26f2f96393 |
feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell
* Update README and gateway config to clarify egress proxy address handling and allocation
* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash
* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading
* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling
* Add conditional compilation for Windows host module
* add unit tests for OPA policy evaluation and identity handling
* remove openshell-supervisor-network from unsupported driver package test exclusion list
* feat(mxc): enable host proxy TLS state generation
Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.
* fix(docs): remove outdated notes on governed egress from docs
* fix(tests): update TLS environment variable paths to use temporary directory
* fix(examples): make run-mxc-e2e harness correct and orphan-free
The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.
Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
(proves the agent ran) while the denied write must be absent. fs-empty
probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
harness never writes to the persistent store and cannot leave orphan
sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
(canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).
Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): probe timeout is milliseconds (10ms->30000ms)
MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): use native paths in real runtime probes
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec
Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:
1. Seed process env from host (driver.rs)
ProcessContainer starts with a completely blank environment -- no PATH,
SystemRoot, or anything. Seed the process env from the gateway host
environment so the agent binary can locate DLLs and run. Skip internal
Windows drive-letter variables (keys starting with '=') which cause
CreateProcessW to return ERROR_ENVVAR_NOT_FOUND. User agent_env entries
and TLS CA vars are applied as overrides on top of the host env.
2. Remove TLS readonly_paths grant (driver.rs)
The release wxc-exec (BaseContainer dispatcher) requires write-DAC
permission on every path in readonly_paths to set up AppContainer ACLs.
Adding the proxy's temp TLS directory caused a DACL error and exit -1.
The CA cert paths remain available to the agent via TLS env vars.
3. Remove allowedHosts from network JSON (mxc.rs)
The release wxc-exec rejects network.allowedHosts / network.blockedHosts
on Windows with "not yet supported". Removed the loopback exemption
attempt (127.0.0.1, ::1, localhost) from the network section.
Intra-container loopback works natively in the release binary without
it -- the spawner can connect to the server at 127.0.0.1:22000 directly.
Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
relay-ready.txt polling before ws-echo to avoid connecting before the
spawner has established the proxy bridge.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)
Four robustness/correctness fixes from CodeRabbit:
1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
throw. If the process is alive but never binds the port, $gw is not yet
assigned in the caller, so the finally block cannot reap it -> orphan
gateway holding the port for the next run.
2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
of a policy rejection (gateway-registration/transport/fixture errors also
exit non-zero and would false-pass). PASS now requires a genuine
rejection signal (network / invalid_argument / network_policies) AND that
it is not an infrastructure failure; other non-zero exits go to FAIL with
output captured.
3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
Wait-File lands the control artifact, so a late denied write (enforcement
regression racing the control write) can no longer be recorded as PASS.
4. -KeepRunning: break out of the scenario loop after the first scenario so
a later scenario does not start a second gateway on the same port
(previously a reliable port collision instead of a usable debug mode).
Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".
Signed-off-by: Akber Raza <akberr@nvidia.com>
* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle
Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.
Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.
Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1
- Require -Scenario when -KeepRunning: the loop breaks after the first
scenario, so a full-suite run would execute only one scenario yet still
report the suite as PASS. Fail fast so a partial run can't be mislabeled
complete.
- Start-Transcript now runs inside the guarded try block with a
$transcriptStarted flag; Stop-Transcript is only called when it actually
started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
(reliable .Count and a proper array for the scenario loop on PS 5.1).
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths
Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.
Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.
Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling
* fix(mxc): reconcile proxy support after rebase
Restore the proxy-enabled OCSF audit example removed by
|
||
|
|
cc780d4e17 |
feat(helm): add BackendTLSPolicy support (#2728)
* feat(helm): add optional BackendTLSPolicy for e2e TLS Add grpcRoute.backendTLSPolicy values to optionally create a BackendTLSPolicy resource that enables end-to-end TLS between the Gateway proxy and the OpenShell gateway pod. The Gateway proxy terminates client-facing TLS and re-encrypts when connecting to the backend, validating the pod's certificate against a user-supplied CA ConfigMap. This removes the requirement to set server.disableTls=true when using HTTPS at the Gateway listener. Supported on OpenShift 4.22+ and other platforms with BackendTLSPolicy support in the Gateway API implementation. Update OpenShift and ingress documentation with e2e TLS instructions. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm,server): auto-create backend CA ConfigMap in certgen hook Extend the generate-certs command with --backend-ca-configmap-name and --backend-ca-source-secret flags. When BackendTLSPolicy is enabled, the certgen pre-install hook creates the CA ConfigMap automatically: - pkiInitJob mode (default): uses the CA from the generated PKI bundle. Fully automatic on first install. - cert-manager mode: reads ca.crt from the server TLS Secret. On first install the Secret does not exist yet (cert-manager reconciles after templates are applied), so the ConfigMap is created on the first helm upgrade. Logs a warning on the initial skip. The caCertificateConfigMapName value now defaults to <fullname>-backend-ca when empty, so users only need to set backendTLSPolicy.enabled=true. Update certgen RBAC to include configmaps get/create when the feature is enabled. Add CLI arg parsing tests for the new flags. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * refactor(helm): add server.tls.enableMtls flag for mTLS control Replace automatic mTLS disabling based on BackendTLSPolicy with an explicit server.tls.enableMtls flag that defaults to true. The user is now responsible for setting this to false when using BackendTLSPolicy, as ingress proxies cannot present client certificates to backends. Updated: - values.yaml: Added server.tls.enableMtls (default true) - gateway-config.yaml: Check enableMtls instead of backendTLSPolicy - _gateway-workload.tpl: Check enableMtls for client CA mount - Tests: Updated to use enableMtls flag - Docs: Added enableMtls=false to BackendTLSPolicy examples - README: Document new flag and BackendTLSPolicy requirement Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(helm): clarify cert-manager backend CA ConfigMap workflow Update documentation to explain the two-step install process required when using cert-manager with BackendTLSPolicy: 1. helm install - cert-manager issues the server certificate, but the certgen hook can't create the backend CA ConfigMap yet (cert-manager reconciles after templates are applied) 2. helm upgrade - certgen hook reads the CA from the cert-manager-issued certificate and creates the ConfigMap Previously, the docs said "created on first upgrade" without explaining why or that the feature won't work until then. The updated docs now: - Explain the timing issue (cert-manager reconciles after chart install) - Provide clear steps for the cert-manager workflow - Note that pkiInitJob (default) creates it immediately on install - Clarify that users must wait for the Certificate to be Ready before running the second upgrade Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): remove incorrect external hostname requirement for BackendTLSPolicy BackendTLSPolicy validates the backend certificate against the service FQDN (e.g., openshell.openshell.svc.cluster.local), not the external hostname. The external hostname only needs to be on the Gateway listener certificate for client-facing TLS. The default certManager.serverDnsNames already includes all required service FQDN variants, so no configuration is needed for BackendTLSPolicy to work. Fixed incorrect documentation that claimed: - "The server certificate SAN list must include the external hostname" - Users need to "configure certManager.serverDnsNames with the external hostname" Removed the unnecessary pkiInitJob.serverDnsNames override from the example and clarified that: - Gateway listener certificate needs the external hostname (for clients) - Backend certificate needs the service FQDN (for Gateway proxy) - The service FQDN is already in the defaults Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs: clarify ACME with LetsEncrypt reference Change all references from "ACME issuer" to "LetsEncrypt/ACME issuer" to help users understand that LetsEncrypt is the most common ACME provider and what ACME means in practice. Updated: - docs/kubernetes/managing-certificates.mdx - docs/kubernetes/openshift.mdx - deploy/helm/openshell/values.yaml - deploy/helm/openshell/README.md - deploy/helm/openshell/ci/values-openshift-route-cert-manager.yaml Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * docs(openshift): restructure end-to-end TLS options and clarify Gateway hostname Reorganize the OpenShift production deployment documentation: 1. Changed main section from "Production Deployments" to "Options for end-to-end TLS" for better clarity 2. Renamed subsections for consistency and clarity: - "End-to-end TLS using Gateway API and BackendTLSPolicy (OpenShift 4.22+)" - "End-to-end TLS using pass-through Route (all OpenShift versions)" 3. Clarified that the Gateway hostname is typically a wildcard: "typically a wildcard like *.openshell-ingress-gw.example.com" 4. Removed the recommendation to copy the cluster's wildcard certificate from openshift-ingress namespace, as this is not a recommended security best practice These changes make it clearer that users have two end-to-end TLS options and help them understand the typical naming pattern for Gateway hostnames. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * feat(helm): eliminate two-stage install for BackendTLSPolicy with cert-manager When using BackendTLSPolicy with cert-manager, the certgen hook now polls for up to 90 seconds waiting for cert-manager to issue the TLS certificate before creating the backend CA ConfigMap. This eliminates the need for a second `helm upgrade` in most cases. The hook polls every 2 seconds with progress logging every 10 seconds. If cert-manager takes longer than 90 seconds, the hook times out gracefully and logs a warning, preserving the fallback to manual ConfigMap creation or a second upgrade. The Job's activeDeadlineSeconds is 120s, so the 90s timeout leaves 30s margin for ConfigMap creation and hook completion. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable timeout for certgen hook Add `pkiInitJob.timeoutSeconds` Helm value (default 120) to control how long the certgen hook Job can run. When using cert-manager with BackendTLSPolicy, the hook polls for (timeoutSeconds - 30) seconds to leave margin for ConfigMap creation and cleanup. This allows users to increase the timeout for environments where cert-manager takes longer than 90 seconds to issue certificates, without requiring code changes. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 # Hook polls for 150 seconds ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): document configurable certgen timeout Update documentation to mention the pkiInitJob.timeoutSeconds value and how it affects the cert-manager polling behavior when using BackendTLSPolicy. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add configurable failure behavior for certgen timeout Add `pkiInitJob.failOnTimeout` Helm value (default false) to control whether the certgen hook fails or succeeds when cert-manager does not issue a certificate within the polling timeout. When false (default), the hook succeeds with a warning and users can run `helm upgrade` after cert-manager issues the certificate to create the backend CA ConfigMap. This provides backwards-compatible behavior. When true, the hook fails immediately if the timeout is reached, providing clear feedback that BackendTLSPolicy is non-functional. This is useful for strict validation requirements where incomplete installs should fail fast. Example usage: ```yaml pkiInitJob: timeoutSeconds: 180 failOnTimeout: true # Fail install if cert-manager takes >150s ``` Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): change failOnTimeout default to true and add troubleshooting docs Change `pkiInitJob.failOnTimeout` default from false to true to provide immediate feedback when cert-manager does not issue certificates within the polling timeout. This prevents silent failures where BackendTLSPolicy is non-functional but the install appears to succeed. Add comprehensive troubleshooting section to docs/kubernetes/ingress.mdx documenting the specific error "TLS error: Secret is not supplied by SDS" that occurs when the backend CA ConfigMap is missing, with step-by-step resolution instructions. Updated comments in values.yaml to clearly document the default behavior and explain when administrators might see connectivity errors if they override the default to failOnTimeout=false. BREAKING CHANGE: pkiInitJob.failOnTimeout now defaults to true. Helm installs will fail if cert-manager takes longer than (timeoutSeconds - 30) seconds to issue certificates. To restore the old behavior of allowing installs to succeed with a warning, set `pkiInitJob.failOnTimeout=false`. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make cert-manager resources pre-install hooks to fix ordering Make Certificate and Issuer resources run as pre-install/pre-upgrade hooks with weight -30, before the certgen hook (weight -20). This fixes the chicken-and-egg problem where the certgen hook was waiting for Secrets created by Certificates that hadn't been created yet. **Hook ordering:** 1. Certificate and Issuer resources created (weight -30) 2. cert-manager issues certificates and creates Secrets 3. certgen hook runs (weight -20), finds Secrets, creates ConfigMap 4. Main resources (StatefulSet, Service, etc.) created Previously, the certgen pre-install hook would run before any resources were created, poll for a non-existent Secret, timeout, and fail. The Certificate resources would never get created because Helm waits for all pre-install hooks to succeed before creating main resources. This fix allows single-stage installs to work reliably as long as cert-manager can issue certificates within the polling timeout. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * feat(helm): add validation to prevent enableMtls with BackendTLSPolicy Add Helm chart validation that fails the install if both server.tls.enableMtls=true and grpcRoute.backendTLSPolicy.enabled=true are set, since this is an invalid configuration. BackendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend. This validation provides immediate, clear feedback at install time rather than allowing the misconfiguration to be discovered through runtime errors. Example error message: ``` Error: grpcRoute.backendTLSPolicy requires mTLS to be disabled because the Gateway proxy cannot present client certificates to the backend; set server.tls.enableMtls=false ``` Also updated documentation to mention this validation check. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(helm): clarify pkiInitJob.timeoutSeconds polling behavior Improve documentation to clearly explain that pkiInitJob.timeoutSeconds controls the Job deadline, but the actual polling timeout is (timeoutSeconds - 30) to reserve 30 seconds for ConfigMap creation and cleanup. Added concrete example: "timeoutSeconds=180 allows 150 seconds of polling" to make the relationship explicit and avoid confusion where users might expect the hook to poll for the full timeout value. Updated both values.yaml inline comments and ingress.mdx documentation for consistency. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * docs(openshift): remove outdated two-stage install instructions Update OpenShift documentation to reflect that single-stage installs now work with cert-manager and BackendTLSPolicy. The Certificate resources run as pre-install hooks (weight -30) before certgen (weight -20), allowing the hook to poll for and find the issued certificates. Removed the outdated two-step process: 1. helm install (cert-manager issues cert, hook logs warning) 2. helm upgrade (hook creates ConfigMap) Replaced with current single-stage behavior: - Certificate resources created as pre-install hooks - certgen hook polls for up to 90 seconds (configurable) - Single helm install succeeds in most cases - Fails fast by default if timeout reached This brings openshift.mdx in line with the already-updated ingress.mdx documentation. Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> * fix(helm): make pkiInitJob.timeoutSeconds the actual polling duration The timeout value now represents the actual polling time that users experience when waiting for cert-manager to issue certificates. The Job activeDeadlineSeconds is set to (timeoutSeconds + 30) to allow buffer time for ConfigMap creation and cleanup. Previously, the hook polled for (timeoutSeconds - 30) seconds, which was confusing when users set timeoutSeconds=180 and only got 150 seconds of actual polling. Updated documentation in values.yaml, ingress.mdx, and openshift.mdx to reflect the clearer behavior. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(helm): update values.yaml and README with correct polling duration Updated the caCertificateConfigMapName description to reflect that the hook polls for exactly pkiInitJob.timeoutSeconds seconds, not (timeoutSeconds - 30) seconds. Regenerated README.md with helm-docs. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): add OIDC configuration to helm install and CLI examples Updated all helm install and openshell gateway add examples in ingress.mdx and openshift.mdx to include OIDC issuer and audience configuration. Examples now use concrete placeholder values: - OIDC issuer: https://keycloak.example.com/realms/openshell - OIDC audience: openshell-cli - Hostname: gateway.example.com - ClusterIssuer: letsencrypt-prod This makes it clearer how to configure OIDC authentication, which is required when using BackendTLSPolicy or HTTPS termination since the Gateway proxy cannot present client certificates to the backend. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * docs(kubernetes): explicitly list OIDC client ID in gateway add examples Added --oidc-client-id openshell-cli to all openshell gateway add commands in ingress.mdx and openshift.mdx, making the default client ID explicit in the examples even though it's the CLI default. This improves clarity and helps users understand the complete OIDC configuration needed for gateway registration. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> * fix(helm): address PR review feedback for BackendTLSPolicy - Read backend CA from the authoritative server Secret instead of the in-memory PKI bundle so enabling BackendTLSPolicy on an existing release uses the CA that actually signed the server certificate. - Reconcile the backend CA ConfigMap on every hook run (compare and update) instead of skipping when it already exists, so CA rotations propagate automatically. - Remove hook annotations from cert-manager Issuer/Certificate resources so they remain regular release objects managed by Helm lifecycle. Split the cert-manager backend CA ConfigMap creation into a separate post-install/post-upgrade hook Job that polls after cert-manager Certificate resources are applied. - Update architecture/gateway.md, docs/reference/gateway-config.mdx, debug-openshell-cluster skill, and helm-dev-environment skill with BackendTLSPolicy, backend CA ConfigMap, enableMtls, and timeout documentation. Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): resolve markdown lint errors in helm README and kubernetes docs Escape inline HTML angle brackets in README.md template placeholders, remove trailing spaces, and add blank lines around fenced code blocks in numbered lists. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Update docs * fix(helm): escape inline HTML in values.yaml descriptions and sync mise lockfile Wrap `<fullname>` and `<namespace>` template placeholders in backticks so markdownlint does not flag them as inline HTML (MD033). Regenerate mise.lock to match current mise.toml after rebase onto main. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * Run 'mise lock' * fix: align mise.lock with CI mise version output The lockfile was regenerated locally with mise 2026.8.10 which resolves uv Linux artifacts to gnu variants and adds provenance_verified fields, but CI uses v2026.4.25 which produces musl variants without those fields. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> * fix(docs): correct cert-manager hook ordering and clientCaSecretName comment Update ingress.mdx and openshift.mdx to describe Certificate resources as regular release objects with a post-install/post-upgrade Job, matching the current implementation and architecture/gateway.md. Fix values.yaml clientCaSecretName comment to state that "" disables client certificate verification, matching the helper and access-control docs. Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> --------- Signed-off-by: Brandon Squizzato <bsquizza@redhat.com> Signed-off-by: Brandon Squizzato <bsquizzato@nvidia.com> Signed-off-by: Brandon Squizzato <bsquizza@nvidia.com> |
||
|
|
0803c4aa4c |
refactor(policy)!: remove NetworkBinary harness field (#3222)
* refactor(policy)!: remove NetworkBinary harness field Closes #3054 Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(policy): preserve unknown fields during migration Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(policy): preserve advisor provenance during merge Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(policy): remove redundant network binary default Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): record harness schema migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ae57979b03 |
feat(mxc): add Windows ETW-to-OCSF audit trail (#3015)
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1) Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019, process 1007, finding 2004). cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux: - openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge thread-local AND stamps sandbox_id+message in one dispatch) plus public set/clear_current_event; OS-aware device (Device::windows/for_current_os) so device.os.name reflects the host instead of a hardcoded Linux stub. - etw_consumer: emit via the routed emit (previously fired a bare info! that never populated the bridge, so the structured event was dropped). - openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR). - device.hostname now resolves to the real gateway machine name. Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid OCSF JSON, per-sandbox attribution intact, disabled state writes nothing. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc): map remaining Sandboxing ETW events to OCSF Close the last three ETW->OCSF gaps so the audit trail covers the full set of events the Sandboxing provider emits (12/12): - ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start; carries the real processId/threadId, the twin of CreateProcessInSandbox which only has the request + command line). - SandboxProxyConfigured -> Device Config State Change [5019] (the one network-plane setup event; surfaces proxyPort, "no proxy" when 0). - SandboxConsoleReferencePlumbed -> Device Config State Change [5019] (console-handle plumbing). map_config_state now handles the full config/hardening/setup family and carries proxyPort/hasConsoleReference/creationFlags as unmapped fields. Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy (SandboxProxyConfigured requires proxy config to fire). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): seed ETW attribution under registry lock + Device tests Address CodeRabbit review on !31: - Prevent stale ETW attribution on a delete/launch race: register the wxc-exec pid while holding the registry lock, and bail if the sandbox entry is already gone. Previously the attribution key could be seeded after `delete` had removed the sandbox, leaving a stale key that could misroute later Sandboxing ETW events to a dead sandbox_id. Lock order (registry -> attribution) matches the delete path, so no deadlock. - Add unit tests for the new Device::windows and Device::for_current_os constructors to harden Windows/Linux OCSF device parity. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): buffer+replay racing events and harden attribution keys Addresses two ETW->OCSF attribution review items (Shailendra #1, #2). #2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved(). #1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute. Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0). Signed-off-by: Akber Raza <akberr@nvidia.com> * docs(mxc-etw): note cmd_line is captured raw with no privacy filtering Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): open ETW trace on caller thread so start_session reports real status Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage). Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace. Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit). Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): guard pending-event replay against PID recycling CodeRabbit flagged that drain_resolved() re-resolved buffered events against the live by_pid map, so if Windows recycled a wxc-exec PID within PENDING_TTL a stale event from the dead sandbox could be emitted under the new owner. Stamp each by_pid registration with its Instant and add resolve_replay(), used only on the buffered/replay path. It (a) never falls back to the recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID match only when the registration is not newer than the buffered event by more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well outside that window, so the stale event ages out instead of misattributing. The legitimate #2 seed race (registration lands ~immediately) still replays. Adds unit tests for the recycle-refusal, in-grace acceptance, and weak-fallback exclusion. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc-etw): surface unexpected ProcessTrace termination (review #4) start_session already returns Err on OpenTraceW failure (runs on the caller thread since e41a7701), closing the first half of Shailendra's #4. This closes the second half: ProcessTrace's result was discarded, so if capture died mid-run the backend had no way to know. Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR and, when the pump returns without a deliberate stop, logs at ERROR that MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping` before teardown so a normal shutdown isn't misreported, and EtwSession::is_capture_alive() exposes the state for status/diagnostics. Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows, JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure). Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1, mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit trail across all four classes (6002/5019/1007/2004). Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's presence already indicates proxy configuration. Verified on-box: 26 events, all mapped ETW event types present. Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1 Improve the ETW to OCSF audit-trail example output and make it safe to ship. Report: - Add an event-type coverage count ("N of M expected event types fired"); the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy). - Split the checklist into expected event types vs anomaly findings (ActivityError/FallbackError), which are reported separately and not counted toward coverage (a clean run may emit none). - Verdict is now coverage-based (all expected types must fire) instead of the looser "at least 3 OCSF classes". - Call out the absolute path to the durable OCSF JSONL log prominently. Client-safety: - Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to opt in. Removes a hardcoded internal share path from a published example. - Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:") in favor of neutral "Results bundle:". - Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior. Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8 event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer fallback) -> reduced set as expected, clean output. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): configure OCSF audit workloads per sandbox - remove unsupported gateway-scoped workload fields from the shipped MXC audit example. - build the command, working directory, and filesystem grant from each run's ShareDir - pass the workload through --driver-config-json. - preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge - require the workload output when determining the audit verdict. Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(mxc): omit command arguments from OCSF audit logs - record only the executable basename for MXC CreateProcessInSandbox audit events - leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs - cover tokens, passwords, signed URLs, and PII with a secret-leak regression test - update the audit example, architecture guidance, and published logging documentation - preserve ETW attribution and future host_connect_proxy enforcement behavior Signed-off-by: Akber Raza <akberr@nvidia.com> * feat(etw): enhance ETW session management with distinct naming for concurrent gateways * fix(etw): bound the audit queue during overload - replace the unbounded ETW callback channel with count- and byte-bounded buffering - keep the ETW callback non-blocking and count records rejected during overload - emit immediate, rate-limited warnings that identify resulting audit coverage gaps - make the audit example fail when queue overload causes dropped ETW records - cover stalled consumers, oversized events, and warning throttling with unit tests Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): harden sandbox audit attribution - remove command-line and persistent per-PID fallback keys from live and replay resolution - retire the driver-owned wxc-exec PID before publishing child completion - retain established identity, activity, and correlation-vector links only for the five-second late-event window - prevent buffered records from crossing rapid PID retirement and reuse boundaries - add resolver and lifecycle coverage and document the attribution trust boundary Signed-off-by: Akber Raza <akberr@nvidia.com> * chore(mxc): address rebase follow-ups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): align OCSF audit example with driver config - remove unsupported egress proxy settings - stop requiring the unavailable proxy audit event - update example documentation for supported event coverage Signed-off-by: Akber Raza <akberr@nvidia.com> * fix(etw): redact command-line secrets in DecodedEtwEvent summary * fix(etw): enhance PID resolution and event attribution logic for ETW records * fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration * address rebase issues * fix(mxc): address ETW audit review feedback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): fail closed across ambiguous PID reuse Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): bind ETW attribution to process generation Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Akber Raza <akberr@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Jamie King <jamiek@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
38f2aef930 |
feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): support selective Windows MXC builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1860010850 |
feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sdk): harden pager edge cases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sdk): cover initial resume token Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): fix all-workspaces pager examples Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
226a83b323 |
fix(supervisor): classify credential placeholders in request bodies (#3246)
* fix(supervisor): classify credential placeholders in request bodies Closes #2904 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(supervisor): preserve same-provider placeholders in request bodies Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
90dbe5454b |
feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(cli): preserve template workspace metadata Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): migrate workspace request selectors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(api): update public schema inventory Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
67374efdf8 |
fix(ocsf): emit schema-valid event identities (#3247)
Previously, every event from a sandbox reused the sandbox ID as its event ID. Consumers deduplicating security records could mistake separate events for the same record, and missing device types or empty image objects could prevent schema validation. Give each event its own ID, retain the sandbox association separately, and classify the environment as Other/Sandbox while keeping the OS separate. Omit unknown container details instead of emitting empty objects. Security tooling can now distinguish events from the same sandbox and read their identity consistently after serialization. Refs #1055 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
ea8eda6d5b |
feat(supervisor): enforce MCP request protocol versions (#3241)
* feat(supervisor): enforce MCP request protocol versions Signed-off-by: Shiju <shiju@nvidia.com> * fix(supervisor): enforce MCP versions across HTTP forwarding Apply shared request-version guards before authorization and after forward-request rewriting. Require version metadata to survive HTTP header cleanup, and cover valid initialization and selected-revision forwarding through middleware. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
7f4bd49a47 |
fix(policy): harden landlock.compatibility validation (#2541)
* fix(policy): reject invalid landlock.compatibility values at parse time Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(policy): abort sandbox startup when hard_requirement has no filesystem paths Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(policy): validate landlock.compatibility at gateway and fix zero-path logging Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(policy): reject invalid landlock.compatibility on serialization Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix: fixed linting error Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> |
||
|
|
8af79a7f4b |
fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup On macOS, Homebrew-installed Podman does not create the default socket path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET override and the podman machine inspect lookup in both the compute drivers reference and the debug-openshell-cluster skill. Fixes #1690 Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(podman): resolve macOS Podman socket dynamically Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: restore debug skill file Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: drop legacy debug skill path Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(podman): trim unrelated e2e changes Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(e2e): harden shell array expansion Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: remove unrelated skill note * ci: retrigger checks Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> --------- Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> |
||
|
|
6e6b3c8905 |
refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com> |
||
|
|
118b250f01 |
feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from flag for VM-backed gateways. The CLI detects the archive extension, validates that the gateway uses the VM compute driver, and passes the tar path through driver_config. The VM driver copies the tar into its staging area and feeds it into the existing rootfs extraction and ext4 disk creation pipeline, skipping the container image pull/export steps. Closes #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): validate rootfs tar path at the VM driver boundary The rootfs_tar_path field in driver_config was passed from the API caller directly to tokio::fs::copy without validation. An authenticated user bypassing the CLI could supply arbitrary host paths (e.g. /dev/zero for disk exhaustion, or readable host files for data exfiltration). Introduce a trusted staging directory that the VM driver creates on startup and advertises via GetCapabilities. The CLI now copies the tar into the staging directory before creating the sandbox, and the driver validates that the received path is a regular file inside the staging root and within a configurable size limit (default 10 GiB) before any I/O. New VmDriverConfig options: - rootfs_tar_staging_dir: override the staging directory (default: <state_dir>/rootfs-tar-staging) - rootfs_tar_max_bytes: override the size limit (default: 10 GiB) Addresses GATOR-28b5152e-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar Tighten the rootfs tar staging flow to address the remaining GATOR-01 obligations: - Request-scoped staging: the CLI creates a unique per-request subdirectory (req-<pid>) under the staging root instead of placing files directly in the shared directory. The driver enforces that the tar path is at depth 2 (staging_root/<subdir>/<file>), preventing cross-request path selection. - Size pre-check: the driver advertises rootfs_tar_max_bytes via GetCapabilities. The CLI reads this limit and rejects oversized files before copying, avoiding disk exhaustion in the staging directory. - Cleanup: the driver removes the request staging subdirectory after consuming the tar (on cache hit, copy success, or copy failure), ensuring staged data does not persist beyond the request. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): restore rootfs-tar sandboxes from persisted image identity On restore or restart, the one-shot staged tar archive has already been cleaned up. Reading the persisted image identity from the sandbox state directory and resolving the cached disk path directly avoids re-accessing the deleted staging path. Addresses GATOR-168b9210-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy Replace PID-based request staging directories with tempfile-generated random names to prevent collisions and make paths unpredictable. Replace bare tokio::fs::copy with a streaming copy loop that enforces the advertised max_bytes limit during transfer, closing the TOCTOU gap between the pre-copy size check and the actual copy. Signed-off-by: Philippe Martin <phmartin@nvidia.com> Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): issue rootfs tar staging slots from the gateway A caller could name any host path in `driver_config.vm.rootfs_tar_path`, which the privileged VM driver then read. The CLI-side locality check did not apply to direct API requests. The gateway now owns staging. `BeginRootfsTarStaging` allocates a request-scoped directory under the driver-advertised staging root and returns an opaque single-use token; `CreateSandbox` carries the token, and the gateway substitutes the path it allocated before dispatching to the driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected outright in request validation, so a caller-supplied path never reaches privileged I/O. Tokens are bound to the issuing workspace and subject, consumed once, and expire after 30 minutes. Outstanding slots are capped per caller and overall, so one caller can neither exhaust the staging filesystem nor starve others. An RAII guard reclaims the directory on every failure path after consumption, and an age-gated sweep runs at startup and on each reconcile pass for directories whose driver died before its own cleanup. The token is stripped from the public sandbox before persistence: the stored copy is returned verbatim by GetSandbox, ListSandboxes and WatchSandbox to every member of the workspace. Also fixes two defects this exposed: - The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but the gateway forwards only `driver_config.<driver_name>`, silently dropping unmatched keys. The archive never reached the VM driver, so the documented `--from ./rootfs.tar` flow did not work at all. Config is now nested under `vm` and deep-merged, so a caller's existing VM settings survive instead of being clobbered by a shallow extend. - Staging previously required `GetGatewayInfo`, which is restricted to `platform_admin`, making the feature unusable for ordinary users on any RBAC-enabled gateway. The new RPC matches CreateSandbox at `sandbox:write` / `workspace_role: user`. `compute_driver.proto` is unchanged; the gateway reads the staging root from the capabilities it already stores. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): derive rootfs tar cache identity from archive contents The prepared-disk cache key combined the archive's full path with an mtime truncated to seconds, then mapped punctuation to `-`. Distinct paths such as `/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each other's disk, a rewrite within the same second kept stale contents, and a long path could exceed filesystem component limits. Identity is now a SHA-256 of the archive contents. This is also what makes the cache work at all now that the gateway allocates a fresh staging directory per request: a path-derived key would miss on every create. The archive is hashed, the cache checked, and only on a miss copied — so a hit skips writing a multi-gigabyte file. The copy is hashed as it is written and rejected if the digest differs from the first pass, which closes the window where the source changes during staging rather than approximating it with a re-stat. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): decompress gzip rootfs tar archives during staging `--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes it was given and the guest image-prep VM extracts the staged file with a plain `tar -xpf`. Compressed sources therefore depended on the guest tar auto-detecting gzip, and the prepared disk was sized from the compressed length, which is far too small for the expanded rootfs. Staging now detects gzip from the archive's magic bytes -- the driver only ever sees a gateway-issued staging path, never the caller's file name -- and writes an uncompressed tar. The digest still covers the source bytes, so the "archive changed while staging" check is unaffected, and expansion is bounded by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk. `extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction path matches. Adds unit coverage for gzip staging, bounded expansion, and gzip extraction, plus an e2e sandbox created from a gzip-compressed export. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: Philippe Martin <phmartin@nvidia.com> |
||
|
|
2ad86c1b2e |
fix(cli): fail closed when OIDC refresh fails (#2817)
* fix(cli): fail closed when OIDC refresh fails Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): fall back to unexpired OIDC token on refresh failure Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> |
||
|
|
1510e2c5a7 |
feat(compute): advertise resource capabilities (#3010)
* refactor(driver-docker): group runtime GPU capabilities Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(compute): advertise resource capabilities Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(compute): describe resource capabilities Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(server): satisfy resource capability clippy lint Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
592df3e014 |
feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata Signed-off-by: Shiju <shiju@nvidia.com> * feat(policy): canonicalize MCP version allowlists Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): align MCP policy tests with current main Signed-off-by: Shiju <shiju@nvidia.com> * fix(policy): canonicalize supervisor protobuf ingress Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
fc0929749c |
fix(policy): harden advisor transport proposals (#3136)
Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
48a8a4bf09 |
feat(ocsf): configurable schema version for SIEM backward compatibility (#2717)
* feat(ocsf): configurable schema version for SIEM backward compatibility Add a gateway-configurable OCSF schema version target that downgrades JSONL output for SIEMs that only support older schema versions. AWS Security Lake requires v1.1.0, Splunk CIM Add-On targets v1.1-v1.3. The downgrade filter strips profile-gated fields (ai_model, container, observation_point_id), removes unknown profiles from metadata.profiles, rewrites metadata.version, and adds an unmapped.downgraded_from breadcrumb so auditors can distinguish "no model involved" from "model attribution stripped." Supported target versions (1.1, 1.3) are enforced by an allow-list in the settings registry. Invalid values are rejected with a clear error. The setting flows to sandboxes via the settings bundle and takes effect on the next poll cycle. The shorthand log output is unaffected. Closes #2662 Signed-off-by: Adel Zaalouk <zanetworker@gmail.com> Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> * fix(ocsf): align downgrade with schema version Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Adel Zaalouk <zanetworker@gmail.com> Signed-off-by: Adel Zaalouk <azaalouk@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
c93b2fa7da |
docs(gateway-config): fix stale community sandbox image path (#2800)
* docs(gateway-config): fix stale community sandbox image path Signed-off-by: Yuedong Wu <dwcn22@outlook.com> * docs(sandbox-image): purge remaining stale image references Rebasing onto main surfaced four more instances of the same dead ghcr.io/nvidia/openshell/sandbox path, introduced by commits merged after this branch was opened: three test fixtures (driver-docker, openshell-ocsf, compute::mod) and one user-facing default in the SPIFFE token-exchange Podman demo README. Correct all four to ghcr.io/nvidia/openshell-community/sandboxes/base, consistent with the rest of this fix. Signed-off-by: Yuedong Wu <dwcn22@outlook.com> --------- Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
7cc9551677 |
feat(server): support EC and EdDSA keys in OIDC JWKS validation (#2593)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
08eac8c46d |
fix(sandbox): detect an available login shell instead of hardcoding /bin/bash (#3147)
* fix(sandbox): detect an available login shell instead of hardcoding /bin/bash The built-in default sandbox command and the interactive SSH session hardcoded /bin/bash. Minimal images such as Alpine ship only /bin/sh (BusyBox ash), so sandbox startup failed with an opaque "No such file or directory (os error 2)" that never named the missing binary. Add openshell-core::shell with shell-path constants and a runtime detect_login_shell() that resolves a shell present in the sandbox image ($SHELL if executable, then bash, then /bin/sh). Use it for: - the built-in default command (only the default is remapped; explicit user commands are never rewritten), resolved in the supervisor so it inspects the sandbox filesystem rather than the gateway's - the SSH interactive shell - the SHELL environment variable Also name the program in the spawn error so a missing shell/binary is diagnosable instead of a bare ENOENT. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): drop $SHELL preference in shell detection $SHELL is image/user-controlled and the detected shell is later invoked with `-lc`, so an executable that is not a compatible shell (e.g. SHELL=/bin/false) would pass the executable check and then break command execution even when /bin/sh is available. Resolve only from known shell paths instead. Also add a USR_BASH constant for /usr/bin/bash rather than a string literal in SHELL_CANDIDATES. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): resolve the default login shell in the supervisor (empty command = default) Addresses review: interactive PTY SSH now uses the detected shell, the shell tests are portable across the Windows lane, and default-shell provenance is carried without a new spec field. An omitted command is left empty end to end and resolved in the supervisor, which is the only place that sees the sandbox image: - The CLI forwards the command as-is; the gateway persists an omitted command as empty (no baked /bin/bash -l) and requests a TTY. - MainProcessConfig carries the command empty (the transport now allows it); the supervisor resolves a login shell that exists in the sandbox image (bash when present, otherwise /bin/sh on minimal images like Alpine) and logs the resolved shell. - Interactive PTY SSH (spawn_pty_shell) uses the detected shell; a shared build_ssh_shell_command helper covers the PTY and non-PTY paths, with a deterministic sh-only regression test. - Unix-only shell tests are gated with cfg(unix). An explicit command is always run verbatim. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> --------- Signed-off-by: Akram <akram.benaissi@gmail.com> |
||
|
|
17171cd933 |
refactor(otel): unify compute driver tracing (#2995)
Centralize compute-driver RPC descriptors, stream instrumentation, provider routing, and standalone installation in openshell-otel. Use typed RPC constants so gateway and in-process driver paths cannot panic on unknown operation strings or repeat runtime method parsing. Emit semantic-convention rpc.service and rpc.method attributes, preserve trace context and resource identity across deployment modes, and route both RPC boundary and backend crate spans to each selected driver provider. Leave consumer-dropped watch spans unset while recording observed terminal status, and avoid reboxing untraced external-driver streams. Derive each driver tracing identity from Cargo package and crate metadata and attach its descriptor to the compute-driver registration, keeping provider selection and target routing tied to the registered implementation. Share tracing setup and round-trip test support across Docker, Podman, Kubernetes, and VM, and update the gateway tracing documentation. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
e64b0352e8 |
feat(middleware): broaden HTTP header mutation authority (#3072)
* feat(middleware): broaden HTTP header mutation authority Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): protect response credential headers from mutation Response middleware could write or remove Set-Cookie, WWW-Authenticate, Authentication-Info, and Proxy-Authentication-Info, letting a stage plant or strip credentials the sandbox client acts on. Protect them in both directions, matching the request profile's treatment of Authorization and Cookie, and reserve the x-openshell-credential prefix for responses too. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): reject credential placeholder writes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
172b65e788 |
docs: fix first-network-policy sandbox lifecycle flow (#3140)
The tutorial told users to exit the sandbox and reconnect later, but exiting the interactive shell stops the sandbox's main process and it is not reconnectable under the default restart policy. Switch to the two-terminal flow already used by the github-sandbox tutorial so the sandbox stays running, matching what examples/sandbox-policy-quickstart/ demo.sh actually does. Related: #2998, #2798 Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
6c3980d01a |
fix(middleware): drain websocket session end streams (#3143)
* fix(middleware): drain websocket session end streams Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): diagram websocket stream shutdown Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): keep shutdown diagram in pull request Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
8bc7955263 |
feat(skills): separate public and contributor workflows (#2899)
* feat(skills): separate public and contributor workflows Closes #2736 Publish the four user-facing OpenShell skills from the top-level skills directory, mark contributor workflows internal, and update portability guidance, validation, and documentation. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): clarify public skill audit scope Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): use markdown documentation links Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): align public and contributor guidance Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
03003cd017 |
fix(cli): require ANSI-capable terminal before colorizing (#3121)
* fix(cli): require ANSI-capable terminal before colorizing Follow-up to #3026, raised in review. `auto` treated any terminal as styleable, so `TERM=dumb openshell ...` still emitted escapes into a terminal that renders them literally. An unset TERM had the same problem. This is partly a regression that #3026 introduced. `console`, which drives indicatif and dialoguer, already refused to colorize when TERM is `dumb` or unset, and miette applies the same check through supports-color. #3026 overrides both with its own switch, so it replaced two working checks rather than only failing to add one. tracing and the owo-colors wrapper never had detection, so those two are a gap rather than a regression. Add the capability check to the `auto` branch only, matching console's unix rule: `dumb` is not capable, and an unset TERM is not capable because nothing identifies a capable terminal. Empty is treated as unset, which diverges from console — it reads `TERM=""` as capable since the value is not `dumb` — because an empty value names no terminal type and every other variable here already treats empty as unset. Because the check sits after the explicit branches, `--color always` and FORCE_COLOR still force styling on a dumb terminal, and `--color never` and NO_COLOR still suppress it on a capable one. TERM is a unix signal; Windows consoles enable virtual terminal processing and do not set it, so the check does not apply there. The existing pty test now pins TERM. It previously inherited the ambient value, which would make its outcome depend on the environment now that capability is consulted — CI runners frequently leave TERM unset. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * refactor(cli): combine stream and terminal capability checks Signed-off-by: Evan Lezar <elezar@nvidia.com> * docs(cli): clarify table color behavior Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(cli): cover redirected status table colors Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
857af42a16 |
feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes (#3090)
* feat(vm): support corporate HTTP forward proxy egress for microVM sandboxes The corporate forward proxy machinery from #1792 is driver-agnostic and already merged: openshell-supervisor-network implements CONNECT chaining, NO_PROXY matching, credentials, https:// proxies and corporate CA trust, and openshell-sandbox exposes it as six argv-only flags. Podman gained the driver half in #2245/#2512 and Kubernetes in #2633; the VM driver had none of it, so VM sandboxes on proxy-only networks could not reach any destination requiring the proxy even when policy allowed it. The blocking piece was not proxy logic but delivery: the VM guest init script runs as PID 1 and execs a fixed supervisor command line, and libkrun's krun_set_exec receives an empty argv, so there was no channel for driver-owned supervisor arguments. The supervisor's proxy flags deliberately have no environment fallback, and build_guest_environment merges user-supplied environment, so the guest env is not a safe transport either. Add a driver-authored argument file, mirroring the existing init.d manifest: the driver writes /opt/openshell/supervisor-args into the overlay upperdir on every launch and the guest reads it verbatim, one argument per line, appending it to every supervisor exec. It is written even when empty, which is what makes the channel unforgeable -- the upperdir always shadows the read-only image layer, so an image can neither supply its own arguments nor disable the operator's by omitting the file. Because both launch backends exec the same init script, this covers libkrun and QEMU without touching either. A microVM has no bind mounts or container secrets, so the credential and CA bundle are staged into the per-sandbox overlay the way the gateway JWT already is: credential root-only at 0600, CA at 0644, both rewritten every launch so a removed setting clears prior material, and both deleted with the sandbox state directory. This places the credential at rest in the overlay image on the gateway host, which differs from the Podman secret model and is documented as an explicit security consideration. Validation is fail-closed and shared: a new openshell_core::driver_utils::validate_upstream_proxy_settings holds the pairing rules the Podman driver established, and both the gateway and the driver call it so an invalid table names the offending key instead of surfacing as an opaque driver-readiness timeout. Guest egress leaves through gvproxy, so a proxy on the gateway host's loopback is reachable only through host.openshell.internal; the guest to gateway callback is unaffected. Closes #3088 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): bound the proxy CA read and scope the host-loopback recipe Two review findings on the corporate forward proxy support for microVM sandboxes. The driver read the operator's proxy_ca_bundle with an unbounded fs::read and accepted it on a substring match for the PEM BEGIN CERTIFICATE marker. A special file such as /dev/zero therefore grew driver memory without bound on every authorized sandbox create, and a PEM block holding invalid DER passed the host check but contributes no trust anchor in the guest, so every supervisor would fail after boot with an error attributed to the sandbox rather than to the setting. Move the read into openshell-core as read_upstream_proxy_ca_bundle_file: it reuses the credential reader's bounded-read path (non-regular files rejected on fstat, size capped, read bounded even if the file grows), then requires at least one anchor that RootCertStore::add_parsable_certificates accepts. The supervisor's own reader now delegates to it, so host acceptance and guest acceptance are the same function and cannot drift. The published host-loopback recipe was written for libkrun only. gvproxy NATs host.openshell.internal to the gateway host's 127.0.0.1, but GPU sandboxes run on the QEMU/TAP backend where that name resolves to the TAP host address and the driver's own nftables input chain accepts only the gateway port from the guest — no proxy on the gateway host is reachable there at any bind address, so an operator following the generic recipe lost all proxy-required egress while configuration validation succeeded. Scope the recipe to libkrun in every reference and reject a gateway-host proxy URL when a launch plan resolves to QEMU, naming the reason, instead of booting a sandbox whose policy-approved CONNECTs all time out. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): match the QEMU proxy preflight to the selected TAP host The gateway-host proxy guard added for the QEMU/TAP backend classified the wrong set of addresses in both directions. It ran at the top of configure_qemu_launch_plan, before the subnet allocation that settles plan.host_ip, so it could not compare against the address the guest actually reaches the host on. An operator pointing https_proxy at the sandbox's own TAP host address, such as 10.0.128.1, passed the check, and the driver's nftables input chain — which accepts only the gateway port from the guest — then dropped every policy-approved CONNECT, which is exactly the silent timeout the guard exists to prevent. In the other direction it rejected 192.168.127.254 unconditionally. That address is special only to libkrun/gvproxy; on QEMU/TAP it is an ordinary address that may be routable through the guest's masqueraded egress, so the guard refused a working configuration. Run the check after the launch plan's network allocation, on both the freshly-allocated and already-complete paths, and compare IP literals with that sandbox's selected TAP host. Loopback literals, localhost, and the documented host aliases that write_host_gateway_aliases seeds to the TAP host still classify as the gateway host, and the failure names the address. The gvproxy host-loopback constant returns to being a documentation anchor. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
74960ebfae |
feat(server): add sandbox templates (#2833)
* feat(server): add sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(go-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(rust-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(python-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(typescript-sdk): add sandbox workload template support Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(agents): document sandbox workload templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): support default GPU requests in sandbox templates Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): cap sandbox templates per workspace Signed-off-by: Gordon Sim <gsim@redhat.com> * docs(architecture): document sandbox workload template boundaries Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(cli+sdk): expose sandbox workload template provenance Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): include sandbox template annotations in output Signed-off-by: Gordon Sim <gsim@redhat.com> * feat(sandbox): add label selectors to template listing Signed-off-by: Gordon Sim <gsim@redhat.com> * test(e2e): cover sandbox template failure paths Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): preserve command and ttl when creating sandbox from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(server): add field coverage test for template merge Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): add pagination support to fake client Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): warn on env vars that looks like secrets Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(sdk-ts): propagate sandbox workspace through lifecycle calls Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(docs): update workspace management docs Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(ts-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): support command and tty when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * test(python-sdk): verify command and tty handling when creating from template Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): update docs and ClientInterface Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(server): validate sandbox create specs before I/O Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(go-sdk): guard empty DNS-1123 label validation Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(python-sdk): allow empty template builder mappings Signed-off-by: Gordon Sim <gsim@redhat.com> * fix(cli): align template GPU JSON default output Signed-off-by: Gordon Sim <gsim@redhat.com> --------- Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
cc4ded2088 |
feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(helm): preserve split chart upgrade compatibility Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology. * fix(ci): preserve VM runtime for E2E The Rust cache restores target/ after VM runtime artifacts are staged, overwriting target/vm-runtime-compressed before openshell-driver-vm is built. Stage the compressed runtime outside target and pass that location through OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor. Also locate the Helm split-ownership test repository root from the script path rather than git rev-parse. The test runs in a container where the GitHub checkout can be owned by a different UID and rejected as dubious ownership. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(ci): install yq for Helm ownership test The split-chart ownership regression uses yq to inspect rendered YAML, but the Helm CI container installs only tools declared in mise. Declare and lock yq so mise install --locked provides the test dependency. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> --------- Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> |
||
|
|
7b64c5c88e |
fix(cli): continue multi-item deletes after failures (#3111)
Signed-off-by: Gordon Sim <gsim@redhat.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
b4afcd8a43 |
fix(cli): suppress ANSI color when stdout is not a terminal (#3026)
* fix(cli): suppress ANSI color when stdout is not a terminal
The CLI colorized output unconditionally. owo-colors is built without
its `supports-colors` feature, so `.green()` and friends emitted escape
sequences regardless of destination, and nothing in the CLI read
NO_COLOR. Piping any command through grep or awk matched against bytes
the caller could not see; `forward list` was the case that surfaced it,
where an escape sits immediately before the STATUS word and defeats a
pattern anchored on whitespace.
Add a `color` module holding a process-wide switch resolved once in
run_async, before any output. Command modules import its `Colorize`
trait in place of `OwoColorize`; the method names match, so the ~450
call sites are unchanged, but each consults the switch when it renders
and delegates to owo-colors so the escape bytes stay identical. The two
traits collide by design: importing both in one module is an ambiguity
error, which keeps unconditional coloring from returning.
owo-colors is not the only styled path, and the rest each carry their
own default, so the switch governs them too:
- tracing_subscriber formats with ANSI on, does no terminal detection,
and writes to stdout, so `openshell -v ... | ...` leaked escapes the
same way the tables did. It now takes the setting via with_ansi.
- indicatif and dialoguer both style through console, which has its
own detection but cannot learn about --color. Overriding console's
global switch covers every progress bar and prompt rather than the
specific ones constructed today. Both the stdout and stderr switches
are set, since prompts and progress bars draw to stderr.
- miette renders errors through its own handler, likewise unaware of
--color, so init installs one built from the setting.
Resolution order: `--color always|never`, then NO_COLOR, then
CLICOLOR_FORCE, then whether stdout is a terminal. The decision is made
against stdout even for stderr text, since stdout is what gets parsed;
`--color always` restores styling when redirecting.
Padding is unaffected — the format spec is forwarded to the inner
Display, so widths measure text rather than text plus escapes.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
* fix(cli): resolve color per output stream
Review feedback on #3026.
Resolving one answer from stdout and handing it to every library meant a
redirected stream inherited the other stream's terminal check. Running
`openshell ... 2> build.log` from a terminal wrote escapes into the log,
because console's stderr switch and miette's handler were both given
stdout's answer. That is worse than the behavior before this branch,
where both libraries did their own per-stream detection.
Resolve `auto` separately for stdout and stderr and hand each library
the answer for the stream it writes to: tracing and console's stdout
switch get stdout, miette and console's stderr switch get stderr. The
owo-colors wrapper is the exception, since its call sites are split
across println! and eprintln! and a Painted value cannot tell which
macro will consume it; it styles only when both streams accept escapes,
erring toward plain text rather than risking a redirected stream.
Existing tests could not catch this: Command::output gives both streams
pipes, so a per-stream decision and a single stdout-derived one look
identical. Add a test that puts stdout on a pty and stderr on a pipe,
which fails when stderr is handed stdout's answer.
Replace CLICOLOR_FORCE with FORCE_COLOR. The clicolors spec does not say
how to treat `0`, and implementations that special-case it disagree with
force-color.org, which keys on presence and non-emptiness only. Using
FORCE_COLOR gives it the same rule as NO_COLOR: set and non-empty means
yes, whatever the value. Nothing depended on CLICOLOR_FORCE, which was
introduced earlier on this branch and never released.
Carry the whole style in an owo_colors::Style rather than dispatching a
local enum through a six-arm match, and merge styles when chaining so
`x.green().bold()` emits one `\x1b[32;1m...\x1b[0m` instead of nesting
two wrappers. No call site styles already-styled text, so merging is
safe; the emitted bytes are shorter and there is a single reset.
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
---------
Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
|
||
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
e508c169ea |
fix(helm): honor empty clientCaSecretName for HTTPS-only mode (#2235)
* fix(helm): honor empty clientCaSecretName for HTTPS-only mode Signed-off-by: Yuedong Wu <dwcn22@outlook.com> * docs(skills): document HTTPS-only clientCaSecretName in debug-openshell-cluster Signed-off-by: Yuedong Wu <dwcn22@outlook.com> --------- Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
d5742e01a3 |
feat(cli): add structured output to list commands (#3067)
Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
8ffc6c2a13 |
fix(policy): compose advisor proposals with provider endpoints (#2935)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
5c541e1e0e |
fix(kubernetes): prevent stop-start relay race (#3064)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
9b6d904e88 |
feat(compute): delegate sandbox authentication to drivers (#2968)
* feat(compute): delegate sandbox authentication to drivers Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(kubernetes): align sandbox identity annotation Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(auth): restore sandbox bootstrap coverage Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(kubernetes): satisfy ownership test lint Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |