mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
316
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
71440b28f4 |
fix(policy): refresh pending proposals when the sandbox policy changes (#3923)
* fix(policy): refresh pending proposals when the sandbox policy changes Approving, removing, or undoing a rule, or updating the sandbox policy, changes the inputs every other pending proposal was evaluated against. Only proposals the new policy covered were reconciled; the rest kept their old prover result and review token. The review surface (GetDraftPolicy) therefore showed a stale evaluation, and the first approval of the next proposal refreshed it and failed with FAILED_PRECONDITION, so approving proposals one after another always failed once. Re-evaluate the remaining pending proposals at each policy change, reusing the cached prover result unless the proposal's inputs changed. Approval still rejects a review token that does not match the stored evaluation, so a reviewer holding a pre-refresh evaluation must still refetch it. When a refresh does happen at approval time (inputs changed between fetch and approve), the CLI now explains that the rule was re-evaluated and how to review it, instead of printing the raw gRPC status. Closes #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> * fix(policy): make pending proposal refresh race-safe and bounded Store refreshed evaluations with a compare-and-swap: the store re-reads the proposal, refuses when its rule name, proposed rule, or review token changed since the evaluation read it, copies only the evaluation fields onto the stored record, and updates only if the payload is still the one it read. A refresh can no longer revert a concurrent edit or observation, and the edit path uses the same guard against a concurrent refresh. Bound each refresh to the 32 newest pending proposals; the rest keep the approval-time recheck, which still refuses a stale review token. Operator decisions (approve, approve-all, remove, undo) refresh before responding. UpdateConfig, which holds the gateway-wide sandbox sync guard, and agent-driven auto-approval refresh in a background task instead, one per sandbox with later changes coalesced into a single rerun. Refs #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> * docs(policy): describe proposal rechecks after approvals and approve-all Explain that approving, removing, or undoing a rule rechecks the other pending proposals so they can be approved one after another, when the recheck is deferred or bounded, and what rule approve reports when a proposal changed after it was listed. Show rule approve-all in Run Your First Agent with its security-flag behavior. Refs #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> * fix(policy): refresh pending proposals after a full policy replacement A full policy UpdateConfig (openshell policy set) re-reads the latest revision after its atomic write, finds the revision it just committed, and returns before reaching the pending-proposal refresh at the end of the handler. Pending proposals kept their stale evaluation, so rule get showed the old candidate and the next approval failed with the refresh precondition. Schedule the background refresh right after the commit. Refs #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> --------- Signed-off-by: fede-kamel <fkamelhar@gmail.com> |
||
|
|
fe38637533 |
fix(runtime): recover SSH relays and bound startup diagnostics (#4011)
* fix(runtime): recover SSH relays and bound startup diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): deliver pending relays once per supervisor session Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): satisfy relay delivery clippy diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): bound relay setup with one absolute deadline Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
e21b7fd8cf |
chore(build): remove bundled Z3 support (#3275)
* chore(build): remove bundled Z3 support Signed-off-by: Simon Scatton <sscatton@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(build): preserve vendored Z3 for local gateway artifacts Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
021400be8a |
refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(auth): clarify gateway mTLS behavior Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(auth): include workspace scope in TLS authorization checks Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): bound service auth sandbox names for large PIDs Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
2935e9731b |
fix(gateway): delete finalized ephemeral sandboxes while connected (#3984)
Start driver cleanup after terminal finalization and retain disconnect fallback. Add detached success and failure e2e coverage across supervisor-based drivers. Closes #3938 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
dde8a9a57f |
fix(server): log polled request responses at debug (#3974)
log_response has always logged every gateway response at INFO, including health probes and the GetSandboxConfig and provider-readiness polls each supervisor makes. #3915 demoted the request spans for those polled paths to DEBUG, which stripped the request{method path} prefix from the log line at INFO but left the line itself, so the gateway log fills with bare 'response status=200' lines several times per second. Follow the span's level: polled requests log their response at DEBUG, or WARN on a 5xx so probe and poll failures stay visible. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
912a077bd6 |
feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough Signed-off-by: Derek Carr <decarr@redhat.com> * docs(sdk): add service authorization migration guide Signed-off-by: Derek Carr <decarr@redhat.com> * fix(server): remove stale version import Signed-off-by: Derek Carr <decarr@redhat.com> * docs(upgrade): remove service authorization SDK guide Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): relabel provider readiness TLS mount Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): stabilize exposed service routing Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): support HTTPS service routing Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
7caff12d3c |
perf(otel): stop exporting spans from steady-state polling (#3915)
Store operation spans and request spans for supervisor-polled RPCs (GetSandboxConfig, ReportProviderReadiness) use DEBUG level, so the default INFO filter no longer exports them. The provider credential refresh worker opens its span only when a state has work. Refs #2698 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
252882f37f |
feat(providers): serve sandbox config files on demand (#3832)
* feat(providers): serve sandbox config files on demand Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(providers): defer managed file documentation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(providers): mark managed file api experimental Signed-off-by: Drew Newberry <anewberry@nvidia.com> * style(go): format provider profile fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(providers): preserve legacy environment with managed files Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
5c0c9e446e |
feat(providers): add OCI Generative AI example provider profile (#3904)
Add an example oci-genai inference profile for Oracle Cloud Infrastructure Generative AI through its OpenAI-compatible endpoint. The profile injects a compartment-scoped Generative AI API key as a bearer token only at the regional OCI inference hosts and only under /openai/v1 with GET, POST, and DELETE, so the sandbox never holds the key and the key cannot reach any other OCI surface. The header comments carry the OCI-side setup (create the IAM policy before the key, least-privilege statement, key creation and rotation), the operations verified through the sandbox proxy with a real key (chat completions with streaming, tool calling, vision input, embeddings, and the Responses API), the OCI error messages operators will meet, and the realm and signed-transport caveats. Scoped to the profile YAML per #3906; the only code change is the entry in the profile listing test, which enumerates providers/*.yaml. Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com> |
||
|
|
a875add234 |
feat(server): write gateway OCSF events to JSONL (#3264)
* feat(server): write gateway OCSF events to JSONL Previously, gateway security activity was available only in diagnostic output, and events not associated with a sandbox, such as TLS certificate reloads, had no independent structured record. Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The destination supports daily or disabled rotation, retention limits, queue bounds, and optional schema downgrade to OCSF 1.1 or 1.3. Additionally, existing gateway emitters (TLS reloads, service routing, and policy approval and auto-approval audits) emit structured events, so they reach the JSONL destination, console shorthand, and the affected sandbox's log stream. Records identify the gateway by its configured name in `device.uid` and `device.name`, shared across replicas, with `device.hostname` identifying the replica and `device.os` the gateway's operating system. Metrics and warnings expose known best-effort losses. Refs #2762 Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(mxc): attribute ETW events to gateway Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
acbac9cb79 |
feat(sandbox): add main restart policy (#2798)
* feat(sandbox): add main restart policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden policy-driven restarts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): address restart review feedback Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(sandbox): port restart policy to current runtime Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(sandbox): restart promptly after terminal delivery Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
1358941b81 |
feat(mcp): inspect requests with Tower-selected protocol profiles (#3335)
* fix(sandbox-backend): sort boundary request objects before hashing Sort boundary request objects recursively before hashing so serde_json's preserve_order feature cannot change digest identity. Cover canonical bytes, envelope round trips, and rejection of modified provider values and operations. Signed-off-by: Shiju <shiju@nvidia.com> * feat(mcp): upgrade tower-mcp-types to 0.22.2 Upgrade tower-mcp-types from 0.12.0 to an exact-pinned 0.22.2 and use its inspection APIs to validate MCP requests against the selected revision. Carry inspection metadata into policy evaluation and validate requests after header rewriting, before forwarding. Add explicit support for the sessionless 2026-07-28 revision while keeping 2025-11-25 as the default. Validate per-request metadata and standard HTTP header mirrors, and support discovery, tools, and subscription requests. Delegate batch availability and parameter schemas to Tower. Share typed request names between policy and HTTP checks, retain the local batch resource cap, and centralize MCP policy version parsing and ordering. Keep supported MCP revisions and shared allowlist parsing in the canonical policy schema; core re-exports those types. Tower owns wire-profile semantics, and every supported policy revision must map to the matching inspector profile. Reject duplicate JSON keys, invalid known-method parameters, unavailable methods, and unsupported batches. Keep exact extension allow rules and deny precedence. Document request inspection boundaries and add unit, forwarding, and sandbox coverage. Refs #2174. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): prove authorization at the forwarding boundary Cover March batch denial in both member orders, valid and malformed controls, and audit behavior across both relay entry paths. Exercise real middleware tool rewrites with matching metadata and assert the exact upstream representation or zero forwarded bytes. Verify legacy bodyless SSE GET remains usable while GET tool bodies and unsupported DELETE cleanup are rejected. Clarify request-selected profile and middleware mutation comments without changing production behavior. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): exercise permitted profiles through the sandbox proxy Cover March and June singleton policies and select November and July separately under one endpoint allowlist. Capture upstream tool receipts to distinguish proxy policy denial from an upstream rejection. Extend middleware rewrite coverage to June and multi-version policies, and preserve the sessionless discovery and subscription checks through the shared fixture helpers. Signed-off-by: Shiju <shiju@nvidia.com> * test(kubernetes): box the admission check future Keep the admission test future below Clippy's size limit when the workspace dependency features are unified. Signed-off-by: Shiju <shiju@nvidia.com> * test(mcp): reuse the forwarding fixture identity cache Share the binary identity cache across protocol-profile cases, matching the proxy lifecycle and avoiding repeated hashes of the test executable. Keep procfs authorization and all forwarding assertions intact. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
9f60f55c6b |
fix(server): retry transient sandbox CAS conflicts (#3501)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d4f5034d7f |
fix(gateway)!: make the WebSocket tunnel opt-in (#3727)
The /_ws_tunnel endpoint pipes a WebSocket into the full gRPC service. On a plaintext loopback gateway any web page could open it, since browsers do not apply CORS to WebSocket upgrades. Mount the tunnel only when enable_websocket_tunnel is set (config file, --enable-websocket-tunnel, or OPENSHELL_ENABLE_WEBSOCKET_TUNNEL; server.enableWebsocketTunnel in Helm). BREAKING CHANGE: gateways behind an authenticating edge proxy must set enable_websocket_tunnel = true for CLI edge-tunnel connections. Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
73a181d32f |
docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs Restructure the README as a short path from overview to quickstart to further reading. Move prerelease install steps into the installation guide and telemetry build flags into a new observability page. Fix broken docs links and outdated runtime and credential descriptions. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): describe 0.1.0 as adding new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): add policy prover as a gateway component Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe OpenShell as a runtime for fleets of autonomous AI agents Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): describe policy prover as formal verification Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): name OpenShell Sandbox in component table Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): fold isolation backend into supervisor row Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): mention formal verification in overview Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): use the published OpenCode image in the first-agent guide Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(inference): correct provider examples and readiness Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(providers): correct Google binding and provider selection Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(readme): sharpen value prop, how it works, and explore further Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): use example Anthropic profile and add policy advisor step Import the example Anthropic profile, which now allows OpenCode, instead of editing it with sed. Add a step that shows how to review and approve mechanistic policy proposals as the agent needs more access. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): add OpenRouter example for OpenCode Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): run OpenCode against OpenRouter with a free model Add an example OpenRouter provider profile scoped to OpenCode and switch the first-agent guide to it, using a free Nemotron model so readers do not need OpenRouter credits. Revert the OpenCode binary added to the example Anthropic profile. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): link first-agent guide and add agent skills section Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
585bcf0b82 |
(Fix) ha sandbox resilience with k8s (#3644)
* fix(server): retry internal store updates on version conflict - Re-read and reapply internal CAS updates (expected version 0) up to 5 times on conflict - Client-supplied versions still fail on conflict - Add concurrent-writer test Signed-off-by: divesh <dgude@nvidia.com> * fix(kubernetes): keep sandboxes running while the supervisor reconnects - Treat a running but not-Ready supervisor Pod as degraded, not unavailable - Report degraded sandboxes as not ready without suspending them - Suspend only when the supervisor Pod is missing, terminated, or deleting - Add availability test Signed-off-by: divesh <dgude@nvidia.com> * fix(kubernetes): skip fence generation check for suspended sandboxes - Check the fence generation only for running or bootstrapping sandboxes - Stop re-suspending stopped sandboxes and logging a warning every reconcile Signed-off-by: divesh <dgude@nvidia.com> * fix(kubernetes): rank degraded supervisor above unknown dependencies - Report a degraded supervisor as unavailable even when another dependency read is unknown - Extract dependency aggregation and readiness mapping into pure helpers - Add mixed degraded and unknown regression test Signed-off-by: divesh <dgude@nvidia.com> --------- Signed-off-by: divesh <dgude@nvidia.com> |
||
|
|
9244868056 |
docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe updated security architecture neutrally Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: highlight new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandboxes): clarify how to disconnect Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align architecture and guides with current navigation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(extensibility): streamline extension authentication guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
4688061882 |
fix(sandbox): deliver complete exec output before success (#3688)
* fix(sandbox): preserve exec output through channel close Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(exec): propagate output delivery failures before exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(cli): clarify exec output delivery failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
7691f88e0e |
perf(server): enable WAL for the SQLite store; relax sync only for SSH session issuance (#3543)
* perf(server): enable WAL and NORMAL sync for the SQLite store On-disk SQLite stores ran with sqlx defaults: rollback journal (`journal_mode=delete`) and `synchronous=FULL`. Every autocommit write paid several fsyncs and blocked readers while it held the lock, so gateway hot paths made of many small writes serialized on disk latency. The clearest case is `openshell forward service`, which mints and revokes an SSH session token around every forwarded TCP connection: two commits per connection, tens of milliseconds each on a virtual disk, wall clock linear in the number of concurrent connections, and enough queueing that bursts hit the per-sandbox connection cap and get refused. Switch on-disk databases to WAL with `synchronous=NORMAL`. The mode change runs once on a single connection before the pool opens: entering WAL needs exclusive access to the file, so doing it up front means pool connections only ever re-apply the pragma to a file already in WAL mode, and a failure surfaces as one clear connect error. The first start after upgrading an existing database therefore needs the file to be otherwise unopened. `synchronous` is applied through the connect options on every pooled connection. In-memory databases keep their defaults. A crash can now roll back the most recent transactions without corrupting the database, which is the standard WAL trade-off and fits the single-node scope of the SQLite backend. Tests cover a fresh store, an existing rollback-journal file that must be switched on connect, sidecar permissions, and concurrent readers under a burst of insert-then-update writes. Architecture, configuration and Helm docs describe the durability trade-off, the sidecar files, and the local filesystem requirement. Signed-off-by: Jason T. Greene <jason.greene@redhat.com> * fix(server): keep SQLite commits durable except SSH session issuance WAL with synchronous=NORMAL can roll back acknowledged commits after a power loss or kernel crash, including SSH session revocations and other authorization-tightening writes. Run the main pool with synchronous=FULL so every acknowledged write is durable; in WAL mode that is a single fsync of the WAL per commit. Add Store::create_relaxed for inserts that are safe to lose, and use it only for SSH session issuance: a dropped token just fails validation. On file-backed SQLite it runs on a dedicated single-connection pool with synchronous=NORMAL. Both pools share one WAL, so the next FULL commit also makes earlier relaxed commits durable. Postgres treats it as an ordinary durable MustCreate insert. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Jason T. Greene <jason.greene@redhat.com> Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Co-authored-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
0518bd4c83 |
fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): propagate non-expiring local sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): distinguish main exit from transport loss Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): bound sandbox connect recovery Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
a8f98ec09d |
fix(auth): remove legacy sandbox JWT admission (#3562)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
0b351c4a9b |
fix(helm)!: reduce gateway Secret privileges (#3616)
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace The Kubernetes Secrets credential driver now stores every credential in its configured namespace in all workspace modes and rejects handles that reference any other namespace before contacting the Kubernetes API. The gateway reaches credential Secrets through the Role in that namespace; this allows removing the Secret rules from the ClusterRole. - Remove the workspace_mode, gateway_id, and allow_reference_namespace driver settings and stop rendering them from Helm. Configurations that set them fail at startup. Existing credential state is not migrated. - Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a dedicated credential namespace. The namespace is kept on uninstall, adopted by a reinstall of the same release, and left untouched when something else owns it. - Update the gateway config reference, Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(helm): reduce gateway Secret permissions Remove the gateway's Secret list permission in every workspace mode and grant source Secret reads through a Role in the sandbox namespace. Bootstrap Secret cleanup deletes Secrets by exact name instead of listing them. - Grant get on the copied client TLS and image-pull Secrets through a Role in the sandbox namespace. The ClusterRole keeps get and patch on those names for the ownership check and server-side apply into workspace namespaces. - Delete sandbox and supervisor bootstrap Secrets by exact name, derived from the runtime generation recorded on the Sandbox and, on restart, the target generation, tolerating 404. The generation annotation is cleared only after cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so garbage collection removes any generation the driver does not name. - Drop Secret list from the ClusterRole and the shared-mode sandbox Role. - Extend the managed e2e RBAC checks to Secret list. - Update the Kubernetes setup and sandbox runtime docs, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(driver-kubernetes): stage workspace Secrets per runtime generation Every Secret the Kubernetes driver writes into a workspace namespace is now scoped to one sandbox runtime generation, immutable, and created with create only, so the gateway never reads, patches, or adopts an existing Secret there. This removes the gateway's cluster-wide get and patch on the copied client TLS and image-pull Secret names. - In managed mode, create an immutable copy of each configured image-pull Secret per generation, named os-pull-<id>-<generation>-<n> and owned by the generation's workload and supervisor Pods. Pods and the restarted Sandbox template reference those names, and generation cleanup deletes them by name. A Secret already holding a generation name fails the create. - Outside shared mode, stage the gateway client TLS material into the supervisor bootstrap Secret instead of copying the client TLS Secret into the workspace namespace. - Remove the fixed-name TLS and image-pull copies, the target ownership read, and the ClusterRole get and patch rule on the copied names. Source reads stay in the sandbox-namespace Role. - Update the managed e2e to expect generation image-pull Secrets and client TLS material in the supervisor bootstrap Secret, and to check that the gateway cannot read the copied names in workspace namespaces. - Update the gateway config and compute driver references, Kubernetes setup and sandbox runtime docs, compute-runtime architecture, driver README, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(helm)!: grant operator-mode Secret permissions through the workspace chart The operator-mode gateway ClusterRole grants no Secret permissions. The openshell-workspace chart Role, installed in each operator-managed namespace, grants the gateway create and delete on Secrets for sandbox runtime generations. Operator-managed namespaces require the workspace chart. - Fail the chart tests on any ClusterRole rule that includes Secrets in operator and shared modes. - Install the workspace chart when the operator e2e provisions a namespace, and assert that the gateway has no Secret permissions in a namespace without it. - Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime architecture, and cluster debugging skill. Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
123d95e2ed |
fix(pagination): document list contract and harden SDK pagers (#3279)
* fix(pagination): document list contract and harden SDK pagers Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): bound pager token history Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): preflight token history limits Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): paginate sandbox providers Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): harden TypeScript pager Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): expose provider pagers in SDKs Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * docs(pagination): describe a uniform list contract Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): use stable provider cursors Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * docs(pagination): clarify mutation semantics Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(cli): expose sandbox provider pagination Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(proto): refresh pagination schema inventory Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(proto): refresh rebased schema inventory Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
95632406cd |
fix(server): make HA sandbox create and HA e2e tests reliable (#3635)
* fix(server): retry runtime identity persistence on sandbox create With multiple gateway replicas, another replica can update a new sandbox record between the create path's read and its compare-and-swap write of the runtime identity. The create path made one attempt, so the conflict failed the request and deleted the backend sandbox. Persist the identity through the same retrying helper that start and startup recovery use, which rereads the record and retries while the sandbox stays in the same generation and a Provisioning or Ready phase. Signed-off-by: Kris Hicks <khicks@nvidia.com> * ci(e2e): run Kubernetes HA tests one at a time The two HA tests scale and roll the shared gateway Deployment. Their in-process lock does not apply under nextest, which runs each test in its own process, so one test could delete a gateway pod while the other was executing through it. Put both tests in a nextest test group limited to one thread in the e2e-kubernetes profile. The override matches test names because a binary() filter fails in workspaces that lack the HA test binary. Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
49df4d7ca7 |
fix(server): drain supervisor ownership cleanup on shutdown (#3547)
Closes #3546 Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
bdffa102c3 |
feat(api): add durable exec launch admission (#3324)
* feat(api): add durable exec launch admission Fence duplicate exec launches with keyed durable admission and producer-owned terminal completion. Keep uncertain launches unresolved and never replay output or interactive input. Part of #3051 (phase 4a). Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(api): fence exec identity across authorization lookups Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
1e34e8c576 |
fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): address resource admission review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): reserve driver-owned admission labels Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): clarify workspace admission label Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): clarify resource admission failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): configure resource admission fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): retry forbidden admission lookups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): preserve external driver admission defaults Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
470a34635d |
fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag. Partially addresses #3055 (cursor/resume follow up separately). Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * refactor(server): group per-sandbox log bus state and stamp sequence numbers Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(proto): add resume cursor fields to sandbox watch API Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): stamp watch cursors from a shared per-sandbox sequence Allocate cursors from a single SeqAllocator shared by the log and platform event buses, so a sandbox's merged watch stream carries unique, strictly increasing cursors. A single resume_after_cursor can then unambiguously locate a client's position across both sources. Rewrite both publish paths to allocate the sequence, stamp event.cursor, send, and append to the tail under one lock. This removes the previous get_mut().expect() TOCTOU race where a concurrent remove() between the two lock sections could panic. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): serve WatchSandbox resume from cursor with gap detection Add tail_after() to the log and platform event buses, returning every buffered event newer than a client's resume cursor. Each PerSandbox now tracks last_trimmed_seq (the highest seq it has evicted) so a resume is reported as an unrecoverable ResumeGap only when this bus dropped an event the client still needs. Judging gaps by evictions, not by the tail's oldest seq, is required under the shared cursor space: each bus's tail is non-contiguous in the global sequence because the other bus owns the missing seqs, so comparing against tail.front() would flag false gaps. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): resume WatchSandbox from cursor across log and platform buses Wire resume_after_cursor into the watch producer. On a non-zero cursor, replay events strictly after it from both the log and platform buses, merge by shared cursor, and emit in order before entering the live loop. A trimmed range on either bus is an unrecoverable gap and terminates the stream with OUT_OF_RANGE carrying the requested and earliest-available cursors, distinct from recoverable lag which warns and continues. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): cover WatchSandbox cursor resume paths Add handler-level tests for the resumable watch stream: replay strictly after the client cursor, merge log and platform events in shared-cursor order, suppress duplicates when resuming at the latest cursor, and terminate with OUT_OF_RANGE when the requested cursor has been trimmed. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * docs(api): document WatchSandbox loss-awareness and resume Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): deliver watch events once and harden cursor teardown Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(sdk): add loss-aware resumable watch_logs to Rust SDK client Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): keep watch cursors monotonic across teardown and restart Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): merge live watch sources by cursor before emission Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): bind watch cursors to a cursor space and merge tail sources Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): revalidate the watch cursor space after collecting replay Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): update the public RPC schema fingerprint for the string cursor Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): synchronize the watch live-order test with the end of initialization Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(sdk): use canonical sandbox name in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(sdk): guard canonical-name addressing in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): fix public rpc schema Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): reconcile watch resume rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): bound interactive relay cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
718dba3430 |
fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
7139df8ca5 |
fix(exec): preserve output after stdin EOF and verify stream completion (#3359)
* fix(exec): preserve output after stdin EOF and verify stream completion Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(exec): cover fair duplex progress and live SDK completion Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(sdk): compare large exec buffers with native equality Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(exec): use workspace-scoped sandbox name in EOF regression Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
50230616d5 |
refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c8a4ff5f19 |
fix(kubernetes): bind bootstrap to runtime identity (#3531)
* fix(kubernetes): bind bootstrap to runtime identity Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(compute): compensate runtime binding failures Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(compute): clean up backend on store failure Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(compute): merge runtime binding after start Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(auth): bind restarted sandbox sessions Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor(compute): fold runtime binding into authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
0a770d9173 |
feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com> |
||
|
|
2493d415c2 |
feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation Closes #3057 Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK. Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): fail fast on negotiation errors Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(extensions): validate gateway handshake metadata Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(go-sdk): re-export extension kind constants Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(extensions): fail fast on credential handshake rejection Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
484f0768fc |
fix(server): serialize sandbox restart authentication (#3485)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
1905069948 |
feat(sandbox): expose services during creation (#3439)
* feat(sandbox): expose services during creation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): refresh Codex credentials in gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): normalize create-time service URLs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): roll back failed service exposure Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(example): simplify Codex provider setup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(example): separate provider setup commands Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden create-time service exposure Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(example): bundle Codex provider profile Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(example): allow npm-installed Codex binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
fa0bfa490e |
fix(server): close SQLite stores in policy tests (#3466)
Signed-off-by: Kaylee Lubick <klubick@nvidia.com> |
||
|
|
17ce738bfb |
fix(ci)!: remove gateway callback listener dependency (#3365)
* fix(ci): repair post-merge release canary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(packaging): bootstrap canary runtime prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): collect macOS VM diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): pin libkrun-compatible macOS runner Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(canary): limit macOS smoke test to package startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute)!: remove gateway callback listeners Run Docker supervisors on host networking so they use the operator-configured primary gateway endpoint. Remove the unused compute-driver callback listener negotiation and listener-scoped routing machinery. BREAKING CHANGE: The ComputeDriver API no longer exposes GetGatewayListenerRequirements or GatewayListenerRequirement. External drivers must regenerate bindings and connect supervisors to the configured primary gateway endpoint. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox runtime image in launcher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(podman): exercise production endpoint selection Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): route supervisors to reachable gateways Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): align Podman endpoint fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host aliases for supervisors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): align sandbox host gateway pin Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): address Docker fixtures by bridge IP Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): serialize sandbox lifecycle cases Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): host Docker TCP fixture with gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use loopback for host-network supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d91b1999a0 |
feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(cli): update forward color fixture for workspace scope Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): use sandbox names for settings lookup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox request fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): harden sandbox mutation handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): update rebased sandbox references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(api)!: standardize canonical entity references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(api): codify protobuf API conventions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve workspace selector semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): restore workspace selector parity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve descriptive name fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): update e2e request fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): omit workspace selector during bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): remove proto convention checker Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): refresh schema fingerprints after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): use canonical provider receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
e38d7254e6 |
fix(policy): reject unknown endpoint security modes (#3187)
* fix(policy): reject unknown endpoint security modes Closes #3046 Validate TLS, enforcement, and access values across policy and provider profile ingress, and prevent runtime parsing from falling back to audit for unknown enforcement values. Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> * fix(policy)!: use enums for endpoint security modes Replace the public TLS, enforcement, and access strings with protobuf enums and carry the typed values through policy composition, provider profiles, drivers, and runtime conversion. Preserve the documented YAML spellings, reject unknown and invalid numeric enum values consistently, and update generated Go bindings, SDK conversions, tests, and policy documentation. Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> --------- Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> |
||
|
|
1d010f4187 |
feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation Closes #3145 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): bound startup failures and preserve activation history Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): enforce deadlines on startup RPC attempts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): isolate provider auto-create policy fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): remove unnecessary fixture string delimiters Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): capture startup logs before ephemeral cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): preserve credential revocation after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): retain provider revision helper for endpoint reports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): import provider object trait in production Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): preserve admission across boundary extraction Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): use existing boundary discovery imports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): distinguish workload and supervisor containers Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission schema with timestamp migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with typed deletion schema Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with mutation request IDs Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(supervisor): adapt local startup fixture to admission state Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): deduplicate startup quarantine diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): show configuration blockers in sandbox notes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): expire provisioning repair attempts after five minutes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): reconcile admission with provider readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): separate configuration summaries from full diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): shorten invalid configuration note Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): refresh schema inventory after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
04146692d9 |
refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules Twelve modules under crates/openshell-providers/src/providers/ were never declared in providers/mod.rs, so they have not been compiled since the plugin registry was narrowed to the two adapters it still registers. Four of them (generic, gitlab, opencode, outlook) key off provider type identifiers that normalize_provider_type already retired. Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new registers. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(providers): load example profiles from providers/ at test time Add an example_profiles module that reads the YAML under providers/ from the source checkout at run time and parses it with the existing profile loader. It is gated behind a new non-default example-profiles feature so it is available to this crate's own tests and, once wired into dev-dependencies, to gateway and CLI tests, while never reaching a release binary. Nothing consumes it yet; later changes move the test fixtures off the compiled catalog and onto these files. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test: source profile fixtures from providers/ instead of the compiled catalog Unit and integration fixtures reached into builtin_profiles() to get a profile to work with, which ties the tests to the compiled catalog rather than to the files an operator would import. Point them at the example_profiles loader instead. The profiles.rs tests keep their coverage unchanged and become explicit golden tests over providers/*.yaml: they are what keeps those files valid once nothing compiles them. builtin_profiles_are_sorted_by_id widens into a test that the whole example set parses, sorts, and lints clean as one catalog. The CLI fake gateways serve the example profiles the way a real gateway serves what an operator imported, through a shared helper. No behavior change: the gateway still loads the built-in source by default and serves the same profiles. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document the example provider profiles Each file under providers/ now opens with a header naming its expected client binary identities, the image layout those paths assume, the credential scope, the endpoint access it grants, and a smoke test. Add a README covering the import commands and why a profile should be copied and edited rather than imported unchanged. Several of these profiles bind network access to paths that only exist in the OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server, /usr/lib/node_modules. Imported unchanged into another image the profile matches nothing: the catalog still advertises it, but the credential is never injected and the traffic is denied. The headers say so where it applies. Comments only; the profile schema has no documentation fields and all fifteen files still lint clean. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(server): seed unit-test state with the example provider profiles Server unit tests inherited the builtin + user source default from ServerState, so around a hundred and forty assertions about github, openai and the rest resolved against the compiled catalog. Point the test state at the user source alone and import the example profiles from providers/ into its store first, the way an operator would. Every one of those assertions keeps passing unchanged, which is the point: it proves the gateway behaves identically with an imported catalog before the default moves. Four tests asserted builtin-source semantics specifically. Under import-only the only profiles a gateway cannot edit are the ones a non-user source vends, so they now exercise a source-managed profile composed with the user source; the read-only guard they cover is the one that still applies to interceptor catalogs. The list test asserts the imported profile's user/platform identity instead of builtin with an empty scope. test_server_state_with_user_only_github_profile is gone: the default test state now is a user-only gateway with github imported. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli): resolve credential suggestions from the gateway catalog The --env credential warning scanned a profile table compiled into the CLI, so its suggestions described the binary rather than the gateway the user is talking to: a profile the gateway does not serve was suggested anyway, and an imported custom profile never was. Fetch the catalog from the connected gateway instead. The warning moves out of argument parsing and into sandbox create and sandbox template create, where a client already exists. A catalog fetch failure is not fatal — the warning degrades to its generic form rather than blocking sandbox creation. Extract the ListProviderProfiles paging loop from provider list-profiles into a shared fetch_provider_profile_catalog, which the profile-driven paths now share. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: infer providers from the gateway catalog Command-to-provider inference went through a hardcoded alias table in openshell-providers, a second copy of the built-in catalog's identifiers. A custom profile could never be inferred no matter what binaries it declared, and the table drifted from the profiles it mirrored. Infer from the connected gateway's catalog instead: match the command's basename against each profile's ID and against the basenames of the binaries the profile authorizes. A profile that names /usr/bin/claude is the profile for running claude. The match must be unique — where several profiles claim a command, the user names one with --provider — and an empty catalog infers nothing. No command is special-cased. `binaries` is the operator's authorization statement, so a profile that declares a binary claims the command that runs it, whatever that binary is; narrowing that belongs in the profile rather than in a list compiled into the CLI, which could never cover an unbounded catalog anyway. Breaking: the retired aliases stop resolving, and commands the old table never listed can now infer. git, pip and uv are declared by the github and pypi example profiles, so they infer where they previously did not. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers,server)!: resolve provider profiles by exact ID normalize_provider_type was a hardcoded alias map — gh to github, claude to claude-code, vertex to google-vertex-ai — and a second copy of the built-in catalog's identifiers, independent of the YAML it mirrored. It let a provider type resolve to a profile the operator never named, and it made the built-in IDs behave as a reserved namespace. Remove it, along with detect_provider_from_command and the alias-normalizing ProviderRegistry::inject_env. A provider type now names a profile exactly: - the effective catalog resolves an ID or reports it absent, with no alias retry - plugins activate only for the ID of a profile the gateway resolved; a provider with no resolvable profile gets no plugin projection, instead of falling back to an alias guess - the CLI surfaces the gateway's not-found instead of retrying under an alias - telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets go with the aliases that were their only source Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve. Use the profile's own ID. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(server): fail closed when a provider's profile is absent A sandbox composed from a provider whose profile the gateway cannot resolve started anyway: the credential and policy builders warned and skipped, so the sandbox came up carrying none of that provider's credentials or network policy. The operator learned about it later, as a denied connection or a missing environment variable, rather than as the configuration error it is. Check at the two composition boundaries — CreateSandbox and AttachSandboxProvider — right beside the catalog snapshot already taken there, and reject with a bounded diagnostic naming the provider, the profile it refers to, and the import command that supplies it. Scope-aware: a platform-scoped provider is told to import with --global. Read paths are untouched. ListProviders, GetProvider and profile export keep working so an operator can see and recover an affected provider, and the shared policy builders keep their warn-and-skip for the diagnostic paths that also reach them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(server): scope the vendor base-URL pin to declared endpoints provider_profile_endpoints_are_active withheld a credential and the provider's policy layer when an openai or anthropic provider pointed its client somewhere other than the public vendor endpoint. The guard was keyed on profile.source == "builtin" and on those two profile IDs, so it protected only profiles OpenShell shipped. Once profiles are import-only no profile is ever builtin, and the guard would silently stop applying — including to an operator who imported providers/openai.yaml verbatim. Key it on the profile instead of on where the profile came from. A profile's endpoints are the boundary its credential is bound to, so if the provider configures a *_BASE_URL pointing at a host the profile does not declare, the profile no longer describes where that credential goes and is treated as endpointless. Host matching reuses the DNS-label-aware matcher in openshell-core, so wildcard endpoints such as Vertex's *-aiplatform.googleapis.com resolve correctly. A profile with no declared endpoints has no boundary to contradict, and a config value that names no host is not a redirect this can reason about. Both keep the profile active. The control now covers every endpoint-bearing profile, including an operator's own, and the last "builtin" string leaves the gateway. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): import example provider profiles during gateway bring-up The e2e suites create providers from github, openai, nvidia, claude-code and google-cloud, and the GitHub lane runs a real git clone through the profile's binary attribution. A gateway serves only the profiles an operator imported, so the lanes have to import them. Add e2e_import_example_provider_profiles to the shared bring-up helpers and call it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is healthy and registered. It runs the same command the upgrade notes give operators, against the repository's own providers/ directory, so the lanes exercise the documented path rather than a test-only shortcut. Lands before the default changes: a same-ID user profile already shadows a built-in, so importing works today. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers)!: make provider profiles import-only OpenShell compiled fifteen provider profile YAML files into every release binary and selected them by default, so a fresh gateway published a catalog it was never configured with. Those profiles are not image-neutral: their binary selectors name paths that exist in the OpenShell Community image, so changing the sandbox image could make a profile inert while the catalog still advertised it. Every endpoint, credential name and binary path in providers/ was also effectively part of the 0.1.0 public contract. A gateway's catalog is now exactly what an operator imported: - providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone from the openshell-providers API - the default provider_profile_sources is [{ type = "user" }], and a gateway with nothing imported reaches ready and serves an empty catalog — an empty catalog is a valid state, not a startup failure - the builtin source type is removed from the configuration schema, and a gateway.toml that still names it is rejected at parse time with the import command rather than an unknown-variant error - nothing reserves the canonical identifiers any more, so github, pypi, anthropic and the rest import at their own IDs and the imported profile is the only definition for that ID; the static_fallback that kept a shipped definition resident behind an imported one is gone with the source that produced it A collision between an interceptor-vended profile and an imported one still fails closed, which is the behavior that shadowing quietly bypassed for built-ins. Operators upgrading should export the profiles their deployment relies on before upgrading, or copy them from the providers/ directory of the matching release tag, then import them at the scope their providers use. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document import-only provider profiles Rewrite the provider profile documentation around a catalog the operator builds rather than one the gateway ships. - gateway-config: the default is [{ type = "user" }], the builtin source type is gone and rejected at startup, and an empty catalog is a valid ready state - profiles: replace the built-in profile table and its shadowing semantics with the import workflow, and say plainly that a profile whose binary paths do not match the image is inert - manage-providers: replace the two fixed provider-type tables with `provider list-profiles`, and describe catalog-driven command inference - inference-routing: drop the "still loads built-in profiles for compatibility" transition text; its migration walkthrough is now the normal path - quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex tutorials: import the profile before creating the provider, since nothing resolves without it - release notes: a 0.1.0 migration section covering export-before-upgrade, importing from the release tag's providers/ directory, what happens to a provider whose profile is missing, and the removal of the legacy type aliases - README, architecture, the openshell-cli and debug-openshell-cluster skills, and the governance interceptor example follow the same change; the example's smoke assertion now imports a profile to prove an authoritative interceptor hides it Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding Addresses two findings from review (GATOR-c1bd9867-02 and -03). The provider profile catalog lookup collapsed its error into an empty catalog, so an authorization, availability or transport failure on ListProviderProfiles was indistinguishable from a gateway that genuinely has no profiles. Command inference then resolved nothing and the sandbox was created without the provider it needed, deferring the failure to the workload. Keep the lookup's outcome instead of discarding it. The credential warning is advisory and still degrades to its generic form, but inference now consults the catalog only when there is a command to resolve and surfaces the lookup failure when there is, naming the gateway and pointing at explicit --provider selection. Two regression tests cover it through a fake gateway whose ListProviderProfiles returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails without sending a provider-less create request, and a sandbox with no trailing command still succeeds because it needs no catalog. The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes deliberately skip gateway registration and start the gateway without a TLS client CA, so no mTLS identity exists and no token has been acquired when the import runs; the wrapper exited during setup. Skipping the import is not sufficient because the provider tests now require the claude-code profile. Add e2e_register_oidc_admin_session, which mints an administrator token with Keycloak's password grant — the same grant the OIDC test helpers use — and writes the gateway metadata and token bundle that an interactive login would have stored, so the import runs as an authenticated administrator. The mTLS lanes keep the direct import unchanged. Both affected wrappers are covered: with-podman-gateway.sh had the same defect as with-docker-gateway.sh. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: remove command-derived provider attachment Addresses GATOR-c1bd9867-01. A profile's `binaries` list authorizes a binary to reach that profile's endpoints. It is not a statement that running the binary asks for the provider, and reading it as attachment intent let a command silently gain provider authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash` resolved an existing aws-s3 provider and attached its credential-backed capability and network policy to the shell and its descendants. Attachment of an already-created provider never prompted, so the confirmation flow did not guard it. The distinction that would make inference sound — whether a declared binary is a profile's client or merely a permitted runtime — cannot be expressed: NetworkBinary carries only a path, and the field that encoded it was removed in 0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant does not help, because uniqueness measures how sparse the catalog is rather than what the user intended, and a compiled list of "generic" commands could never cover an unbounded operator catalog. Remove trailing-command inference rather than approximate it. A provider is attached only when named with --provider, which still creates a missing provider from local discovery when the name matches an imported profile ID. Sandboxes with no providers remain a normal, fully supported state. Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving authority from the catalog, its only remaining use is the advisory credential warning, so a failed lookup degrades that warning instead of blocking creation. The regression test now asserts that an unreachable catalog still creates the sandbox and attaches nothing. While repurposing the deduplication test, a pre-existing defect surfaced: repeating a name in --provider auto-created it twice and the second attempt failed with "provider already exists", because the explicit pass never consulted the set of names it had already handled. Guard it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): load the OIDC token and trust the gateway when seeding profiles Addresses the carried finding GATOR-c1bd9867-03. e2e_register_oidc_admin_session established a session the CLI never used. The CLI decides whether to load a stored bearer token by matching on the auth_mode field of the gateway metadata alone; the helper omitted that field, so the metadata fell through to the default arm and oidc_token.json stayed on disk unread. The profile import went out unauthenticated exactly as it had before the helper existed. The helper also installed no trust anchor, leaving the CLI unable to verify the gateway's self-signed serving certificate. Write auth_mode = "oidc" so the stored token is loaded, and install the CA at <gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or key on disk the CLI falls back to CA-only server verification and authenticates with the bearer token, which is what these lanes need because they start the gateway without --tls-client-ca. Certificate verification stays on; no insecure transport override is introduced. Assert the session before anything depends on it. ListProviderProfiles is annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded and accepted. The helper now fails at that point, naming the gateway config directory and echoing the CLI output, rather than letting the defect surface later as an opaque profile import error. Both affected wrappers pass the PKI directory and CLI binary the helper needs; the mTLS lanes are untouched. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): request the OpenShell scopes when minting the admin token The OIDC lanes failed at the session assertion added in b25dd0298 with PERMISSION_DENIED and "scope 'provider:read' required". Authentication was working -- the gateway logged ListProviderProfiles at gRPC status 7, which it can only reach once the bearer token has been loaded and accepted. The token simply carried no OpenShell scope. sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional client scopes on the openshell-cli client in scripts/keycloak-realm.json, so Keycloak mints them only when the request asks for them. The password grant here asked for nothing, leaving the realm defaults (openid, profile, email, roles, web-origins, acr) and an access token that authorizes no RPC. Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already does for its administrator tokens. openshell:all is SCOPE_ALL in crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls without enumerating them. Verified against the realm: the token goes from "email profile openid" to "email profile openshell:all openid". Record the same scopes in metadata.json. The helper writes the bundle openshell gateway login would have stored, and oidc_scopes is the field that login path reads back, so leaving it out would misdescribe the stored token. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): seed provider profiles once per VM lane gateway The VM lanes failed on the second test target with "custom provider profile 'anthropic' already exists" for all fifteen example profiles, and the import exited non-zero. e2e_import_example_provider_profiles was called from run_e2e_test, so it ran once per target -- four times against one long-lived gateway. That was harmless while import overwrote silently, but profiles are now import-only: ImportProviderProfiles is create-only and reports an existing id as an error-severity diagnostic, with no overwrite flag on the request. The first import therefore succeeds and every later one fails. Hoist the call to just after the conformance run, which is where the docker, podman and kube lanes already seed their catalogs. The profiles persist for the gateway's lifetime, so every target still finds them, and the E2E_TEST_OVERRIDE path is covered by the same single call. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): name the provider explicitly in the OIDC workspace-user test user_can_create_sandbox_with_inferred_provider_command reached the gateway once the admin session was fixed, and then failed: it asserted a missing-provider error, but the sandbox was created and died provisioning with "failed to spawn sandbox entrypoint process 'claude-code'". The test drove provider resolution by passing claude-code as the trailing command and relying on the CLI to infer the provider type from it. That inference is what this branch removed, so the trailing word is now nothing but an entrypoint, and the image has no such binary. The regression the test guards is not inference itself: it is that a workspace user resolving a provider is not gated behind Platform Admin. Name the provider with --provider, the only remaining way to attach one. The CLI still has to fetch the claude-code profile before it can auto-create the provider, so the lookup a workspace user must be allowed to make still happens, and auto-creation still stops at the non-interactive branch with "missing required provider". Both assertions therefore keep their meaning. Rename the test and rework its comments to describe what it now exercises; the old name would otherwise outlive the behavior it was named for. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(providers): reconcile import-only profiles with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
a72351d370 |
fix(policy)!: require explicit L7 append targets and scope (#3380)
Make allow and deny appends identify a rule and endpoint and declare every affected binary and port. Reject incomplete, stale, ambiguous, and provider targets atomically so a small append cannot silently change a broader scope. Update CLI previews, wire requests, Go types and generated bindings, SDK regressions, operator documentation, and live policy-update coverage. BREAKING CHANGE: AddAllowRules and AddDenyRules require L7RuleTarget instead of host and port. CLI appends require a rule name and explicit binary scope. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
9708ba9999 |
feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes Record exact provider mutation targets in shared configuration operations. Require authenticated evidence that credentials, effective policy, and the workload launch environment have been installed before reporting readiness. Add bounded CLI and Rust SDK status and wait support, preserving ordinary revision-scoped references for existing processes. Verify new-client rotation and acknowledged detach revocation without external-stable resolver changes. Signed-off-by: Shiju <shiju@nvidia.com> * fix(providers): align readiness times with protobuf contracts Represent readiness receipts, status, and operation times with Timestamp and report intervals with Duration. Reserve the scalar field tags, update all consumers and generated bindings, and preserve timestamp presence and nanosecond identity through storage and client validation. Qualify both empty-map constructors in the Linux boundary test so its module compiles while retaining the explicit default required by Clippy. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): preserve provider mutation storage uncertainty Recognize the gateway's exact structured storage-uncertainty reason for provider attach, detach, and update. Explain that the change may already be saved and must be reconciled before retrying, without exposing server messages or metadata. Preserve uncertainty ahead of generic retry hints. Exercise saved mutations through the CLI and verify single submission, redaction, missing receipt handling, and untrusted error-detail rejection. Document the recovery guidance for users and the public CLI skill. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): explain denied provider profile lookups Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations. Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
a316fd7832 |
feat(api): add durable workspace mutation admission and replay (#3321)
* feat(api): add durable workspace mutation admission and replay Part of #3051 (phase 3a). Preserve unresolved admissions, reauthorize replay, and protect same-name replacements. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * feat(api): extend mutation replay through gateway interceptors (#3323) Add typed durable replay receipts for 24 ordinary unary mutations, protect sensitive payload fingerprints, and revalidate intercepted retries without repeating post-commit observation. Part of #3051 (phase 3b). Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(api): scope workspace request IDs by target Include the requested workspace name in create/delete admission keys while leaving workspace UUID guards unset. Cover cross-target UUID reuse, replay, and missing targets with server and live gateway regressions. Merge the latest phase-two SDK fixes and preserve the approved interceptor replay changes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
c502be9fd7 |
feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(sdk): bind deletion waits to sandbox identity Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures. Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): accept asynchronous stopped sandbox deletion Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
d68b7069c3 |
refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time migration behavior Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve timestamp boundary semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): convert sandbox token expiry to timestamp Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time compatibility semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve exact endpoint and profile times Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): use duration for interactive exec timeout Signed-off-by: Derek Carr <decarr@redhat.com> * fix(sdk-go)!: remove legacy profile duration fields BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
2ccef97769 |
feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(policy-schema): add canonical authored policy model Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): use canonical authored schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): project canonical policy documents Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe shared schema boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): close reviewed parser gaps Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): fail closed on unsupported fields Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): preserve partial process identities Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): harden authored policy inspection Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): rename policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * revert(policy): restore policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(e2e): serialize OIDC PKCE scenarios Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |