mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 23:51:07 +08:00
dev
218
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
36819f476d |
fix(cli): stop uploads when Git filtering fails or selects no files (#3957)
Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
71440b28f4 |
fix(policy): refresh pending proposals when the sandbox policy changes (#3923)
* fix(policy): refresh pending proposals when the sandbox policy changes Approving, removing, or undoing a rule, or updating the sandbox policy, changes the inputs every other pending proposal was evaluated against. Only proposals the new policy covered were reconciled; the rest kept their old prover result and review token. The review surface (GetDraftPolicy) therefore showed a stale evaluation, and the first approval of the next proposal refreshed it and failed with FAILED_PRECONDITION, so approving proposals one after another always failed once. Re-evaluate the remaining pending proposals at each policy change, reusing the cached prover result unless the proposal's inputs changed. Approval still rejects a review token that does not match the stored evaluation, so a reviewer holding a pre-refresh evaluation must still refetch it. When a refresh does happen at approval time (inputs changed between fetch and approve), the CLI now explains that the rule was re-evaluated and how to review it, instead of printing the raw gRPC status. Closes #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> * fix(policy): make pending proposal refresh race-safe and bounded Store refreshed evaluations with a compare-and-swap: the store re-reads the proposal, refuses when its rule name, proposed rule, or review token changed since the evaluation read it, copies only the evaluation fields onto the stored record, and updates only if the payload is still the one it read. A refresh can no longer revert a concurrent edit or observation, and the edit path uses the same guard against a concurrent refresh. Bound each refresh to the 32 newest pending proposals; the rest keep the approval-time recheck, which still refuses a stale review token. Operator decisions (approve, approve-all, remove, undo) refresh before responding. UpdateConfig, which holds the gateway-wide sandbox sync guard, and agent-driven auto-approval refresh in a background task instead, one per sandbox with later changes coalesced into a single rerun. Refs #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> * docs(policy): describe proposal rechecks after approvals and approve-all Explain that approving, removing, or undoing a rule rechecks the other pending proposals so they can be approved one after another, when the recheck is deferred or bounded, and what rule approve reports when a proposal changed after it was listed. Show rule approve-all in Run Your First Agent with its security-flag behavior. Refs #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> * fix(policy): refresh pending proposals after a full policy replacement A full policy UpdateConfig (openshell policy set) re-reads the latest revision after its atomic write, finds the revision it just committed, and returns before reaching the pending-proposal refresh at the end of the handler. Pending proposals kept their stale evaluation, so rule get showed the old candidate and the next approval failed with the refresh precondition. Schedule the background refresh right after the commit. Refs #3884 Signed-off-by: fede-kamel <fkamelhar@gmail.com> --------- Signed-off-by: fede-kamel <fkamelhar@gmail.com> |
||
|
|
ffcbe6280c |
fix(cli): start sandbox exec without waiting for piped stdin EOF (#4006)
With a non-terminal stdin, sandbox exec read stdin to EOF before it sent the exec request. A pipe that never closes (CI runners, supervisors, agent harnesses) blocked the CLI forever in read(2) without the gateway ever seeing the request, and a slow producer delayed the command until EOF. Collect piped stdin on a detached reader thread for at most 200 ms. Input that reaches EOF within that window still travels in the single request that older gateways need. If the pipe is still open, start the command through the streaming RPC and forward the collected prefix plus the rest of stdin as it arrives, closing remote stdin at EOF. The 4 MiB cap covers the prefix and the streamed remainder together. Closes #3993 Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com> |
||
|
|
fe38637533 |
fix(runtime): recover SSH relays and bound startup diagnostics (#4011)
* fix(runtime): recover SSH relays and bound startup diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): deliver pending relays once per supervisor session Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): satisfy relay delivery clippy diagnostics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): bound relay setup with one absolute deadline Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
912a077bd6 |
feat(service): add bearer authorization passthrough (#3796)
* feat(service): add bearer authorization passthrough Signed-off-by: Derek Carr <decarr@redhat.com> * docs(sdk): add service authorization migration guide Signed-off-by: Derek Carr <decarr@redhat.com> * fix(server): remove stale version import Signed-off-by: Derek Carr <decarr@redhat.com> * docs(upgrade): remove service authorization SDK guide Signed-off-by: Derek Carr <decarr@redhat.com> * fix(e2e): relabel provider readiness TLS mount Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): stabilize exposed service routing Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): support HTTPS service routing Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
07a486d751 |
fix(cli): accept sandbox name before -- in exec (#3901)
* fix(cli): accept sandbox name before -- in exec Closes #3882 Signed-off-by: Eric Curtin <eric.curtin@docker.com> * fix(cli): define exec grammar in clap Signed-off-by: Eric Curtin <eric.curtin@docker.com> * docs(sandboxes): remove exec overview change from PR Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Eric Curtin <eric.curtin@docker.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ba16b9f2c7 |
fix(mcp): explain revision-scoped policy and rejections (#3850)
* fix(mcp): explain revision-scoped policy and rejections Explain the selected-revision method set in profile output and policy docs. Distinguish protocol and policy rejection causes and give a next step while preserving authorization, response statuses, error codes and YAML keys. Cover CLI serialization, revision selection, exact extension rules, deny precedence and rejection before forwarding with focused regressions. Signed-off-by: Shiju <shiju@nvidia.com> * docs(mcp): correct HTTP cancellation revision support Limit notifications/cancelled to the three 2025 revisions in the core method matrix. State that MCP 2026-07-28 HTTP cancellation closes the response stream, matching the runtime rejection and sessionless docs. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
12ef86c285 |
fix(cli): keep policy and provider diagnostics readable (#3444)
Reuse character-safe truncation for policy history errors so a multibyte character cannot panic the table renderer. Distinguish unavailable provider-profile YAML from an absent profile, display a bounded diagnostic, and preserve strict serialization and redacted object navigation. Cover the actual CLI renderer and TUI display/navigation paths, including invalid and absent profiles, Unicode input, and redacted errors. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
acbac9cb79 |
feat(sandbox): add main restart policy (#2798)
* feat(sandbox): add main restart policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden policy-driven restarts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): address restart review feedback Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(sandbox): port restart policy to current runtime Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(sandbox): restart promptly after terminal delivery Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
e63cfa1182 |
fix(cli): keep SSH forwards owned by spawned process (#3759)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
36b0386c92 |
feat(cli): detach sandbox sessions with Ctrl-D (#3744)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c63f8ce564 |
test(cli): migrate gateway-free smoke coverage (#3641)
* test(cli): migrate gateway-free smoke coverage Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(e2e): remove migrated tests from podman CI Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
9244868056 |
docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe updated security architecture neutrally Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: highlight new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandboxes): clarify how to disconnect Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align architecture and guides with current navigation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(extensibility): streamline extension authentication guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c93fd94a5f |
feat(cli): import provider profiles from HTTP URLs (#3706)
* feat(cli): import provider profiles from HTTP URLs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): initialize crypto provider for remote profiles Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
7a50c0899f |
fix(cli): stream piped exec stdin beyond gRPC request limit (#3687)
* fix(cli): stream piped exec stdin across gRPC messages Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): preserve small exec requests and surface stdin errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): preserve exec stdin limit across gRPC streaming Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
0518bd4c83 |
fix(sandbox): preserve local sessions across host sleep (#3573)
* fix(cli): recover sandbox connect transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): propagate non-expiring local sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): distinguish main exit from transport loss Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): bound sandbox connect recovery Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
123d95e2ed |
fix(pagination): document list contract and harden SDK pagers (#3279)
* fix(pagination): document list contract and harden SDK pagers Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): bound pager token history Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): preflight token history limits Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): paginate sandbox providers Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): harden TypeScript pager Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): expose provider pagers in SDKs Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * docs(pagination): describe a uniform list contract Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(pagination): use stable provider cursors Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * docs(pagination): clarify mutation semantics Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(cli): expose sandbox provider pagination Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(proto): refresh pagination schema inventory Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(proto): refresh rebased schema inventory Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
bdffa102c3 |
feat(api): add durable exec launch admission (#3324)
* feat(api): add durable exec launch admission Fence duplicate exec launches with keyed durable admission and producer-owned terminal completion. Keep uncertain launches unresolved and never replay output or interactive input. Part of #3051 (phase 4a). Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(api): fence exec identity across authorization lookups Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
1e34e8c576 |
fix(drivers): require admission labels for external resources (#3538)
* fix(drivers): require admission labels for external resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): address resource admission review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): reserve driver-owned admission labels Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): clarify workspace admission label Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(drivers): clarify resource admission failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): configure resource admission fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): retry forbidden admission lookups Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): preserve external driver admission defaults Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
470a34635d |
fix(api): make WatchSandbox loss-aware and resumable (#3209)
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag. Partially addresses #3055 (cursor/resume follow up separately). Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * refactor(server): group per-sandbox log bus state and stamp sequence numbers Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(proto): add resume cursor fields to sandbox watch API Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): stamp watch cursors from a shared per-sandbox sequence Allocate cursors from a single SeqAllocator shared by the log and platform event buses, so a sandbox's merged watch stream carries unique, strictly increasing cursors. A single resume_after_cursor can then unambiguously locate a client's position across both sources. Rewrite both publish paths to allocate the sequence, stamp event.cursor, send, and append to the tail under one lock. This removes the previous get_mut().expect() TOCTOU race where a concurrent remove() between the two lock sections could panic. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): serve WatchSandbox resume from cursor with gap detection Add tail_after() to the log and platform event buses, returning every buffered event newer than a client's resume cursor. Each PerSandbox now tracks last_trimmed_seq (the highest seq it has evicted) so a resume is reported as an unrecoverable ResumeGap only when this bus dropped an event the client still needs. Judging gaps by evictions, not by the tail's oldest seq, is required under the shared cursor space: each bus's tail is non-contiguous in the global sequence because the other bus owns the missing seqs, so comparing against tail.front() would flag false gaps. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(server): resume WatchSandbox from cursor across log and platform buses Wire resume_after_cursor into the watch producer. On a non-zero cursor, replay events strictly after it from both the log and platform buses, merge by shared cursor, and emit in order before entering the live loop. A trimmed range on either bus is an unrecoverable gap and terminates the stream with OUT_OF_RANGE carrying the requested and earliest-available cursors, distinct from recoverable lag which warns and continues. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): cover WatchSandbox cursor resume paths Add handler-level tests for the resumable watch stream: replay strictly after the client cursor, merge log and platform events in shared-cursor order, suppress duplicates when resuming at the latest cursor, and terminate with OUT_OF_RANGE when the requested cursor has been trimmed. Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * docs(api): document WatchSandbox loss-awareness and resume Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): deliver watch events once and harden cursor teardown Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * feat(sdk): add loss-aware resumable watch_logs to Rust SDK client Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): keep watch cursors monotonic across teardown and restart Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): merge live watch sources by cursor before emission Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): bind watch cursors to a cursor space and merge tail sources Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): revalidate the watch cursor space after collecting replay Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): update the public RPC schema fingerprint for the string cursor Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(server): synchronize the watch live-order test with the end of initialization Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(sdk): use canonical sandbox name in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * test(sdk): guard canonical-name addressing in watch_logs Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(server): fix public rpc schema Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> * fix(api): reconcile watch resume rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): bound interactive relay cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Artem Lytvyn <alytvyn@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
718dba3430 |
fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
50230616d5 |
refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
0a770d9173 |
feat(kubernetes): support HA gateway rebalancing (#1868)
* feat(kubernetes): support HA gateway rebalancing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(server): cache peer connections, tokens, and owner lookups Every forwarded relay rebuilt its setup from scratch: an owner lookup, a blocking read of the peer token, a TLS connect to the owning replica, and a TokenReview plus Pod GET on the receiving side. Sandbox service routing does this per HTTP request, so the apiserver calls scaled with traffic. Cache all of it on ServerState: - peer channels pooled per endpoint, so relays multiplex over one connection instead of redialing - peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit cannot accept an expired token - owner records for 3s against a 45s ownership TTL, still freshness checked before use Entries are evicted when a relay fails. Also raise HTTP/2 max_concurrent_streams to 1024, since pooling funnels every relay between two replicas onto one connection and hyper's default of 200 sits below the 256 pending-relay budget. Signed-off-by: divesh <dgude@nvidia.com> * perf(server): pool upstream connections for sandbox services Each HTTP request to a sandbox service opened its own supervisor relay, paying a new TCP connection and HTTP/1 handshake every time. Worse, it counted against the 32 in-flight relay cap, so a service handling more than 32 concurrent requests failed outright. Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is safe because the pool only returns a connection hyper reports as ready, and HTTP/1 cannot start a request until the previous body has drained. Upgrades are never pooled since they take the connection over, and a failed send evicts that endpoint. Pruning is bounded per key, with the full sweep limited to once per 30s. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): address HA gateway review findings (#3449) - Let a gateway own supervisor sessions without a peer endpoint. Requiring one whenever the store is PostgreSQL broke every single-instance PostgreSQL deployment, because no sandbox supervisor could connect. A cross-replica request to an owner that advertises no endpoint now fails immediately naming the cause, instead of retrying until the wait timeout. - Close a supervisor session on heartbeat only when another replica owns it, or after renewals fail for the ownership TTL. A database error no longer drops every session heartbeating during an outage. - Clamp owner record ages at zero so a skewed or corrupt stored timestamp cannot produce a negative age. - Bound the cross-object advisory lock with a lock timeout, so a stuck holder fails instead of blocking every mutation in the fleet. - Refuse to start when a peer endpoint is configured on a multi-replica backend but peer authentication is unavailable, and warn when a multi-replica backend has no peer endpoint at all. - Reject a plaintext peer endpoint when the gateway serves TLS. - Skip the sandbox watch poller on single-replica backends, where the local update bus already sees every write. - Rate-limit the peer owner cache sweep so an insert no longer scans the whole map under the lock. - Retry GET and HEAD on a pooled upstream the sandbox closed, instead of returning 502, and drop an emptied endpoint from the pool right away. - Document the gateway peer environment variables and the post-rollout ownership skew operators should expect. Signed-off-by: divesh <dgude@nvidia.com> * fix(server): harden HA supervisor ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: divesh <dgude@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: divesh <dgude@nvidia.com> Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com> |
||
|
|
2493d415c2 |
feat(extensions)!: normalize protocol negotiation (#3352)
* feat(extensions)!: normalize protocol negotiation Closes #3057 Introduce a shared extension handshake, enforce protocol and capability compatibility across extension families, and expose immutable negotiated snapshots through gateway info and the Go SDK. Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): fail fast on negotiation errors Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(extensions): validate gateway handshake metadata Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(go-sdk): re-export extension kind constants Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(extensions): fail fast on credential handshake rejection Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Seth Jennings <sjenning@redhat.com> |
||
|
|
fa8f6d3949 |
feat(cli): promote profile commands to top level (#3258)
* feat(cli): promote profile commands to top level Add profile discovery and management commands with shared handlers for the existing provider entry points. List a flat catalog across scopes and follow continuation tokens through full and short pages. Describe metadata, credentials, endpoints, TLS inspection, and MCP access settings while preserving complete JSON/YAML definitions. Cover parser equivalence, scope forwarding, pagination, and inspection settings with focused unit and compiled-CLI integration tests. Update docs, public skills, examples, and E2E command invocations. Refs #2588 Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): remove redundant workspace selector qualification Use the imported WorkspaceSelector in the provider integration helper so the target passes Clippy with warnings denied. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
1905069948 |
feat(sandbox): expose services during creation (#3439)
* feat(sandbox): expose services during creation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): refresh Codex credentials in gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): normalize create-time service URLs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): roll back failed service exposure Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(example): simplify Codex provider setup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(example): separate provider setup commands Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden create-time service exposure Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(example): bundle Codex provider profile Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(example): allow npm-installed Codex binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d91b1999a0 |
feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(cli): update forward color fixture for workspace scope Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): use sandbox names for settings lookup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox request fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): harden sandbox mutation handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): update rebased sandbox references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(api)!: standardize canonical entity references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(api): codify protobuf API conventions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve workspace selector semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): restore workspace selector parity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve descriptive name fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): update e2e request fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): omit workspace selector during bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): remove proto convention checker Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): refresh schema fingerprints after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): use canonical provider receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
903d9a0e7a |
chore(license): align repository compliance text (#3467)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
e38d7254e6 |
fix(policy): reject unknown endpoint security modes (#3187)
* fix(policy): reject unknown endpoint security modes Closes #3046 Validate TLS, enforcement, and access values across policy and provider profile ingress, and prevent runtime parsing from falling back to audit for unknown enforcement values. Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> * fix(policy)!: use enums for endpoint security modes Replace the public TLS, enforcement, and access strings with protobuf enums and carry the typed values through policy composition, provider profiles, drivers, and runtime conversion. Preserve the documented YAML spellings, reject unknown and invalid numeric enum values consistently, and update generated Go bindings, SDK conversions, tests, and policy documentation. Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> --------- Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com> |
||
|
|
1d010f4187 |
feat(sandbox): validate configuration before workload activation (#3259)
* feat(sandbox): validate configuration before workload activation Closes #3145 Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): bound startup failures and preserve activation history Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): enforce deadlines on startup RPC attempts Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): isolate provider auto-create policy fixture Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): remove unnecessary fixture string delimiters Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): capture startup logs before ephemeral cleanup Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): preserve credential revocation after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): retain provider revision helper for endpoint reports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(server): import provider object trait in production Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): preserve admission across boundary extraction Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(sandbox): use existing boundary discovery imports Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(sandbox): distinguish workload and supervisor containers Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission schema with timestamp migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with typed deletion schema Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): reconcile admission with mutation request IDs Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(supervisor): adapt local startup fixture to admission state Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): deduplicate startup quarantine diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): show configuration blockers in sandbox notes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat(sandbox): expire provisioning repair attempts after five minutes Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(supervisor): reconcile admission with provider readiness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): separate configuration summaries from full diagnostics Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(tui): shorten invalid configuration note Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(server): refresh schema inventory after rebase Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
04146692d9 |
refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules Twelve modules under crates/openshell-providers/src/providers/ were never declared in providers/mod.rs, so they have not been compiled since the plugin registry was narrowed to the two adapters it still registers. Four of them (generic, gitlab, opencode, outlook) key off provider type identifiers that normalize_provider_type already retired. Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new registers. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(providers): load example profiles from providers/ at test time Add an example_profiles module that reads the YAML under providers/ from the source checkout at run time and parses it with the existing profile loader. It is gated behind a new non-default example-profiles feature so it is available to this crate's own tests and, once wired into dev-dependencies, to gateway and CLI tests, while never reaching a release binary. Nothing consumes it yet; later changes move the test fixtures off the compiled catalog and onto these files. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test: source profile fixtures from providers/ instead of the compiled catalog Unit and integration fixtures reached into builtin_profiles() to get a profile to work with, which ties the tests to the compiled catalog rather than to the files an operator would import. Point them at the example_profiles loader instead. The profiles.rs tests keep their coverage unchanged and become explicit golden tests over providers/*.yaml: they are what keeps those files valid once nothing compiles them. builtin_profiles_are_sorted_by_id widens into a test that the whole example set parses, sorts, and lints clean as one catalog. The CLI fake gateways serve the example profiles the way a real gateway serves what an operator imported, through a shared helper. No behavior change: the gateway still loads the built-in source by default and serves the same profiles. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document the example provider profiles Each file under providers/ now opens with a header naming its expected client binary identities, the image layout those paths assume, the credential scope, the endpoint access it grants, and a smoke test. Add a README covering the import commands and why a profile should be copied and edited rather than imported unchanged. Several of these profiles bind network access to paths that only exist in the OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server, /usr/lib/node_modules. Imported unchanged into another image the profile matches nothing: the catalog still advertises it, but the credential is never injected and the traffic is denied. The headers say so where it applies. Comments only; the profile schema has no documentation fields and all fifteen files still lint clean. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(server): seed unit-test state with the example provider profiles Server unit tests inherited the builtin + user source default from ServerState, so around a hundred and forty assertions about github, openai and the rest resolved against the compiled catalog. Point the test state at the user source alone and import the example profiles from providers/ into its store first, the way an operator would. Every one of those assertions keeps passing unchanged, which is the point: it proves the gateway behaves identically with an imported catalog before the default moves. Four tests asserted builtin-source semantics specifically. Under import-only the only profiles a gateway cannot edit are the ones a non-user source vends, so they now exercise a source-managed profile composed with the user source; the read-only guard they cover is the one that still applies to interceptor catalogs. The list test asserts the imported profile's user/platform identity instead of builtin with an empty scope. test_server_state_with_user_only_github_profile is gone: the default test state now is a user-only gateway with github imported. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli): resolve credential suggestions from the gateway catalog The --env credential warning scanned a profile table compiled into the CLI, so its suggestions described the binary rather than the gateway the user is talking to: a profile the gateway does not serve was suggested anyway, and an imported custom profile never was. Fetch the catalog from the connected gateway instead. The warning moves out of argument parsing and into sandbox create and sandbox template create, where a client already exists. A catalog fetch failure is not fatal — the warning degrades to its generic form rather than blocking sandbox creation. Extract the ListProviderProfiles paging loop from provider list-profiles into a shared fetch_provider_profile_catalog, which the profile-driven paths now share. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: infer providers from the gateway catalog Command-to-provider inference went through a hardcoded alias table in openshell-providers, a second copy of the built-in catalog's identifiers. A custom profile could never be inferred no matter what binaries it declared, and the table drifted from the profiles it mirrored. Infer from the connected gateway's catalog instead: match the command's basename against each profile's ID and against the basenames of the binaries the profile authorizes. A profile that names /usr/bin/claude is the profile for running claude. The match must be unique — where several profiles claim a command, the user names one with --provider — and an empty catalog infers nothing. No command is special-cased. `binaries` is the operator's authorization statement, so a profile that declares a binary claims the command that runs it, whatever that binary is; narrowing that belongs in the profile rather than in a list compiled into the CLI, which could never cover an unbounded catalog anyway. Breaking: the retired aliases stop resolving, and commands the old table never listed can now infer. git, pip and uv are declared by the github and pypi example profiles, so they infer where they previously did not. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers,server)!: resolve provider profiles by exact ID normalize_provider_type was a hardcoded alias map — gh to github, claude to claude-code, vertex to google-vertex-ai — and a second copy of the built-in catalog's identifiers, independent of the YAML it mirrored. It let a provider type resolve to a profile the operator never named, and it made the built-in IDs behave as a reserved namespace. Remove it, along with detect_provider_from_command and the alias-normalizing ProviderRegistry::inject_env. A provider type now names a profile exactly: - the effective catalog resolves an ID or reports it absent, with no alias retry - plugins activate only for the ID of a profile the gateway resolved; a provider with no resolvable profile gets no plugin projection, instead of falling back to an alias guess - the CLI surfaces the gateway's not-found instead of retrying under an alias - telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets go with the aliases that were their only source Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve. Use the profile's own ID. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(server): fail closed when a provider's profile is absent A sandbox composed from a provider whose profile the gateway cannot resolve started anyway: the credential and policy builders warned and skipped, so the sandbox came up carrying none of that provider's credentials or network policy. The operator learned about it later, as a denied connection or a missing environment variable, rather than as the configuration error it is. Check at the two composition boundaries — CreateSandbox and AttachSandboxProvider — right beside the catalog snapshot already taken there, and reject with a bounded diagnostic naming the provider, the profile it refers to, and the import command that supplies it. Scope-aware: a platform-scoped provider is told to import with --global. Read paths are untouched. ListProviders, GetProvider and profile export keep working so an operator can see and recover an affected provider, and the shared policy builders keep their warn-and-skip for the diagnostic paths that also reach them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(server): scope the vendor base-URL pin to declared endpoints provider_profile_endpoints_are_active withheld a credential and the provider's policy layer when an openai or anthropic provider pointed its client somewhere other than the public vendor endpoint. The guard was keyed on profile.source == "builtin" and on those two profile IDs, so it protected only profiles OpenShell shipped. Once profiles are import-only no profile is ever builtin, and the guard would silently stop applying — including to an operator who imported providers/openai.yaml verbatim. Key it on the profile instead of on where the profile came from. A profile's endpoints are the boundary its credential is bound to, so if the provider configures a *_BASE_URL pointing at a host the profile does not declare, the profile no longer describes where that credential goes and is treated as endpointless. Host matching reuses the DNS-label-aware matcher in openshell-core, so wildcard endpoints such as Vertex's *-aiplatform.googleapis.com resolve correctly. A profile with no declared endpoints has no boundary to contradict, and a config value that names no host is not a redirect this can reason about. Both keep the profile active. The control now covers every endpoint-bearing profile, including an operator's own, and the last "builtin" string leaves the gateway. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): import example provider profiles during gateway bring-up The e2e suites create providers from github, openai, nvidia, claude-code and google-cloud, and the GitHub lane runs a real git clone through the profile's binary attribution. A gateway serves only the profiles an operator imported, so the lanes have to import them. Add e2e_import_example_provider_profiles to the shared bring-up helpers and call it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is healthy and registered. It runs the same command the upgrade notes give operators, against the repository's own providers/ directory, so the lanes exercise the documented path rather than a test-only shortcut. Lands before the default changes: a same-ID user profile already shadows a built-in, so importing works today. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers)!: make provider profiles import-only OpenShell compiled fifteen provider profile YAML files into every release binary and selected them by default, so a fresh gateway published a catalog it was never configured with. Those profiles are not image-neutral: their binary selectors name paths that exist in the OpenShell Community image, so changing the sandbox image could make a profile inert while the catalog still advertised it. Every endpoint, credential name and binary path in providers/ was also effectively part of the 0.1.0 public contract. A gateway's catalog is now exactly what an operator imported: - providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone from the openshell-providers API - the default provider_profile_sources is [{ type = "user" }], and a gateway with nothing imported reaches ready and serves an empty catalog — an empty catalog is a valid state, not a startup failure - the builtin source type is removed from the configuration schema, and a gateway.toml that still names it is rejected at parse time with the import command rather than an unknown-variant error - nothing reserves the canonical identifiers any more, so github, pypi, anthropic and the rest import at their own IDs and the imported profile is the only definition for that ID; the static_fallback that kept a shipped definition resident behind an imported one is gone with the source that produced it A collision between an interceptor-vended profile and an imported one still fails closed, which is the behavior that shadowing quietly bypassed for built-ins. Operators upgrading should export the profiles their deployment relies on before upgrading, or copy them from the providers/ directory of the matching release tag, then import them at the scope their providers use. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document import-only provider profiles Rewrite the provider profile documentation around a catalog the operator builds rather than one the gateway ships. - gateway-config: the default is [{ type = "user" }], the builtin source type is gone and rejected at startup, and an empty catalog is a valid ready state - profiles: replace the built-in profile table and its shadowing semantics with the import workflow, and say plainly that a profile whose binary paths do not match the image is inert - manage-providers: replace the two fixed provider-type tables with `provider list-profiles`, and describe catalog-driven command inference - inference-routing: drop the "still loads built-in profiles for compatibility" transition text; its migration walkthrough is now the normal path - quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex tutorials: import the profile before creating the provider, since nothing resolves without it - release notes: a 0.1.0 migration section covering export-before-upgrade, importing from the release tag's providers/ directory, what happens to a provider whose profile is missing, and the removal of the legacy type aliases - README, architecture, the openshell-cli and debug-openshell-cluster skills, and the governance interceptor example follow the same change; the example's smoke assertion now imports a profile to prove an authoritative interceptor hides it Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding Addresses two findings from review (GATOR-c1bd9867-02 and -03). The provider profile catalog lookup collapsed its error into an empty catalog, so an authorization, availability or transport failure on ListProviderProfiles was indistinguishable from a gateway that genuinely has no profiles. Command inference then resolved nothing and the sandbox was created without the provider it needed, deferring the failure to the workload. Keep the lookup's outcome instead of discarding it. The credential warning is advisory and still degrades to its generic form, but inference now consults the catalog only when there is a command to resolve and surfaces the lookup failure when there is, naming the gateway and pointing at explicit --provider selection. Two regression tests cover it through a fake gateway whose ListProviderProfiles returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails without sending a provider-less create request, and a sandbox with no trailing command still succeeds because it needs no catalog. The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes deliberately skip gateway registration and start the gateway without a TLS client CA, so no mTLS identity exists and no token has been acquired when the import runs; the wrapper exited during setup. Skipping the import is not sufficient because the provider tests now require the claude-code profile. Add e2e_register_oidc_admin_session, which mints an administrator token with Keycloak's password grant — the same grant the OIDC test helpers use — and writes the gateway metadata and token bundle that an interactive login would have stored, so the import runs as an authenticated administrator. The mTLS lanes keep the direct import unchanged. Both affected wrappers are covered: with-podman-gateway.sh had the same defect as with-docker-gateway.sh. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: remove command-derived provider attachment Addresses GATOR-c1bd9867-01. A profile's `binaries` list authorizes a binary to reach that profile's endpoints. It is not a statement that running the binary asks for the provider, and reading it as attachment intent let a command silently gain provider authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash` resolved an existing aws-s3 provider and attached its credential-backed capability and network policy to the shell and its descendants. Attachment of an already-created provider never prompted, so the confirmation flow did not guard it. The distinction that would make inference sound — whether a declared binary is a profile's client or merely a permitted runtime — cannot be expressed: NetworkBinary carries only a path, and the field that encoded it was removed in 0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant does not help, because uniqueness measures how sparse the catalog is rather than what the user intended, and a compiled list of "generic" commands could never cover an unbounded operator catalog. Remove trailing-command inference rather than approximate it. A provider is attached only when named with --provider, which still creates a missing provider from local discovery when the name matches an imported profile ID. Sandboxes with no providers remain a normal, fully supported state. Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving authority from the catalog, its only remaining use is the advisory credential warning, so a failed lookup degrades that warning instead of blocking creation. The regression test now asserts that an unreachable catalog still creates the sandbox and attaches nothing. While repurposing the deduplication test, a pre-existing defect surfaced: repeating a name in --provider auto-created it twice and the second attempt failed with "provider already exists", because the explicit pass never consulted the set of names it had already handled. Guard it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): load the OIDC token and trust the gateway when seeding profiles Addresses the carried finding GATOR-c1bd9867-03. e2e_register_oidc_admin_session established a session the CLI never used. The CLI decides whether to load a stored bearer token by matching on the auth_mode field of the gateway metadata alone; the helper omitted that field, so the metadata fell through to the default arm and oidc_token.json stayed on disk unread. The profile import went out unauthenticated exactly as it had before the helper existed. The helper also installed no trust anchor, leaving the CLI unable to verify the gateway's self-signed serving certificate. Write auth_mode = "oidc" so the stored token is loaded, and install the CA at <gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or key on disk the CLI falls back to CA-only server verification and authenticates with the bearer token, which is what these lanes need because they start the gateway without --tls-client-ca. Certificate verification stays on; no insecure transport override is introduced. Assert the session before anything depends on it. ListProviderProfiles is annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded and accepted. The helper now fails at that point, naming the gateway config directory and echoing the CLI output, rather than letting the defect surface later as an opaque profile import error. Both affected wrappers pass the PKI directory and CLI binary the helper needs; the mTLS lanes are untouched. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): request the OpenShell scopes when minting the admin token The OIDC lanes failed at the session assertion added in b25dd0298 with PERMISSION_DENIED and "scope 'provider:read' required". Authentication was working -- the gateway logged ListProviderProfiles at gRPC status 7, which it can only reach once the bearer token has been loaded and accepted. The token simply carried no OpenShell scope. sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional client scopes on the openshell-cli client in scripts/keycloak-realm.json, so Keycloak mints them only when the request asks for them. The password grant here asked for nothing, leaving the realm defaults (openid, profile, email, roles, web-origins, acr) and an access token that authorizes no RPC. Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already does for its administrator tokens. openshell:all is SCOPE_ALL in crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls without enumerating them. Verified against the realm: the token goes from "email profile openid" to "email profile openshell:all openid". Record the same scopes in metadata.json. The helper writes the bundle openshell gateway login would have stored, and oidc_scopes is the field that login path reads back, so leaving it out would misdescribe the stored token. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): seed provider profiles once per VM lane gateway The VM lanes failed on the second test target with "custom provider profile 'anthropic' already exists" for all fifteen example profiles, and the import exited non-zero. e2e_import_example_provider_profiles was called from run_e2e_test, so it ran once per target -- four times against one long-lived gateway. That was harmless while import overwrote silently, but profiles are now import-only: ImportProviderProfiles is create-only and reports an existing id as an error-severity diagnostic, with no overwrite flag on the request. The first import therefore succeeds and every later one fails. Hoist the call to just after the conformance run, which is where the docker, podman and kube lanes already seed their catalogs. The profiles persist for the gateway's lifetime, so every target still finds them, and the E2E_TEST_OVERRIDE path is covered by the same single call. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): name the provider explicitly in the OIDC workspace-user test user_can_create_sandbox_with_inferred_provider_command reached the gateway once the admin session was fixed, and then failed: it asserted a missing-provider error, but the sandbox was created and died provisioning with "failed to spawn sandbox entrypoint process 'claude-code'". The test drove provider resolution by passing claude-code as the trailing command and relying on the CLI to infer the provider type from it. That inference is what this branch removed, so the trailing word is now nothing but an entrypoint, and the image has no such binary. The regression the test guards is not inference itself: it is that a workspace user resolving a provider is not gated behind Platform Admin. Name the provider with --provider, the only remaining way to attach one. The CLI still has to fetch the claude-code profile before it can auto-create the provider, so the lookup a workspace user must be allowed to make still happens, and auto-creation still stops at the non-interactive branch with "missing required provider". Both assertions therefore keep their meaning. Rename the test and rework its comments to describe what it now exercises; the old name would otherwise outlive the behavior it was named for. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(providers): reconcile import-only profiles with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
a72351d370 |
fix(policy)!: require explicit L7 append targets and scope (#3380)
Make allow and deny appends identify a rule and endpoint and declare every affected binary and port. Reject incomplete, stale, ambiguous, and provider targets atomically so a small append cannot silently change a broader scope. Update CLI previews, wire requests, Go types and generated bindings, SDK regressions, operator documentation, and live policy-update coverage. BREAKING CHANGE: AddAllowRules and AddDenyRules require L7RuleTarget instead of host and port. CLI appends require a rule name and explicit binary scope. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
9708ba9999 |
feat(providers): report applied sandbox provider changes (#3391)
* feat(providers): report applied sandbox provider changes Record exact provider mutation targets in shared configuration operations. Require authenticated evidence that credentials, effective policy, and the workload launch environment have been installed before reporting readiness. Add bounded CLI and Rust SDK status and wait support, preserving ordinary revision-scoped references for existing processes. Verify new-client rotation and acknowledged detach revocation without external-stable resolver changes. Signed-off-by: Shiju <shiju@nvidia.com> * fix(providers): align readiness times with protobuf contracts Represent readiness receipts, status, and operation times with Timestamp and report intervals with Duration. Reserve the scalar field tags, update all consumers and generated bindings, and preserve timestamp presence and nanosecond identity through storage and client validation. Qualify both empty-map constructors in the Linux boundary test so its module compiles while retaining the explicit default required by Clippy. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): preserve provider mutation storage uncertainty Recognize the gateway's exact structured storage-uncertainty reason for provider attach, detach, and update. Explain that the change may already be saved and must be reconciled before retrying, without exposing server messages or metadata. Preserve uncertainty ahead of generic retry hints. Exercise saved mutations through the CLI and verify single submission, redaction, missing receipt handling, and untrusted error-detail rejection. Document the recovery guidance for users and the public CLI skill. Signed-off-by: Shiju <shiju@nvidia.com> * fix(cli): explain denied provider profile lookups Report exact and alias profile lookup denials with fixed permission and workspace guidance. Keep backend details redacted and stop before provider mutations. Cover denied create and update calls through the CLI. Verify the complete provider list independently in the cross-workspace OIDC regression, extracting its JSON object from surrounding startup diagnostics. Signed-off-by: Shiju <shiju@nvidia.com> --------- Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
a316fd7832 |
feat(api): add durable workspace mutation admission and replay (#3321)
* feat(api): add durable workspace mutation admission and replay Part of #3051 (phase 3a). Preserve unresolved admissions, reauthorize replay, and protect same-name replacements. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * feat(api): extend mutation replay through gateway interceptors (#3323) Add typed durable replay receipts for 24 ordinary unary mutations, protect sensitive payload fingerprints, and revalidate intercepted retries without repeating post-commit observation. Part of #3051 (phase 3b). Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(api): scope workspace request IDs by target Include the requested workspace name in create/delete admission keys while leaving workspace UUID guards unset. Cover cross-target UUID reuse, replay, and missing targets with server and live gateway regressions. Merge the latest phase-two SDK fixes and preserve the approved interceptor replay changes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
c502be9fd7 |
feat(api): return typed deletion outcomes with explicit missing-target semantics (#3317)
* feat(api): return typed deletion outcomes with explicit missing-target semantics Implement phase 2 of #3051 across the public gateway API and first-party SDKs. Preserve asynchronous sandbox deletion and observed resource identities, reserve legacy wire fields, and document the coordinated migration. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(sdk): bind deletion waits to sandbox identity Track accepted deletions by original sandbox identity in Rust and TypeScript. Preserve name-only waits and cover replacement races and lookup failures. Merge the phase-one cleanup fix and adapt its interceptor regressions to typed deletion outcomes. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): accept asynchronous stopped sandbox deletion Allow either completed or accepted deletion output, then continue polling for actual sandbox absence. This preserves the lifecycle assertion under the phase-two deletion contract. Refs #3051. Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
d68b7069c3 |
refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time migration behavior Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve timestamp boundary semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): convert sandbox token expiry to timestamp Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve time compatibility semantics Signed-off-by: Derek Carr <decarr@redhat.com> * fix(proto): preserve exact endpoint and profile times Signed-off-by: Derek Carr <decarr@redhat.com> * test(e2e): use duration for interactive exec timeout Signed-off-by: Derek Carr <decarr@redhat.com> * fix(sdk-go)!: remove legacy profile duration fields BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields. Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
2ccef97769 |
feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(policy-schema): add canonical authored policy model Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): use canonical authored schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): project canonical policy documents Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe shared schema boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): close reviewed parser gaps Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): fail closed on unsupported fields Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): preserve partial process identities Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): harden authored policy inspection Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): rename policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * revert(policy): restore policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(e2e): serialize OIDC PKCE scenarios Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
dfd5238d0d |
fix(gator): make supervised lifecycle sandbox-native (#3343)
* fix(gator): run supervisor as sandbox main process Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * feat(gator): persist supervised state history Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * refactor(gator): remove obsolete background launch mode Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
fd3fd9cf74 |
feat(sandbox): explain failed calls to external tool servers (#3207)
Show configured tool server addresses and their last observed connection results together in sandbox status. Keep sandbox lifecycle readiness separate so an external connection failure does not mark the sandbox unready. Expose direct endpoint records through the CLI and SDKs, with plain-language failure explanations and gateway acceptance times. Keep observation tracking, runtime reporting, and gateway validation in dedicated endpoint status modules. Preserve bounded reporting, request attribution, retry ordering, and configuration and supervisor authority checks. Clear obsolete observations while retaining the configured addresses, and document the distinction between an observed HTTP response, current availability, and tool success. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
0803c4aa4c |
refactor(policy)!: remove NetworkBinary harness field (#3222)
* refactor(policy)!: remove NetworkBinary harness field Closes #3054 Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * test(policy): preserve unknown fields during migration Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(policy): preserve advisor provenance during merge Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(policy): remove redundant network binary default Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): record harness schema migration Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
33bbda3d33 |
refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address continuation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(tui): recover completed list refreshes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): address review scalability findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(pagination): repair branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(go): use page size in template example Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
90dbe5454b |
feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(cli): preserve template workspace metadata Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(e2e): migrate workspace request selectors Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * test(api): update public schema inventory Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
3ea0ce896c |
fix(cli): reconcile provisional container exits (#3204)
Keep watching briefly when a compute driver reports ContainerExited without a canonical main-process result. This prevents --no-keep cleanup from deleting the sandbox before the supervisor publishes the authoritative exit status. Signed-off-by: Shiju <shiju@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
48c449d8c8 |
chore(deps): replace ring with AWS-LC (#3243)
* chore(deps): replace ring with AWS-LC Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(lint): address warnings after dependency upgrades Signed-off-by: Simon Scatton <sscatton@nvidia.com> * fix(tls): limit provider initialization to reqwest clients Signed-off-by: Simon Scatton <sscatton@nvidia.com> --------- Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
6e6b3c8905 |
refactor(cli): remove local Dockerfile image builds (#3214)
Signed-off-by: Evie Howard <evhoward@redhat.com> |
||
|
|
118b250f01 |
feat(sandbox): support rootfs tar as --from source for VM driver (#2863)
* feat(sandbox): support rootfs tar as --from source for VM driver Accept flat rootfs tar archives (.tar, .tar.gz, .tgz) via the --from flag for VM-backed gateways. The CLI detects the archive extension, validates that the gateway uses the VM compute driver, and passes the tar path through driver_config. The VM driver copies the tar into its staging area and feeds it into the existing rootfs extraction and ext4 disk creation pipeline, skipping the container image pull/export steps. Closes #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): validate rootfs tar path at the VM driver boundary The rootfs_tar_path field in driver_config was passed from the API caller directly to tokio::fs::copy without validation. An authenticated user bypassing the CLI could supply arbitrary host paths (e.g. /dev/zero for disk exhaustion, or readable host files for data exfiltration). Introduce a trusted staging directory that the VM driver creates on startup and advertises via GetCapabilities. The CLI now copies the tar into the staging directory before creating the sandbox, and the driver validates that the received path is a regular file inside the staging root and within a configurable size limit (default 10 GiB) before any I/O. New VmDriverConfig options: - rootfs_tar_staging_dir: override the staging directory (default: <state_dir>/rootfs-tar-staging) - rootfs_tar_max_bytes: override the size limit (default: 10 GiB) Addresses GATOR-28b5152e-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): request-scoped staging, size pre-check, and cleanup for rootfs tar Tighten the rootfs tar staging flow to address the remaining GATOR-01 obligations: - Request-scoped staging: the CLI creates a unique per-request subdirectory (req-<pid>) under the staging root instead of placing files directly in the shared directory. The driver enforces that the tar path is at depth 2 (staging_root/<subdir>/<file>), preventing cross-request path selection. - Size pre-check: the driver advertises rootfs_tar_max_bytes via GetCapabilities. The CLI reads this limit and rejects oversized files before copying, avoiding disk exhaustion in the staging directory. - Cleanup: the driver removes the request staging subdirectory after consuming the tar (on cache hit, copy success, or copy failure), ensuring staged data does not persist beyond the request. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): restore rootfs-tar sandboxes from persisted image identity On restore or restart, the one-shot staged tar archive has already been cleaned up. Reading the persisted image identity from the sandbox state directory and resolving the cached disk path directly avoids re-accessing the deleted staging path. Addresses GATOR-168b9210-01. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli): use random staging dirs and enforce byte limit during rootfs tar copy Replace PID-based request staging directories with tempfile-generated random names to prevent collisions and make paths unpredictable. Replace bare tokio::fs::copy with a streaming copy loop that enforces the advertised max_bytes limit during transfer, closing the TOCTOU gap between the pre-copy size check and the actual copy. Signed-off-by: Philippe Martin <phmartin@nvidia.com> Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): issue rootfs tar staging slots from the gateway A caller could name any host path in `driver_config.vm.rootfs_tar_path`, which the privileged VM driver then read. The CLI-side locality check did not apply to direct API requests. The gateway now owns staging. `BeginRootfsTarStaging` allocates a request-scoped directory under the driver-advertised staging root and returns an opaque single-use token; `CreateSandbox` carries the token, and the gateway substitutes the path it allocated before dispatching to the driver. `template.driver_config.<driver>.rootfs_tar_path` is rejected outright in request validation, so a caller-supplied path never reaches privileged I/O. Tokens are bound to the issuing workspace and subject, consumed once, and expire after 30 minutes. Outstanding slots are capped per caller and overall, so one caller can neither exhaust the staging filesystem nor starve others. An RAII guard reclaims the directory on every failure path after consumption, and an age-gated sweep runs at startup and on each reconcile pass for directories whose driver died before its own cleanup. The token is stripped from the public sandbox before persistence: the stored copy is returned verbatim by GetSandbox, ListSandboxes and WatchSandbox to every member of the workspace. Also fixes two defects this exposed: - The CLI wrote `rootfs_tar_path` at the top level of `driver_config`, but the gateway forwards only `driver_config.<driver_name>`, silently dropping unmatched keys. The archive never reached the VM driver, so the documented `--from ./rootfs.tar` flow did not work at all. Config is now nested under `vm` and deep-merged, so a caller's existing VM settings survive instead of being clobbered by a shallow extend. - Staging previously required `GetGatewayInfo`, which is restricted to `platform_admin`, making the feature unusable for ordinary users on any RBAC-enabled gateway. The new RPC matches CreateSandbox at `sandbox:write` / `workspace_role: user`. `compute_driver.proto` is unchanged; the gateway reads the staging root from the capabilities it already stores. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): derive rootfs tar cache identity from archive contents The prepared-disk cache key combined the archive's full path with an mtime truncated to seconds, then mapped punctuation to `-`. Distinct paths such as `/tmp/a/b.tar` and `/tmp/a-b.tar` collapsed onto the same key and reused each other's disk, a rewrite within the same second kept stale contents, and a long path could exceed filesystem component limits. Identity is now a SHA-256 of the archive contents. This is also what makes the cache work at all now that the gateway allocates a fresh staging directory per request: a path-derived key would miss on every create. The archive is hashed, the cache checked, and only on a miss copied — so a hit skips writing a multi-gigabyte file. The copy is hashed as it is written and rejected if the digest differs from the first pass, which closes the window where the source changes during staging rather than approximating it with a re-stat. Refs #2175 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(vm): decompress gzip rootfs tar archives during staging `--from` accepts `.tar.gz` and `.tgz`, but the driver staged whatever bytes it was given and the guest image-prep VM extracts the staged file with a plain `tar -xpf`. Compressed sources therefore depended on the guest tar auto-detecting gzip, and the prepared disk was sized from the compressed length, which is far too small for the expanded rootfs. Staging now detects gzip from the archive's magic bytes -- the driver only ever sees a gateway-issued staging path, never the caller's file name -- and writes an uncompressed tar. The digest still covers the source bytes, so the "archive changed while staging" check is unaffected, and expansion is bounded by `rootfs_tar_max_bytes` so a compression bomb cannot fill the host disk. `extract_rootfs_archive_to` sniffs gzip as well, so the host-side extraction path matches. Adds unit coverage for gzip staging, bounded expansion, and gzip extraction, plus an e2e sandbox created from a gzip-compressed export. Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: Philippe Martin <phmartin@nvidia.com> |
||
|
|
2ad86c1b2e |
fix(cli): fail closed when OIDC refresh fails (#2817)
* fix(cli): fail closed when OIDC refresh fails Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(cli): fall back to unexpired OIDC token on refresh failure Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> |
||
|
|
08eac8c46d |
fix(sandbox): detect an available login shell instead of hardcoding /bin/bash (#3147)
* fix(sandbox): detect an available login shell instead of hardcoding /bin/bash The built-in default sandbox command and the interactive SSH session hardcoded /bin/bash. Minimal images such as Alpine ship only /bin/sh (BusyBox ash), so sandbox startup failed with an opaque "No such file or directory (os error 2)" that never named the missing binary. Add openshell-core::shell with shell-path constants and a runtime detect_login_shell() that resolves a shell present in the sandbox image ($SHELL if executable, then bash, then /bin/sh). Use it for: - the built-in default command (only the default is remapped; explicit user commands are never rewritten), resolved in the supervisor so it inspects the sandbox filesystem rather than the gateway's - the SSH interactive shell - the SHELL environment variable Also name the program in the spawn error so a missing shell/binary is diagnosable instead of a bare ENOENT. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): drop $SHELL preference in shell detection $SHELL is image/user-controlled and the detected shell is later invoked with `-lc`, so an executable that is not a compatible shell (e.g. SHELL=/bin/false) would pass the executable check and then break command execution even when /bin/sh is available. Resolve only from known shell paths instead. Also add a USR_BASH constant for /usr/bin/bash rather than a string literal in SHELL_CANDIDATES. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> * fix(sandbox): resolve the default login shell in the supervisor (empty command = default) Addresses review: interactive PTY SSH now uses the detected shell, the shell tests are portable across the Windows lane, and default-shell provenance is carried without a new spec field. An omitted command is left empty end to end and resolved in the supervisor, which is the only place that sees the sandbox image: - The CLI forwards the command as-is; the gateway persists an omitted command as empty (no baked /bin/bash -l) and requests a TTY. - MainProcessConfig carries the command empty (the transport now allows it); the supervisor resolves a login shell that exists in the sandbox image (bash when present, otherwise /bin/sh on minimal images like Alpine) and logs the resolved shell. - Interactive PTY SSH (spawn_pty_shell) uses the detected shell; a shared build_ssh_shell_command helper covers the PTY and non-PTY paths, with a deterministic sh-only regression test. - Unix-only shell tests are gated with cfg(unix). An explicit command is always run verbatim. Refs #3146 Signed-off-by: Akram <akram.benaissi@gmail.com> --------- Signed-off-by: Akram <akram.benaissi@gmail.com> |
||
|
|
52b1d78e76 |
refactor(cli): extract provider commands into commands/provider module (#2605)
Move all provider-related functions, helpers, constants, and tests from the monolithic run.rs (~2,700 lines) into a dedicated commands/provider.rs module. This is PR3 of the CLI refactor series (issue #2304). The extraction follows the same pattern established in PR2 (gateway): - Self-contained module with own imports - pub use re-exports in run.rs so callers (main.rs) are unchanged - Inline #[cfg(test)] mod tests Signed-off-by: Varsha Prasad Narsing <vnarsing@nvidia.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |