Rewrite legacy NetworkEndpoint string fields before protobuf decode so DEB and RPM upgrades from v0.1.0-pre.4 can load persisted sandboxes.
Signed-off-by: Evan Lezar <elezar@nvidia.com>
CA certificates are PEM-encoded text. Using String instead of Vec<u8>
avoids base64 overhead in the JSON wire format, keeping large CA bundles
within the 1 MiB control frame limit without needing to raise it.
Signed-off-by: Florent Benoit <fbenoit@redhat.com>
* fix(sandbox): fall back to plain seccomp listener when WAIT_KILLABLE_RECV is unavailable
SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV was added in Linux 5.19. On older
kernels (for example RHEL 9.x / 5.14 nodes such as RHCOS on OpenShift) the
flag is rejected with EINVAL, which made the capability-free sandbox fail to
start during the notification probe with "notification launcher disappeared".
Install the notification listener with WAIT_KILLABLE_RECV when the kernel
supports it and fall back to a plain NEW_LISTENER on EINVAL. The fallback
listener records wait_killable_recv = false: its notification receive is
uninterruptible, but the sandbox is otherwise fully functional.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* fix(sandbox): address review — legacy read-only listener mode + cancellation invariant
Follow-up to the PR review (GATOR-de00bfcc-01 / mrunalp): make the < 5.19
fallback cancellation-safe instead of racing broker writes.
- Record an explicit ListenerMode (Killable vs LegacyReadOnly); add
writes_disabled()/mode() and emit the selected mode in qualification output
(seccomp_listener_mode).
- Centralize task-memory output writes behind NotificationListener::
write_task_output; in LegacyReadOnly mode getpeername, accept/accept4 with a
non-null address, and sendmmsg length write-backs fail closed with EOPNOTSUPP.
accept with a null address, socket/connect/bind/listen/sendto/sendmsg keep
working (copied inputs, scalar responses, atomic ADDFD_SEND).
- Enforce the launch invariant `cancellation || task_memory_writes_disabled` in
SandboxConfirmEvidence::validate() rather than dropping cancellation
unconditionally; add task_memory_writes_disabled to SeccompEvidence.
- Add per-path fail-closed tests (write_task_output, write_socket_addr, a real
plain listener installed on a modern kernel) and confirmation-invariant tests.
- Correct the flag-semantics comments and document both modes plus the reduced
legacy syscall compatibility in architecture/sandbox.md.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* docs: document legacy read-only sandbox mode on kernels before 5.19
Kernels without SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV (< 5.19, e.g. RHEL 9.x /
RHCOS 5.14) run the sandbox in a legacy read-only cancellation mode where the
broker fails closed with EOPNOTSUPP on the mediated operations that write results
back into workload memory (getpeername, accept/accept4 with a non-null address,
sendmmsg length write-backs). Document this observable behavior and its syscall
limitations in the public Fern docs: the support-matrix kernel requirements and
the OpenShift runtime guidance.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* style(sandbox): satisfy rustfmt and clippy doc_markdown
Match the pinned rustfmt (Rust 1.95.0) line-wrapping for the write_task_output
call, and backtick `legacy_read_only` in the qualification-report doc comment
so clippy::doc_markdown (-D warnings) passes.
Signed-off-by: Akram <akram.benaissi@gmail.com>
* fix(sandbox): migrate task_memory_writes_disabled into backend protocol
Add the `task_memory_writes_disabled` field to `SeccompEvidence` in the
backend protocol contract and relax the validation from requiring
`cancellation` to accepting `cancellation || task_memory_writes_disabled`.
This completes the rebase migration missed by the isolation-interface
refactor: the sandbox reports this field but the backend struct lacked
it, and the validator rejected every pre-5.19 legacy listener.
Signed-off-by: Akram <akram.benaissi@gmail.com>
---------
Signed-off-by: Akram <akram.benaissi@gmail.com>
* fix(api): emit warning on WatchSandbox broadcast lag instead of terminating
Broadcast lag on the status, log, and platform receivers was converted to a RESOURCE_EXHAUSTED status that terminated the whole watch stream. Lag is recoverable: the receiver resumes at the oldest surviving message. Emit a SandboxStreamWarning and continue streaming instead; keep terminating on Closed. Add helpers and unit tests covering the warning payload and receiver recovery after lag.
Partially addresses #3055 (cursor/resume follow up separately).
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* refactor(server): group per-sandbox log bus state and stamp sequence numbers
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(proto): add resume cursor fields to sandbox watch API
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(server): stamp watch cursors from a shared per-sandbox sequence
Allocate cursors from a single SeqAllocator shared by the log and
platform event buses, so a sandbox's merged watch stream carries
unique, strictly increasing cursors. A single resume_after_cursor can
then unambiguously locate a client's position across both sources.
Rewrite both publish paths to allocate the sequence, stamp
event.cursor, send, and append to the tail under one lock. This
removes the previous get_mut().expect() TOCTOU race where a concurrent
remove() between the two lock sections could panic.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(server): serve WatchSandbox resume from cursor with gap detection
Add tail_after() to the log and platform event buses, returning every
buffered event newer than a client's resume cursor. Each PerSandbox now
tracks last_trimmed_seq (the highest seq it has evicted) so a resume is
reported as an unrecoverable ResumeGap only when this bus dropped an
event the client still needs.
Judging gaps by evictions, not by the tail's oldest seq, is required
under the shared cursor space: each bus's tail is non-contiguous in the
global sequence because the other bus owns the missing seqs, so
comparing against tail.front() would flag false gaps.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(server): resume WatchSandbox from cursor across log and platform buses
Wire resume_after_cursor into the watch producer. On a non-zero cursor,
replay events strictly after it from both the log and platform buses,
merge by shared cursor, and emit in order before entering the live loop.
A trimmed range on either bus is an unrecoverable gap and terminates the
stream with OUT_OF_RANGE carrying the requested and earliest-available
cursors, distinct from recoverable lag which warns and continues.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(server): cover WatchSandbox cursor resume paths
Add handler-level tests for the resumable watch stream: replay strictly
after the client cursor, merge log and platform events in shared-cursor
order, suppress duplicates when resuming at the latest cursor, and
terminate with OUT_OF_RANGE when the requested cursor has been trimmed.
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* docs(api): document WatchSandbox loss-awareness and resume
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): deliver watch events once and harden cursor teardown
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* feat(sdk): add loss-aware resumable watch_logs to Rust SDK client
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): keep watch cursors monotonic across teardown and restart
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): merge live watch sources by cursor before emission
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(api): bind watch cursors to a cursor space and merge tail sources
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): revalidate the watch cursor space after collecting replay
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(server): update the public RPC schema fingerprint for the string cursor
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): hold watch events above the publication watermark and emit the watch lag warning before its batch
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(server): synchronize the watch live-order test with the end of initialization
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(sdk): use canonical sandbox name in watch_logs
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* test(sdk): guard canonical-name addressing in watch_logs
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(server): fix public rpc schema
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
* fix(api): reconcile watch resume rebase
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(server): bound interactive relay cleanup
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: Artem Lytvyn <alytvyn@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(sandbox): reclaim socket descriptors before exhaustion
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(sandbox): separate socket and descriptor limits
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(sandbox): account for existing broker descriptors
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* feat(kubernetes): support corporate proxy CA bundle
The Kubernetes driver had no way to supply a CA bundle for the corporate
egress proxy, so an `https://` proxy with a private CA, or a TLS-intercepting
proxy, could not be used. Podman and VM already expose `proxy_ca_bundle`.
Add `proxy_ca_bundle` to `[openshell.drivers.kubernetes]` as a path the
gateway Pod reads. The gateway stages the PEM into the existing
per-generation supervisor bootstrap Secret and passes
`--upstream-proxy-ca-bundle` on the supervisor argv. That Secret is already
immutable, owner-referenced and garbage-collected, and its volume mounts
every key at /.openshell/supervisor with no items filter, so this needs no
new object kind, volume, mount, or RBAC verb, and works in shared, managed
and operator workspace modes.
The bundle is deliberately read from the gateway's filesystem rather than
referenced as an object in the sandbox namespace. It becomes a trust anchor
for every upstream the sandbox reaches, so it must stay in the gateway's
trust domain; the immutable staging Secret also keeps the anchor from
changing underneath a running sandbox.
Bound the staged bundle at 256 KiB. The shared reader's limit is exactly the
apiserver's own Secret limit and the bootstrap Secret carries four other
keys, so a bundle between the two would pass gateway startup and then fail
every sandbox create with an opaque `data: Too long`.
Delegate the URL, no_proxy, connect_by_hostname and ca_bundle rules to the
shared validate_upstream_proxy_settings, keeping the Secret-specific
credential block local: this driver accepts an explicit
`proxy_auth_allow_insecure = false` without credentials, which the shared
rules reject. This also fixes the acknowledgement being demanded for an
`https://` proxy, where the credential travels inside the verified TLS
session. Add auth_setting_label so the inline-credential diagnostic names
the Secret keys instead of proxy_auth_file, which this driver rejects as an
unknown key.
Document that the bundle should carry only the CA that signs the proxy's
certificate, or that an intercepting proxy re-signs upstream certificates
with. Public roots already reach the sandbox through the supervisor image and
its TLS stack, and the bundle is concatenated with that system store into a
single boundary control frame, so a full merged trust bundle spends the frame
budget on duplicated roots. The frame, not the apiserver Secret limit, is the
tighter of the two ceilings in practice; raising the staging bound requires
checking it.
Closes#3443
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(helm): quote proxy CA ConfigMap references
Signed-off-by: Philippe Martin <phmartin@redhat.com>
---------
Signed-off-by: Philippe Martin <phmartin@redhat.com>
The RFC-0012 architecture migration preserved sandbox OCSF JSON enablement but
introduced a regression: ocsf_schema_version was no longer honored.
This fixes the regression so that schema downgrades occur again.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(sandbox): default to official Alpine sandbox image
default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(deploy): default deployment configs to the official Alpine sandbox image
Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(driver): default to numeric non-root identity for USER-less images
With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.
Part of #3116.
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* test(conformance): use Alpine workload image
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* refactor(policy): drop community image /app path from default policy
The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* docs(config): document Alpine default images
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): report early sandbox termination
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): initialize rootless workspace ownership
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(sandbox): qualify NVIDIA Ubuntu default
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(podman): initialize rootful default workspace
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(sftp): add native sandbox adapter
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): gate runtime helper support to Linux
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): support standard OpenSSH file operations
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): harden rename and special file handling
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(runtime): remove community image dependencies
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(e2e): build provider readiness tool fixture
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): use a dedicated Noble fixture for Docker tests
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* feat(kubernetes): support HA gateway rebalancing
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* perf(server): cache peer connections, tokens, and owner lookups
Every forwarded relay rebuilt its setup from scratch: an owner lookup, a
blocking read of the peer token, a TLS connect to the owning replica, and
a TokenReview plus Pod GET on the receiving side. Sandbox service routing
does this per HTTP request, so the apiserver calls scaled with traffic.
Cache all of it on ServerState:
- peer channels pooled per endpoint, so relays multiplex over one
connection instead of redialing
- peer tokens keyed by SHA-256, expiring at min(ttl, token exp) so a hit
cannot accept an expired token
- owner records for 3s against a 45s ownership TTL, still freshness
checked before use
Entries are evicted when a relay fails. Also raise HTTP/2
max_concurrent_streams to 1024, since pooling funnels every relay between
two replicas onto one connection and hyper's default of 200 sits below
the 256 pending-relay budget.
Signed-off-by: divesh <dgude@nvidia.com>
* perf(server): pool upstream connections for sandbox services
Each HTTP request to a sandbox service opened its own supervisor relay,
paying a new TCP connection and HTTP/1 handshake every time. Worse, it
counted against the 32 in-flight relay cap, so a service handling more
than 32 concurrent requests failed outright.
Pool idle upstreams per endpoint and port, up to 8 each for 15s. Reuse is
safe because the pool only returns a connection hyper reports as ready,
and HTTP/1 cannot start a request until the previous body has drained.
Upgrades are never pooled since they take the connection over, and a
failed send evicts that endpoint. Pruning is bounded per key, with the
full sweep limited to once per 30s.
Signed-off-by: divesh <dgude@nvidia.com>
* fix(server): address HA gateway review findings (#3449)
- Let a gateway own supervisor sessions without a peer endpoint. Requiring
one whenever the store is PostgreSQL broke every single-instance
PostgreSQL deployment, because no sandbox supervisor could connect.
A cross-replica request to an owner that advertises no endpoint now fails
immediately naming the cause, instead of retrying until the wait timeout.
- Close a supervisor session on heartbeat only when another replica owns it,
or after renewals fail for the ownership TTL. A database error no longer
drops every session heartbeating during an outage.
- Clamp owner record ages at zero so a skewed or corrupt stored timestamp
cannot produce a negative age.
- Bound the cross-object advisory lock with a lock timeout, so a stuck holder
fails instead of blocking every mutation in the fleet.
- Refuse to start when a peer endpoint is configured on a multi-replica
backend but peer authentication is unavailable, and warn when a
multi-replica backend has no peer endpoint at all.
- Reject a plaintext peer endpoint when the gateway serves TLS.
- Skip the sandbox watch poller on single-replica backends, where the local
update bus already sees every write.
- Rate-limit the peer owner cache sweep so an insert no longer scans the
whole map under the lock.
- Retry GET and HEAD on a pooled upstream the sandbox closed, instead of
returning 502, and drop an emptied endpoint from the pool right away.
- Document the gateway peer environment variables and the post-rollout
ownership skew operators should expect.
Signed-off-by: divesh <dgude@nvidia.com>
* fix(server): harden HA supervisor ownership
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: divesh <dgude@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: divesh <dgude@nvidia.com>
Co-authored-by: Divesh Chowdary <47188680+FrostGod@users.noreply.github.com>
* fix(vm): unpack registry images correctly and validate prepared disks
The registry image-prep path expected `umoci raw unpack` to produce a
bundle-style rootfs/ subdirectory, but it extracts the image filesystem
directly into the target. Every registry prep therefore failed after a
successful unpack. Guest init exit codes do not survive the libkrun
boundary, so the failure looked like success and the broken disk was
cached, making every later sandbox for that image fail with "prepared
image disk missing /image-rootfs".
VM E2E started hitting this after the bootstrap image moved to
nvcr.io/nvidia/base/ubuntu:24.04: `--from base` no longer matches the
bootstrap image, so it now goes through registry prep.
- Accept umoci's direct extraction layout in the guest prep script.
- Build the image rootfs under a partial directory and rename it to
/image-rootfs only after every prep step succeeds.
- Check the prepared disk for /image-rootfs before caching it. On
failure, leave the cache untouched and report the prep console tail.
- Size the prep disk to hold the payload and the unpacked rootfs at the
same time. The community base image needs 1.40 GB + 3.32 GB, which
did not fit in the old payload*3 + 512 MiB.
Fixes#2358
Co-authored-by: s2cube <26961336+s2cube@users.noreply.github.com>
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
* test(e2e): run tool-dependent VM tests from the community base image
The host_gateway_alias and vm_corporate_proxy workloads run curl and
python3. The VM driver now defaults to nvcr.io/nvidia/base/ubuntu:24.04,
which ships neither, so these tests fail in VM E2E with "command not
found". Request the community base image explicitly with `--from base`.
Docker, Podman, and Kubernetes E2E already default to that image, so
their behavior is unchanged.
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
---------
Signed-off-by: Emilien Macchi <emacchi@redhat.com>
Co-authored-by: s2cube <26961336+s2cube@users.noreply.github.com>
* feat(cli): promote profile commands to top level
Add profile discovery and management commands with shared handlers for
the existing provider entry points. List a flat catalog across scopes
and follow continuation tokens through full and short pages.
Describe metadata, credentials, endpoints, TLS inspection, and MCP access
settings while preserving complete JSON/YAML definitions. Cover parser
equivalence, scope forwarding, pagination, and inspection settings with
focused unit and compiled-CLI integration tests.
Update docs, public skills, examples, and E2E command invocations.
Refs #2588
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(cli): remove redundant workspace selector qualification
Use the imported WorkspaceSelector in the provider integration helper so
the target passes Clippy with warnings denied.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
Previously, Network Activity could be constructed without a source or
destination endpoint, allowing connection, accept, relay, and configuration
events to violate the OCSF 1.8 endpoint constraint.
Now, NetworkActivityBuilder requires a source or destination endpoint at
compile time. Connection failures identify the workload peer or genuine
transparent destination, listener failures identify the listening endpoint,
and mediation-lane failures use Application Lifecycle rather than fabricated
network endpoints. Malformed forward requests use HTTP Activity with a
method-only request, generated 400 response, and workload peer.
Additionally, Unix relay-channel events use Base Event, policy-validation
warnings use Config State Change, and the unused bypass monitor is removed
because the current isolation architecture no longer uses it.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(policy): reject unknown endpoint security modes
Closes#3046
Validate TLS, enforcement, and access values across policy and provider profile ingress, and prevent runtime parsing from falling back to audit for unknown enforcement values.
Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>
* fix(policy)!: use enums for endpoint security modes
Replace the public TLS, enforcement, and access strings with protobuf enums and carry the typed values through policy composition, provider profiles, drivers, and runtime conversion.
Preserve the documented YAML spellings, reject unknown and invalid numeric enum values consistently, and update generated Go bindings, SDK conversions, tests, and policy documentation.
Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>
---------
Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>
The libkrun and libkrunfw projects moved from the containers/ GitHub
org to libkrun/. Update all clone URLs and version pins accordingly.
libkrun v1.19.4 introduces krun_add_virtiofs4 with a permissions
semantics parameter, enabling correct file ownership mapping on
host-to-guest bind mounts — required for issue #2585.
- libkrun: v1.17.4 → v1.19.4 (libkrun/libkrun)
- libkrunfw: 463f717b → v5.6.1 (libkrun/libkrunfw)
- Centralise LIBKRUN_REF in pins.env instead of hardcoding in each
build script
- Source pins.env in the macOS build script for consistency with the
Linux build script
- Remove the init/init make target step — since libkrun/libkrun@05c4eb7
the init blob moved into the init_blob crate and is built by cargo
Ref: https://github.com/NVIDIA/OpenShell/issues/2585
Signed-off-by: Florent Benoit <fbenoit@redhat.com>
* chore(providers): remove dead provider plugin modules
Twelve modules under crates/openshell-providers/src/providers/ were never
declared in providers/mod.rs, so they have not been compiled since the plugin
registry was narrowed to the two adapters it still registers. Four of them
(generic, gitlab, opencode, outlook) key off provider type identifiers that
normalize_provider_type already retired.
Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new
registers.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* test(providers): load example profiles from providers/ at test time
Add an example_profiles module that reads the YAML under providers/ from the
source checkout at run time and parses it with the existing profile loader. It
is gated behind a new non-default example-profiles feature so it is available to
this crate's own tests and, once wired into dev-dependencies, to gateway and CLI
tests, while never reaching a release binary.
Nothing consumes it yet; later changes move the test fixtures off the compiled
catalog and onto these files.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* test: source profile fixtures from providers/ instead of the compiled catalog
Unit and integration fixtures reached into builtin_profiles() to get a profile
to work with, which ties the tests to the compiled catalog rather than to the
files an operator would import. Point them at the example_profiles loader
instead.
The profiles.rs tests keep their coverage unchanged and become explicit golden
tests over providers/*.yaml: they are what keeps those files valid once nothing
compiles them. builtin_profiles_are_sorted_by_id widens into a test that the
whole example set parses, sorts, and lints clean as one catalog.
The CLI fake gateways serve the example profiles the way a real gateway serves
what an operator imported, through a shared helper.
No behavior change: the gateway still loads the built-in source by default and
serves the same profiles.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* docs(providers): document the example provider profiles
Each file under providers/ now opens with a header naming its expected client
binary identities, the image layout those paths assume, the credential scope,
the endpoint access it grants, and a smoke test. Add a README covering the
import commands and why a profile should be copied and edited rather than
imported unchanged.
Several of these profiles bind network access to paths that only exist in the
OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server,
/usr/lib/node_modules. Imported unchanged into another image the profile matches
nothing: the catalog still advertises it, but the credential is never injected
and the traffic is denied. The headers say so where it applies.
Comments only; the profile schema has no documentation fields and all fifteen
files still lint clean.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* test(server): seed unit-test state with the example provider profiles
Server unit tests inherited the builtin + user source default from ServerState,
so around a hundred and forty assertions about github, openai and the rest
resolved against the compiled catalog. Point the test state at the user source
alone and import the example profiles from providers/ into its store first, the
way an operator would. Every one of those assertions keeps passing unchanged,
which is the point: it proves the gateway behaves identically with an imported
catalog before the default moves.
Four tests asserted builtin-source semantics specifically. Under import-only the
only profiles a gateway cannot edit are the ones a non-user source vends, so
they now exercise a source-managed profile composed with the user source; the
read-only guard they cover is the one that still applies to interceptor
catalogs. The list test asserts the imported profile's user/platform identity
instead of builtin with an empty scope.
test_server_state_with_user_only_github_profile is gone: the default test state
now is a user-only gateway with github imported.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* refactor(cli): resolve credential suggestions from the gateway catalog
The --env credential warning scanned a profile table compiled into the CLI, so
its suggestions described the binary rather than the gateway the user is talking
to: a profile the gateway does not serve was suggested anyway, and an imported
custom profile never was.
Fetch the catalog from the connected gateway instead. The warning moves out of
argument parsing and into sandbox create and sandbox template create, where a
client already exists. A catalog fetch failure is not fatal — the warning
degrades to its generic form rather than blocking sandbox creation.
Extract the ListProviderProfiles paging loop from provider list-profiles into a
shared fetch_provider_profile_catalog, which the profile-driven paths now share.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* refactor(cli)!: infer providers from the gateway catalog
Command-to-provider inference went through a hardcoded alias table in
openshell-providers, a second copy of the built-in catalog's identifiers. A
custom profile could never be inferred no matter what binaries it declared, and
the table drifted from the profiles it mirrored.
Infer from the connected gateway's catalog instead: match the command's basename
against each profile's ID and against the basenames of the binaries the profile
authorizes. A profile that names /usr/bin/claude is the profile for running
claude. The match must be unique — where several profiles claim a command, the
user names one with --provider — and an empty catalog infers nothing.
No command is special-cased. `binaries` is the operator's authorization
statement, so a profile that declares a binary claims the command that runs it,
whatever that binary is; narrowing that belongs in the profile rather than in a
list compiled into the CLI, which could never cover an unbounded catalog anyway.
Breaking: the retired aliases stop resolving, and commands the old table never
listed can now infer. git, pip and uv are declared by the github and pypi
example profiles, so they infer where they previously did not.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* refactor(providers,server)!: resolve provider profiles by exact ID
normalize_provider_type was a hardcoded alias map — gh to github, claude to
claude-code, vertex to google-vertex-ai — and a second copy of the built-in
catalog's identifiers, independent of the YAML it mirrored. It let a provider
type resolve to a profile the operator never named, and it made the built-in IDs
behave as a reserved namespace.
Remove it, along with detect_provider_from_command and the alias-normalizing
ProviderRegistry::inject_env. A provider type now names a profile exactly:
- the effective catalog resolves an ID or reports it absent, with no alias retry
- plugins activate only for the ID of a profile the gateway resolved; a provider
with no resolvable profile gets no plugin projection, instead of falling back
to an alias guess
- the CLI surfaces the gateway's not-found instead of retrying under an alias
- telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets
go with the aliases that were their only source
Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve.
Use the profile's own ID.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* feat(server): fail closed when a provider's profile is absent
A sandbox composed from a provider whose profile the gateway cannot resolve
started anyway: the credential and policy builders warned and skipped, so the
sandbox came up carrying none of that provider's credentials or network policy.
The operator learned about it later, as a denied connection or a missing
environment variable, rather than as the configuration error it is.
Check at the two composition boundaries — CreateSandbox and
AttachSandboxProvider — right beside the catalog snapshot already taken there,
and reject with a bounded diagnostic naming the provider, the profile it refers
to, and the import command that supplies it. Scope-aware: a platform-scoped
provider is told to import with --global.
Read paths are untouched. ListProviders, GetProvider and profile export keep
working so an operator can see and recover an affected provider, and the shared
policy builders keep their warn-and-skip for the diagnostic paths that also
reach them.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* refactor(server): scope the vendor base-URL pin to declared endpoints
provider_profile_endpoints_are_active withheld a credential and the provider's
policy layer when an openai or anthropic provider pointed its client somewhere
other than the public vendor endpoint. The guard was keyed on
profile.source == "builtin" and on those two profile IDs, so it protected only
profiles OpenShell shipped. Once profiles are import-only no profile is ever
builtin, and the guard would silently stop applying — including to an operator
who imported providers/openai.yaml verbatim.
Key it on the profile instead of on where the profile came from. A profile's
endpoints are the boundary its credential is bound to, so if the provider
configures a *_BASE_URL pointing at a host the profile does not declare, the
profile no longer describes where that credential goes and is treated as
endpointless. Host matching reuses the DNS-label-aware matcher in
openshell-core, so wildcard endpoints such as Vertex's
*-aiplatform.googleapis.com resolve correctly.
A profile with no declared endpoints has no boundary to contradict, and a config
value that names no host is not a redirect this can reason about. Both keep the
profile active. The control now covers every endpoint-bearing profile, including
an operator's own, and the last "builtin" string leaves the gateway.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* test(e2e): import example provider profiles during gateway bring-up
The e2e suites create providers from github, openai, nvidia, claude-code and
google-cloud, and the GitHub lane runs a real git clone through the profile's
binary attribution. A gateway serves only the profiles an operator imported, so
the lanes have to import them.
Add e2e_import_example_provider_profiles to the shared bring-up helpers and call
it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is
healthy and registered. It runs the same command the upgrade notes give
operators, against the repository's own providers/ directory, so the lanes
exercise the documented path rather than a test-only shortcut.
Lands before the default changes: a same-ID user profile already shadows a
built-in, so importing works today.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* refactor(providers)!: make provider profiles import-only
OpenShell compiled fifteen provider profile YAML files into every release binary
and selected them by default, so a fresh gateway published a catalog it was
never configured with. Those profiles are not image-neutral: their binary
selectors name paths that exist in the OpenShell Community image, so changing
the sandbox image could make a profile inert while the catalog still advertised
it. Every endpoint, credential name and binary path in providers/ was also
effectively part of the 0.1.0 public contract.
A gateway's catalog is now exactly what an operator imported:
- providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone
from the openshell-providers API
- the default provider_profile_sources is [{ type = "user" }], and a gateway
with nothing imported reaches ready and serves an empty catalog — an empty
catalog is a valid state, not a startup failure
- the builtin source type is removed from the configuration schema, and a
gateway.toml that still names it is rejected at parse time with the import
command rather than an unknown-variant error
- nothing reserves the canonical identifiers any more, so github, pypi,
anthropic and the rest import at their own IDs and the imported profile is the
only definition for that ID; the static_fallback that kept a shipped
definition resident behind an imported one is gone with the source that
produced it
A collision between an interceptor-vended profile and an imported one still
fails closed, which is the behavior that shadowing quietly bypassed for
built-ins.
Operators upgrading should export the profiles their deployment relies on
before upgrading, or copy them from the providers/ directory of the matching
release tag, then import them at the scope their providers use.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* docs(providers): document import-only provider profiles
Rewrite the provider profile documentation around a catalog the operator builds
rather than one the gateway ships.
- gateway-config: the default is [{ type = "user" }], the builtin source type is
gone and rejected at startup, and an empty catalog is a valid ready state
- profiles: replace the built-in profile table and its shadowing semantics with
the import workflow, and say plainly that a profile whose binary paths do not
match the image is inert
- manage-providers: replace the two fixed provider-type tables with
`provider list-profiles`, and describe catalog-driven command inference
- inference-routing: drop the "still loads built-in profiles for compatibility"
transition text; its migration walkthrough is now the normal path
- quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex
tutorials: import the profile before creating the provider, since nothing
resolves without it
- release notes: a 0.1.0 migration section covering export-before-upgrade,
importing from the release tag's providers/ directory, what happens to a
provider whose profile is missing, and the removal of the legacy type aliases
- README, architecture, the openshell-cli and debug-openshell-cluster skills,
and the governance interceptor example follow the same change; the example's
smoke assertion now imports a profile to prove an authoritative interceptor
hides it
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding
Addresses two findings from review (GATOR-c1bd9867-02 and -03).
The provider profile catalog lookup collapsed its error into an empty
catalog, so an authorization, availability or transport failure on
ListProviderProfiles was indistinguishable from a gateway that genuinely has
no profiles. Command inference then resolved nothing and the sandbox was
created without the provider it needed, deferring the failure to the workload.
Keep the lookup's outcome instead of discarding it. The credential warning is
advisory and still degrades to its generic form, but inference now consults the
catalog only when there is a command to resolve and surfaces the lookup failure
when there is, naming the gateway and pointing at explicit --provider selection.
Two regression tests cover it through a fake gateway whose ListProviderProfiles
returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails
without sending a provider-less create request, and a sandbox with no trailing
command still succeeds because it needs no catalog.
The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes
deliberately skip gateway registration and start the gateway without a TLS
client CA, so no mTLS identity exists and no token has been acquired when the
import runs; the wrapper exited during setup. Skipping the import is not
sufficient because the provider tests now require the claude-code profile.
Add e2e_register_oidc_admin_session, which mints an administrator token with
Keycloak's password grant — the same grant the OIDC test helpers use — and
writes the gateway metadata and token bundle that an interactive login would
have stored, so the import runs as an authenticated administrator. The mTLS
lanes keep the direct import unchanged. Both affected wrappers are covered:
with-podman-gateway.sh had the same defect as with-docker-gateway.sh.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* refactor(cli)!: remove command-derived provider attachment
Addresses GATOR-c1bd9867-01.
A profile's `binaries` list authorizes a binary to reach that profile's
endpoints. It is not a statement that running the binary asks for the provider,
and reading it as attachment intent let a command silently gain provider
authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash`
resolved an existing aws-s3 provider and attached its credential-backed
capability and network policy to the shell and its descendants. Attachment of
an already-created provider never prompted, so the confirmation flow did not
guard it.
The distinction that would make inference sound — whether a declared binary is
a profile's client or merely a permitted runtime — cannot be expressed:
NetworkBinary carries only a path, and the field that encoded it was removed in
0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant
does not help, because uniqueness measures how sparse the catalog is rather
than what the user intended, and a compiled list of "generic" commands could
never cover an unbounded operator catalog.
Remove trailing-command inference rather than approximate it. A provider is
attached only when named with --provider, which still creates a missing
provider from local discovery when the name matches an imported profile ID.
Sandboxes with no providers remain a normal, fully supported state.
Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving
authority from the catalog, its only remaining use is the advisory credential
warning, so a failed lookup degrades that warning instead of blocking creation.
The regression test now asserts that an unreachable catalog still creates the
sandbox and attaches nothing.
While repurposing the deduplication test, a pre-existing defect surfaced:
repeating a name in --provider auto-created it twice and the second attempt
failed with "provider already exists", because the explicit pass never
consulted the set of names it had already handled. Guard it.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(e2e): load the OIDC token and trust the gateway when seeding profiles
Addresses the carried finding GATOR-c1bd9867-03.
e2e_register_oidc_admin_session established a session the CLI never used. The
CLI decides whether to load a stored bearer token by matching on the auth_mode
field of the gateway metadata alone; the helper omitted that field, so the
metadata fell through to the default arm and oidc_token.json stayed on disk
unread. The profile import went out unauthenticated exactly as it had before
the helper existed. The helper also installed no trust anchor, leaving the CLI
unable to verify the gateway's self-signed serving certificate.
Write auth_mode = "oidc" so the stored token is loaded, and install the CA at
<gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or
key on disk the CLI falls back to CA-only server verification and authenticates
with the bearer token, which is what these lanes need because they start the
gateway without --tls-client-ca. Certificate verification stays on; no insecure
transport override is introduced.
Assert the session before anything depends on it. ListProviderProfiles is
annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded
and accepted. The helper now fails at that point, naming the gateway config
directory and echoing the CLI output, rather than letting the defect surface
later as an opaque profile import error.
Both affected wrappers pass the PKI directory and CLI binary the helper needs;
the mTLS lanes are untouched.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(e2e): request the OpenShell scopes when minting the admin token
The OIDC lanes failed at the session assertion added in b25dd0298 with
PERMISSION_DENIED and "scope 'provider:read' required". Authentication was
working -- the gateway logged ListProviderProfiles at gRPC status 7, which it
can only reach once the bearer token has been loaded and accepted. The token
simply carried no OpenShell scope.
sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional
client scopes on the openshell-cli client in scripts/keycloak-realm.json, so
Keycloak mints them only when the request asks for them. The password grant
here asked for nothing, leaving the realm defaults (openid, profile, email,
roles, web-origins, acr) and an access token that authorizes no RPC.
Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already
does for its administrator tokens. openshell:all is SCOPE_ALL in
crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls
without enumerating them. Verified against the realm: the token goes from
"email profile openid" to "email profile openshell:all openid".
Record the same scopes in metadata.json. The helper writes the bundle
openshell gateway login would have stored, and oidc_scopes is the field that
login path reads back, so leaving it out would misdescribe the stored token.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(e2e): seed provider profiles once per VM lane gateway
The VM lanes failed on the second test target with "custom provider profile
'anthropic' already exists" for all fifteen example profiles, and the import
exited non-zero.
e2e_import_example_provider_profiles was called from run_e2e_test, so it ran
once per target -- four times against one long-lived gateway. That was
harmless while import overwrote silently, but profiles are now import-only:
ImportProviderProfiles is create-only and reports an existing id as an
error-severity diagnostic, with no overwrite flag on the request. The first
import therefore succeeds and every later one fails.
Hoist the call to just after the conformance run, which is where the docker,
podman and kube lanes already seed their catalogs. The profiles persist for
the gateway's lifetime, so every target still finds them, and the
E2E_TEST_OVERRIDE path is covered by the same single call.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* test(e2e): name the provider explicitly in the OIDC workspace-user test
user_can_create_sandbox_with_inferred_provider_command reached the gateway
once the admin session was fixed, and then failed: it asserted a
missing-provider error, but the sandbox was created and died provisioning
with "failed to spawn sandbox entrypoint process 'claude-code'".
The test drove provider resolution by passing claude-code as the trailing
command and relying on the CLI to infer the provider type from it. That
inference is what this branch removed, so the trailing word is now nothing
but an entrypoint, and the image has no such binary.
The regression the test guards is not inference itself: it is that a
workspace user resolving a provider is not gated behind Platform Admin.
Name the provider with --provider, the only remaining way to attach one.
The CLI still has to fetch the claude-code profile before it can auto-create
the provider, so the lookup a workspace user must be allowed to make still
happens, and auto-creation still stops at the non-interactive branch with
"missing required provider". Both assertions therefore keep their meaning.
Rename the test and rework its comments to describe what it now exercises;
the old name would otherwise outlive the behavior it was named for.
Signed-off-by: Philippe Martin <phmartin@redhat.com>
* fix(providers): reconcile import-only profiles with main
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>