Commit Graph
84 Commits
Author SHA1 Message Date
Drew Newberry 021400be8a refactor(auth): separate sandbox identity from TLS (#3110)
* refactor(auth): separate sandbox identity from TLS

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(auth): clarify gateway mTLS behavior

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(auth): include workspace scope in TLS authorization checks

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): bound service auth sandbox names for large PIDs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-10-01 04:33:25 +00:00
Mrunal Patel bdffa102c3 feat(api): add durable exec launch admission (#3324)
* feat(api): add durable exec launch admission

Fence duplicate exec launches with keyed durable admission and producer-owned terminal completion. Keep uncertain launches unresolved and never replay output or interactive input.

Part of #3051 (phase 4a).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): fence exec identity across authorization lookups

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-22 21:52:19 +00:00
Varsha 7139df8ca5 fix(exec): preserve output after stdin EOF and verify stream completion (#3359)
* fix(exec): preserve output after stdin EOF and verify stream completion

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(exec): cover fair duplex progress and live SDK completion

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(sdk): compare large exec buffers with native equality

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* test(exec): use workspace-scoped sandbox name in EOF regression

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

---------

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
2026-09-22 14:11:06 +00:00
50230616d5 refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image

default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(deploy): default deployment configs to the official Alpine sandbox image

Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* feat(driver): default to numeric non-root identity for USER-less images

With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.

Part of #3116.

Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(conformance): use Alpine workload image

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(policy): drop community image /app path from default policy

The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.

Part of #3116.

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>

* docs(config): document Alpine default images

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): report early sandbox termination

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): initialize rootless workspace ownership

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(sandbox): qualify NVIDIA Ubuntu default

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): initialize rootful default workspace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(sftp): add native sandbox adapter

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): gate runtime helper support to Linux

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): support standard OpenSSH file operations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sftp): harden rename and special file handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(runtime): remove community image dependencies

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): build provider readiness tool fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): use a dedicated Noble fixture for Docker tests

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-09-22 14:43:51 +02:00
Drew Newberry d91b1999a0 feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(cli): update forward color fixture for workspace scope

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(supervisor): use sandbox names for settings lookup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox request fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): use canonical sandbox receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): harden sandbox mutation handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): update rebased sandbox references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(api)!: standardize canonical entity references

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(api): codify protobuf API conventions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve workspace selector semantics

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): restore workspace selector parity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): preserve descriptive name fields

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(api): update e2e request fixtures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(core): omit workspace selector during bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): remove proto convention checker

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(api): refresh schema fingerprints after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(cli): use canonical provider receipt field

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-18 11:39:55 -07:00
Krzysztof Malczuk e38d7254e6 fix(policy): reject unknown endpoint security modes (#3187)
* fix(policy): reject unknown endpoint security modes

Closes #3046

Validate TLS, enforcement, and access values across policy and provider profile ingress, and prevent runtime parsing from falling back to audit for unknown enforcement values.

Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>

* fix(policy)!: use enums for endpoint security modes

Replace the public TLS, enforcement, and access strings with protobuf enums and carry the typed values through policy composition, provider profiles, drivers, and runtime conversion.

Preserve the documented YAML spellings, reject unknown and invalid numeric enum values consistently, and update generated Go bindings, SDK conversions, tests, and policy documentation.

Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>

---------

Signed-off-by: Krzysztof Malczuk <kmalczuk@redhat.com>
2026-09-18 16:48:33 +00:00
Philippe MartinandJohn Myers 04146692d9 refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules

Twelve modules under crates/openshell-providers/src/providers/ were never
declared in providers/mod.rs, so they have not been compiled since the plugin
registry was narrowed to the two adapters it still registers. Four of them
(generic, gitlab, opencode, outlook) key off provider type identifiers that
normalize_provider_type already retired.

Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new
registers.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(providers): load example profiles from providers/ at test time

Add an example_profiles module that reads the YAML under providers/ from the
source checkout at run time and parses it with the existing profile loader. It
is gated behind a new non-default example-profiles feature so it is available to
this crate's own tests and, once wired into dev-dependencies, to gateway and CLI
tests, while never reaching a release binary.

Nothing consumes it yet; later changes move the test fixtures off the compiled
catalog and onto these files.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test: source profile fixtures from providers/ instead of the compiled catalog

Unit and integration fixtures reached into builtin_profiles() to get a profile
to work with, which ties the tests to the compiled catalog rather than to the
files an operator would import. Point them at the example_profiles loader
instead.

The profiles.rs tests keep their coverage unchanged and become explicit golden
tests over providers/*.yaml: they are what keeps those files valid once nothing
compiles them. builtin_profiles_are_sorted_by_id widens into a test that the
whole example set parses, sorts, and lints clean as one catalog.

The CLI fake gateways serve the example profiles the way a real gateway serves
what an operator imported, through a shared helper.

No behavior change: the gateway still loads the built-in source by default and
serves the same profiles.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document the example provider profiles

Each file under providers/ now opens with a header naming its expected client
binary identities, the image layout those paths assume, the credential scope,
the endpoint access it grants, and a smoke test. Add a README covering the
import commands and why a profile should be copied and edited rather than
imported unchanged.

Several of these profiles bind network access to paths that only exist in the
OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server,
/usr/lib/node_modules. Imported unchanged into another image the profile matches
nothing: the catalog still advertises it, but the credential is never injected
and the traffic is denied. The headers say so where it applies.

Comments only; the profile schema has no documentation fields and all fifteen
files still lint clean.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(server): seed unit-test state with the example provider profiles

Server unit tests inherited the builtin + user source default from ServerState,
so around a hundred and forty assertions about github, openai and the rest
resolved against the compiled catalog. Point the test state at the user source
alone and import the example profiles from providers/ into its store first, the
way an operator would. Every one of those assertions keeps passing unchanged,
which is the point: it proves the gateway behaves identically with an imported
catalog before the default moves.

Four tests asserted builtin-source semantics specifically. Under import-only the
only profiles a gateway cannot edit are the ones a non-user source vends, so
they now exercise a source-managed profile composed with the user source; the
read-only guard they cover is the one that still applies to interceptor
catalogs. The list test asserts the imported profile's user/platform identity
instead of builtin with an empty scope.

test_server_state_with_user_only_github_profile is gone: the default test state
now is a user-only gateway with github imported.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli): resolve credential suggestions from the gateway catalog

The --env credential warning scanned a profile table compiled into the CLI, so
its suggestions described the binary rather than the gateway the user is talking
to: a profile the gateway does not serve was suggested anyway, and an imported
custom profile never was.

Fetch the catalog from the connected gateway instead. The warning moves out of
argument parsing and into sandbox create and sandbox template create, where a
client already exists. A catalog fetch failure is not fatal — the warning
degrades to its generic form rather than blocking sandbox creation.

Extract the ListProviderProfiles paging loop from provider list-profiles into a
shared fetch_provider_profile_catalog, which the profile-driven paths now share.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: infer providers from the gateway catalog

Command-to-provider inference went through a hardcoded alias table in
openshell-providers, a second copy of the built-in catalog's identifiers. A
custom profile could never be inferred no matter what binaries it declared, and
the table drifted from the profiles it mirrored.

Infer from the connected gateway's catalog instead: match the command's basename
against each profile's ID and against the basenames of the binaries the profile
authorizes. A profile that names /usr/bin/claude is the profile for running
claude. The match must be unique — where several profiles claim a command, the
user names one with --provider — and an empty catalog infers nothing.

No command is special-cased. `binaries` is the operator's authorization
statement, so a profile that declares a binary claims the command that runs it,
whatever that binary is; narrowing that belongs in the profile rather than in a
list compiled into the CLI, which could never cover an unbounded catalog anyway.

Breaking: the retired aliases stop resolving, and commands the old table never
listed can now infer. git, pip and uv are declared by the github and pypi
example profiles, so they infer where they previously did not.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers,server)!: resolve provider profiles by exact ID

normalize_provider_type was a hardcoded alias map — gh to github, claude to
claude-code, vertex to google-vertex-ai — and a second copy of the built-in
catalog's identifiers, independent of the YAML it mirrored. It let a provider
type resolve to a profile the operator never named, and it made the built-in IDs
behave as a reserved namespace.

Remove it, along with detect_provider_from_command and the alias-normalizing
ProviderRegistry::inject_env. A provider type now names a profile exactly:

- the effective catalog resolves an ID or reports it absent, with no alias retry
- plugins activate only for the ID of a profile the gateway resolved; a provider
  with no resolvable profile gets no plugin projection, instead of falling back
  to an alias guess
- the CLI surfaces the gateway's not-found instead of retrying under an alias
- telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets
  go with the aliases that were their only source

Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve.
Use the profile's own ID.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* feat(server): fail closed when a provider's profile is absent

A sandbox composed from a provider whose profile the gateway cannot resolve
started anyway: the credential and policy builders warned and skipped, so the
sandbox came up carrying none of that provider's credentials or network policy.
The operator learned about it later, as a denied connection or a missing
environment variable, rather than as the configuration error it is.

Check at the two composition boundaries — CreateSandbox and
AttachSandboxProvider — right beside the catalog snapshot already taken there,
and reject with a bounded diagnostic naming the provider, the profile it refers
to, and the import command that supplies it. Scope-aware: a platform-scoped
provider is told to import with --global.

Read paths are untouched. ListProviders, GetProvider and profile export keep
working so an operator can see and recover an affected provider, and the shared
policy builders keep their warn-and-skip for the diagnostic paths that also
reach them.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(server): scope the vendor base-URL pin to declared endpoints

provider_profile_endpoints_are_active withheld a credential and the provider's
policy layer when an openai or anthropic provider pointed its client somewhere
other than the public vendor endpoint. The guard was keyed on
profile.source == "builtin" and on those two profile IDs, so it protected only
profiles OpenShell shipped. Once profiles are import-only no profile is ever
builtin, and the guard would silently stop applying — including to an operator
who imported providers/openai.yaml verbatim.

Key it on the profile instead of on where the profile came from. A profile's
endpoints are the boundary its credential is bound to, so if the provider
configures a *_BASE_URL pointing at a host the profile does not declare, the
profile no longer describes where that credential goes and is treated as
endpointless. Host matching reuses the DNS-label-aware matcher in
openshell-core, so wildcard endpoints such as Vertex's
*-aiplatform.googleapis.com resolve correctly.

A profile with no declared endpoints has no boundary to contradict, and a config
value that names no host is not a redirect this can reason about. Both keep the
profile active. The control now covers every endpoint-bearing profile, including
an operator's own, and the last "builtin" string leaves the gateway.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): import example provider profiles during gateway bring-up

The e2e suites create providers from github, openai, nvidia, claude-code and
google-cloud, and the GitHub lane runs a real git clone through the profile's
binary attribution. A gateway serves only the profiles an operator imported, so
the lanes have to import them.

Add e2e_import_example_provider_profiles to the shared bring-up helpers and call
it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is
healthy and registered. It runs the same command the upgrade notes give
operators, against the repository's own providers/ directory, so the lanes
exercise the documented path rather than a test-only shortcut.

Lands before the default changes: a same-ID user profile already shadows a
built-in, so importing works today.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(providers)!: make provider profiles import-only

OpenShell compiled fifteen provider profile YAML files into every release binary
and selected them by default, so a fresh gateway published a catalog it was
never configured with. Those profiles are not image-neutral: their binary
selectors name paths that exist in the OpenShell Community image, so changing
the sandbox image could make a profile inert while the catalog still advertised
it. Every endpoint, credential name and binary path in providers/ was also
effectively part of the 0.1.0 public contract.

A gateway's catalog is now exactly what an operator imported:

- providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone
  from the openshell-providers API
- the default provider_profile_sources is [{ type = "user" }], and a gateway
  with nothing imported reaches ready and serves an empty catalog — an empty
  catalog is a valid state, not a startup failure
- the builtin source type is removed from the configuration schema, and a
  gateway.toml that still names it is rejected at parse time with the import
  command rather than an unknown-variant error
- nothing reserves the canonical identifiers any more, so github, pypi,
  anthropic and the rest import at their own IDs and the imported profile is the
  only definition for that ID; the static_fallback that kept a shipped
  definition resident behind an imported one is gone with the source that
  produced it

A collision between an interceptor-vended profile and an imported one still
fails closed, which is the behavior that shadowing quietly bypassed for
built-ins.

Operators upgrading should export the profiles their deployment relies on
before upgrading, or copy them from the providers/ directory of the matching
release tag, then import them at the scope their providers use.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* docs(providers): document import-only provider profiles

Rewrite the provider profile documentation around a catalog the operator builds
rather than one the gateway ships.

- gateway-config: the default is [{ type = "user" }], the builtin source type is
  gone and rejected at startup, and an empty catalog is a valid ready state
- profiles: replace the built-in profile table and its shadowing semantics with
  the import workflow, and say plainly that a profile whose binary paths do not
  match the image is inert
- manage-providers: replace the two fixed provider-type tables with
  `provider list-profiles`, and describe catalog-driven command inference
- inference-routing: drop the "still loads built-in profiles for compatibility"
  transition text; its migration walkthrough is now the normal path
- quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex
  tutorials: import the profile before creating the provider, since nothing
  resolves without it
- release notes: a 0.1.0 migration section covering export-before-upgrade,
  importing from the release tag's providers/ directory, what happens to a
  provider whose profile is missing, and the removal of the legacy type aliases
- README, architecture, the openshell-cli and debug-openshell-cluster skills,
  and the governance interceptor example follow the same change; the example's
  smoke assertion now imports a profile to prove an authoritative interceptor
  hides it

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding

Addresses two findings from review (GATOR-c1bd9867-02 and -03).

The provider profile catalog lookup collapsed its error into an empty
catalog, so an authorization, availability or transport failure on
ListProviderProfiles was indistinguishable from a gateway that genuinely has
no profiles. Command inference then resolved nothing and the sandbox was
created without the provider it needed, deferring the failure to the workload.

Keep the lookup's outcome instead of discarding it. The credential warning is
advisory and still degrades to its generic form, but inference now consults the
catalog only when there is a command to resolve and surfaces the lookup failure
when there is, naming the gateway and pointing at explicit --provider selection.
Two regression tests cover it through a fake gateway whose ListProviderProfiles
returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails
without sending a provider-less create request, and a sandbox with no trailing
command still succeeds because it needs no catalog.

The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes
deliberately skip gateway registration and start the gateway without a TLS
client CA, so no mTLS identity exists and no token has been acquired when the
import runs; the wrapper exited during setup. Skipping the import is not
sufficient because the provider tests now require the claude-code profile.

Add e2e_register_oidc_admin_session, which mints an administrator token with
Keycloak's password grant — the same grant the OIDC test helpers use — and
writes the gateway metadata and token bundle that an interactive login would
have stored, so the import runs as an authenticated administrator. The mTLS
lanes keep the direct import unchanged. Both affected wrappers are covered:
with-podman-gateway.sh had the same defect as with-docker-gateway.sh.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* refactor(cli)!: remove command-derived provider attachment

Addresses GATOR-c1bd9867-01.

A profile's `binaries` list authorizes a binary to reach that profile's
endpoints. It is not a statement that running the binary asks for the provider,
and reading it as attachment intent let a command silently gain provider
authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash`
resolved an existing aws-s3 provider and attached its credential-backed
capability and network policy to the shell and its descendants. Attachment of
an already-created provider never prompted, so the confirmation flow did not
guard it.

The distinction that would make inference sound — whether a declared binary is
a profile's client or merely a permitted runtime — cannot be expressed:
NetworkBinary carries only a path, and the field that encoded it was removed in
0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant
does not help, because uniqueness measures how sparse the catalog is rather
than what the user intended, and a compiled list of "generic" commands could
never cover an unbounded operator catalog.

Remove trailing-command inference rather than approximate it. A provider is
attached only when named with --provider, which still creates a missing
provider from local discovery when the name matches an imported profile ID.
Sandboxes with no providers remain a normal, fully supported state.

Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving
authority from the catalog, its only remaining use is the advisory credential
warning, so a failed lookup degrades that warning instead of blocking creation.
The regression test now asserts that an unreachable catalog still creates the
sandbox and attaches nothing.

While repurposing the deduplication test, a pre-existing defect surfaced:
repeating a name in --provider auto-created it twice and the second attempt
failed with "provider already exists", because the explicit pass never
consulted the set of names it had already handled. Guard it.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): load the OIDC token and trust the gateway when seeding profiles

Addresses the carried finding GATOR-c1bd9867-03.

e2e_register_oidc_admin_session established a session the CLI never used. The
CLI decides whether to load a stored bearer token by matching on the auth_mode
field of the gateway metadata alone; the helper omitted that field, so the
metadata fell through to the default arm and oidc_token.json stayed on disk
unread. The profile import went out unauthenticated exactly as it had before
the helper existed. The helper also installed no trust anchor, leaving the CLI
unable to verify the gateway's self-signed serving certificate.

Write auth_mode = "oidc" so the stored token is loaded, and install the CA at
<gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or
key on disk the CLI falls back to CA-only server verification and authenticates
with the bearer token, which is what these lanes need because they start the
gateway without --tls-client-ca. Certificate verification stays on; no insecure
transport override is introduced.

Assert the session before anything depends on it. ListProviderProfiles is
annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded
and accepted. The helper now fails at that point, naming the gateway config
directory and echoing the CLI output, rather than letting the defect surface
later as an opaque profile import error.

Both affected wrappers pass the PKI directory and CLI binary the helper needs;
the mTLS lanes are untouched.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): request the OpenShell scopes when minting the admin token

The OIDC lanes failed at the session assertion added in b25dd0298 with
PERMISSION_DENIED and "scope 'provider:read' required". Authentication was
working -- the gateway logged ListProviderProfiles at gRPC status 7, which it
can only reach once the bearer token has been loaded and accepted. The token
simply carried no OpenShell scope.

sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional
client scopes on the openshell-cli client in scripts/keycloak-realm.json, so
Keycloak mints them only when the request asks for them. The password grant
here asked for nothing, leaving the realm defaults (openid, profile, email,
roles, web-origins, acr) and an access token that authorizes no RPC.

Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already
does for its administrator tokens. openshell:all is SCOPE_ALL in
crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls
without enumerating them. Verified against the realm: the token goes from
"email profile openid" to "email profile openshell:all openid".

Record the same scopes in metadata.json. The helper writes the bundle
openshell gateway login would have stored, and oidc_scopes is the field that
login path reads back, so leaving it out would misdescribe the stored token.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(e2e): seed provider profiles once per VM lane gateway

The VM lanes failed on the second test target with "custom provider profile
'anthropic' already exists" for all fifteen example profiles, and the import
exited non-zero.

e2e_import_example_provider_profiles was called from run_e2e_test, so it ran
once per target -- four times against one long-lived gateway. That was
harmless while import overwrote silently, but profiles are now import-only:
ImportProviderProfiles is create-only and reports an existing id as an
error-severity diagnostic, with no overwrite flag on the request. The first
import therefore succeeds and every later one fails.

Hoist the call to just after the conformance run, which is where the docker,
podman and kube lanes already seed their catalogs. The profiles persist for
the gateway's lifetime, so every target still finds them, and the
E2E_TEST_OVERRIDE path is covered by the same single call.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* test(e2e): name the provider explicitly in the OIDC workspace-user test

user_can_create_sandbox_with_inferred_provider_command reached the gateway
once the admin session was fixed, and then failed: it asserted a
missing-provider error, but the sandbox was created and died provisioning
with "failed to spawn sandbox entrypoint process 'claude-code'".

The test drove provider resolution by passing claude-code as the trailing
command and relying on the CLI to infer the provider type from it. That
inference is what this branch removed, so the trailing word is now nothing
but an entrypoint, and the image has no such binary.

The regression the test guards is not inference itself: it is that a
workspace user resolving a provider is not gated behind Platform Admin.
Name the provider with --provider, the only remaining way to attach one.
The CLI still has to fetch the claude-code profile before it can auto-create
the provider, so the lookup a workspace user must be allowed to make still
happens, and auto-creation still stops at the non-interactive branch with
"missing required provider". Both assertions therefore keep their meaning.

Rename the test and rework its comments to describe what it now exercises;
the old name would otherwise outlive the behavior it was named for.

Signed-off-by: Philippe Martin <phmartin@redhat.com>

* fix(providers): reconcile import-only profiles with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Philippe Martin <phmartin@redhat.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:59:27 +00:00
John T. Myers b7a932a4d0 docs(inference): remove stale managed endpoint references (#3428)
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-17 22:54:20 +00:00
Mrunal Patel a316fd7832 feat(api): add durable workspace mutation admission and replay (#3321)
* feat(api): add durable workspace mutation admission and replay

Part of #3051 (phase 3a). Preserve unresolved admissions, reauthorize replay, and protect same-name replacements.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* feat(api): extend mutation replay through gateway interceptors (#3323)

Add typed durable replay receipts for 24 ordinary unary mutations, protect sensitive payload fingerprints, and revalidate intercepted retries without repeating post-commit observation.

Part of #3051 (phase 3b).

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(api): scope workspace request IDs by target

Include the requested workspace name in create/delete admission keys while leaving workspace UUID guards unset. Cover cross-target UUID reuse, replay, and missing targets with server and live gateway regressions.

Merge the latest phase-two SDK fixes and preserve the approved interceptor replay changes.

Refs #3051.

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-09-17 14:51:56 +00:00
Derek Carr d68b7069c3 refactor(proto)!: use well-known time types (#3113)
* refactor(proto)!: use well-known time types

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time migration behavior

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve timestamp boundary semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): convert sandbox token expiry to timestamp

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve time compatibility semantics

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proto): preserve exact endpoint and profile times

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(e2e): use duration for interactive exec timeout

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(sdk-go)!: remove legacy profile duration fields

BREAKING CHANGE: Go provider profile callers must use RefreshBefore, MaxLifetime, and CacheTTL with ProfileDuration instead of the whole-second fields.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-09-16 19:51:45 +00:00
Drew Newberry c1f2e7189f feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>

* refactor(isolation): name the interface crate explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): expose trusted host gateway

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(agents): inventory the MXC driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add mediated DNS transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): tighten interface error and digest contracts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): remove unrelated driver inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): define capability-free launch contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): seal confirmed boundary state

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate confirmation for external backend implementations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): clarify mediated DNS identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): unify typed network mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): bind launches to sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): initialize extended sandbox status

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add boundary protocol and Linux primitives

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden signals and separate process status from transport

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate remote confirmation through public contract

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): validate wire state and propagate snapshot failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(isolation): import owned agent specification explicitly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(isolation): describe mediated DNS channel

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): bound mediation attach without nested retries

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): generalize loopback protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add transport-neutral session authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): separate sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime boundary controls

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): add terminal boundary operation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): split supervisor and sandbox runtimes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): harden boundary isolation and lifecycle ownership

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): reject private root redirects and adopt typed errors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): preserve accept thread ownership on musl

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): isolate credential probes from filtered threads

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): return retained exec exit status to independent waiters

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound network mediation and preserve socket authorization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bound control admission and retire stale mediation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): select migrated drivers per stack layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): implement loopback connector

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(isolation): authenticate the Sandbox Protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): consume dedicated backend crate

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sandbox): align topology session fixture

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): align projected bootstrap bundle

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): validate refreshed credentials before rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): fail closed across supervisor disconnects

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): repair rebased sandbox CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* build(runtime): publish separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(config): configure the sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate sandbox binary linkage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(isolation): use backend and runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(sandbox): use a scratch runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): refresh schema and dependency policy

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(sandbox): bind reconnects to supervisor process

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: align runtime split operational guidance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(security): document Kubernetes runtime RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): enforce runtime lifecycle invariants

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(compute): identify sandbox start generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(server): restore sandbox launch sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): support authenticated runtime replacement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): bind sandbox session successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): retry pending sandbox successors

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(vm): run the supervisor outside the guest workload

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): use unified build toolchain

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): use sandbox runtime terminology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): own guest network bootstrap

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): expose guest init version

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): select native supervisor artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): guard guest init Linux symbols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): scope Linux test imports

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): avoid guest interface casts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): reconcile admitted sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): share resolved sandbox identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): surface host supervisor failures

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): include guest logs on supervisor exit

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate and clean runtime generations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): keep shared paths in the base layer

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): isolate workloads behind the host supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve host gateway alias resolution

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(docker): use separate sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): restore startup validation after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): narrow supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): close companion isolation gaps

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): align mediated network expectations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(docker): exercise mediated network paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): attach supervisor to managed network

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): defer supervisor recovery until gateway is ready

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): preserve workloads during session rotation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(docker): remove unrelated configuration RFC changes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): add proxy-pod isolation topology

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): use stable sandbox service authority

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(kubernetes): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): adapt proxy pods to current runtime APIs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): describe the single runtime placement

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(kubernetes): simplify sandbox orchestration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): validate deployment prerequisites

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): update Trivy Helm profile inventory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(kubernetes): update Trivy scan inventory count

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): reuse preloaded runtime images in e2e

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): type and clean runtime resources

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): make sandbox restarts recoverable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): preserve supervisor egress

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): adopt isolated sandbox and supervisor containers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): stage bootstrap archives at named volume destinations

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): rotate launch-scoped authentication

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use sandbox backend protocol

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): use host networking for supervisor

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(podman): split sandbox and supervisor images

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): repair rebase integration

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(podman): name the sandbox runtime directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provision supervisor CA runtime storage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): address isolation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): inspect Debian supervisor provenance

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): use libpod-compatible tmpfs options

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind verified sandbox runtime binary

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): provide external driver data directory

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): start sandbox before joining user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): separate supervisor user namespace

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make sandbox starts generation-aware

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): rotate restored sandbox sessions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): bind sandbox session lineage

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* perf(isolation): add TCP and DNS benchmark harnesses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): align benchmark timing and supported protocols

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): report TCP benchmark metrics accurately

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(perf): cancel failed worker startup

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(docker): build matching local supervisor image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): make local sandbox smoke test runnable

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): wire local sandbox runtime image

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(kubernetes): narrow sandbox service RBAC

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): validate split runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): harden runtime session handling

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(supervisor): add standalone network proxy role

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(rfc): remove implementation companion notes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* refactor(vm): standardize runtime release name

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(vm): pin renamed runtime artifacts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): persist sandbox runtime identity

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(runtime): restore branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(isolation): reconcile main after rebase

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(network): close unframed HTTP 1.0 responses

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* chore(isolation): preserve upstream OCSF updates

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(security): close credential and TLS replay paths

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(auth): make sandbox refresh retries idempotent

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-16 00:49:14 +00:00
Drew Newberry 1860010850 feat(sdk): add lazy pagination pagers (#3256)
* feat(sdk): add lazy pagination pagers

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): harden pager edge cases

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(sdk): cover initial resume token

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): fix all-workspaces pager examples

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:08 +00:00
Drew Newberry 33bbda3d33 refactor(persistence): adopt continuation-token pagination (#3249)
* refactor(persistence): adopt continuation-token pagination

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address continuation review findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(tui): recover completed list refreshes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): address review scalability findings

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(pagination): repair branch validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(go): use page size in template example

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-09-11 00:02:07 +00:00
Mrunal PatelandJohn Myers 90dbe5454b feat(api): add typed workspace selectors (#3245)
* feat(api)!: add typed workspace selectors

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(cli): preserve template workspace metadata

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* test(e2e): migrate workspace request selectors

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(api): update public schema inventory

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-10 20:38:50 +00:00
John T. Myers f4dc6be4b2 refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes

Closes #3172

Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve alternate upstream isolation

Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-09-09 18:47:22 +00:00
Shiju 592df3e014 feat(policy): preserve exact MCP revision allowlists (#3027)
* feat(mcp): add version-aware wire profile metadata

Signed-off-by: Shiju <shiju@nvidia.com>

* feat(policy): canonicalize MCP version allowlists

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): align MCP policy tests with current main

Signed-off-by: Shiju <shiju@nvidia.com>

* fix(policy): canonicalize supervisor protobuf ingress

Materialize defaultable MCP revisions before ambiguity checks, OPA construction, and sidecar delivery. Reject invalid sidecar policies with bounded errors.

Signed-off-by: Shiju <shiju@nvidia.com>

---------

Signed-off-by: Shiju <shiju@nvidia.com>
2026-09-05 04:24:49 +00:00
John T. MyersandJohn Myers 07df822090 feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative

Closes #1988

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): move profiles into provider navigation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): clarify provider attachment lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(tui): scroll provider profile picker

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): honor profile credential semantics

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): prefer exact profile IDs

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): harden authoritative profile adoption

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(oidc): align provider fixtures with profiles

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): preserve authoritative profile lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-01 19:22:07 +00:00
Seth Jennings 5206bc51b2 feat(sdk): add OAuth Client Credentials support to SDKs (#2907)
* feat(sdk): add renewable client credentials auth

Implement lazy OAuth client-credentials acquisition and renewal for the Python, TypeScript, and Go SDK clients, with shared security conformance coverage and service-account documentation.\n\nCloses #2803

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

* fix(sdk): address client credentials review

Signed-off-by: Seth Jennings <sjenning@redhat.com>

---------

Signed-off-by: Seth Jennings <sjenning@redhat.com>
2026-08-25 17:53:47 +00:00
Mrunal Patel 8d67250a5d fix(providers): keep refresh credential handles stable (#2780)
* fix(providers): keep refresh credential handles stable

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

* fix(providers): protect refresh-owned credentials

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>

---------

Signed-off-by: Mrunal Patel <mrunalp@gmail.com>
2026-08-18 20:45:36 +00:00
John T. MyersandJohn Myers 0120535efc feat(proxy): bind static credentials to provider endpoints (#2510)
* feat(proxy): bind static credentials to provider endpoints

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(e2e): verify static credential endpoint isolation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(e2e): use valid endpoint isolation fixtures

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(provider): explain static credential endpoint binding

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(credentials): preserve binding identity across rotations

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): enforce bindings across request lifecycle

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): close credential relay gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): hash selected provider profile scope

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): resolve credentials after request admission

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify binding failure diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(proxy): align single-route credential denials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): harden endpoint-bound rotation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce identity and authority binding

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): snapshot provider environment atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): include authority port in query proxy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): close credential revocation gaps

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(proxy): explain authority mismatch diagnostics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): enforce binding lifecycle invariants

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): reject credential config collisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): capture credential scope atomically

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): distinguish origin and absolute targets

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(provider): isolate endpointless profile credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): normalize IPv6 request authorities

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): clarify endpointless profile isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): bind endpointless provider credentials

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): use current GCP placeholder revision

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(providers): explain policy credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover endpointless fail-closed invariant

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(policy): expect ambiguity rejection at creation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): share credential mismatch finding builder

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): cover malformed binding metadata

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(credentials): verify multi-key endpoint isolation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(e2e): cover same-host credential path denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(credentials): document serialized refresh contract

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(proxy): consolidate L7 log formatting

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): precompile endpoint binding patterns

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* perf(credentials): share identity epoch revisions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): require explicit request default ports

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): validate SigV4 credential sources

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(credentials): preserve endpoint bindings for credential handles

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(go-sdk): expose network credential bindings

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-08-10 19:19:45 +00:00
John T. MyersandJohn Myers 905b554c7c refactor(network): consolidate proxy egress pipeline (#2373)
* refactor(network): introduce shared egress pipeline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover shared proxy egress paths

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): make destination authorization explicit

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): pin proxy relay policy context

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): lock relay generation contracts

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): establish phase zero compatibility baseline

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(policy): detect ambiguous network endpoints

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* feat(sandbox): fail closed on invalid policy updates

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* refactor(network): invalidate relays on policy changes

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): document validation failure posture

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): cover validation and middleware egress

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(config): move policy failure mode to gateway toml

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): name proxy contracts by behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): align overlap validation with endpoint selection

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): expect hard loopback denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): match declared endpoint denial

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): preserve path-specific endpoint overrides

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(network): respect hard-blocked host gateways

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): reconcile proxy refactor with main

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(network): preserve CONNECT policy generation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): cover runtime endpoint glob semantics

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): reject ambiguous policies before persistence

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(policy): explain ambiguity preflight behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(policy): compare body limits within protocol

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(proxy): avoid global tracing capture race

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* chore(server): format rebased provider tests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* test(server): authenticate rebased policy requests

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): retain runtime on middleware outage

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): preflight provider composition activation

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): distinguish runtime failure transitions

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-07-31 20:59:00 +00:00
Evan LezarandDrew Newberry d220d89468 feat(compute): negotiate gateway callback listeners (#2492)
* feat(compute): query gateway listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* feat(compute): add Podman listener requirements

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(docker): use default gateway bind address

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(gateway): avoid wildcard primary listener

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate callback listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): support split dual-stack listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): support legacy rootless listener discovery

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): accept loopback plaintext rejection

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(agent): add callback listener diagnostics

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(server): restrict compute callback listeners

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): validate local callback port

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(server): clarify callback listener contract

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(podman): require pasta for local callbacks

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* docs(gateway): document RPM listener default

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* refactor(server): keep listener provenance diagnostic-only

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* fix(compute): preserve callback listener isolation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): remove Podman callback relay

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(packaging): preserve Podman callback loopback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): run VM smoke on nested-virt runner

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): gate VM smoke on usable KVM

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): probe KVM through VM driver

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* ci(e2e): tolerate hosted KVM denial

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(server): close traced futures before assertions

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* revert: remove tracing test stabilization

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
2026-07-31 16:41:06 +00:00
Derek Carr 9c019a93f5 Wire authorization into workspace model (#2445)
* feat(auth): implement RFC 0011 Phase 2 workspace authorization

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address PR review feedback on workspace authorization

- Docker e2e: add --health-port and switch readiness probe from
  `openshell status` to `curl /healthz`, fixing a false-positive
  readiness check in OIDC mode where the CLI exited 0 without
  actually contacting the gateway

- ListWorkspaces: move membership filtering from post-query N+1
  lookups into a SQL EXISTS subquery so pagination applies to the
  visible set, not the global ordering. Add generic
  list_with_membership to the persistence layer.

- Descriptor validator: reject role/scope fields on unauthenticated
  and sandbox auth modes, and allow-list workspace_role as
  user/admin and global_role as platform_admin to catch typos at
  startup

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(server): use authed request in delete telemetry test

The workspace authorization added by the Phase 2 auth changes requires
a Principal on every delete request. The delete-telemetry test was still
using a bare Request::new, so extract_principal failed before the handler
could acquire the delete gate, causing a 5-second timeout flake.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator review findings for workspace authorization

- Inject unauthenticated-local-dev principal in no-auth gateway mode so
  handlers that call extract_principal() always find one.
- Cap label-selector membership query at MAX_PAGE_SIZE instead of
  u32::MAX to bound the in-memory read.
- Authorize workspace membership before resolving workspace existence in
  all sandbox RPCs to prevent workspace-name enumeration by non-members.
- Remove dead_code allow on AuthorizedWorkspace.workspace now that
  callers use the normalized name from the authz result.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): close workspace-name oracle and label-selector truncation

Swap authorize-before-resolve ordering in 27 handlers across
provider.rs, service.rs, policy.rs, and workspace.rs to prevent
CWE-203 workspace-name enumeration by non-members.

Add combined membership+label SQL query (list_with_membership_and_selector)
to both persistence backends so ListWorkspaces with label selectors no
longer silently drops results beyond the first page of membership matches.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): add non-member rejection and membership+label persistence tests

Add comprehensive test coverage for workspace authorization changes:
- Non-member rejection tests across all 44 workspace-scoped handlers
  (sandbox, provider, service, policy, workspace, inference) verifying
  PERMISSION_DENIED is returned instead of NOT_FOUND to prevent
  CWE-203 workspace-name oracle
- Persistence test for list_with_membership_and_selector verifying
  SQL-level membership EXISTS + label filtering, multiple predicates,
  no-match cases, and pagination

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): format merged import line in sandbox tests

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address gator re-review findings on workspace authorization

- Fix TUI unconditionally setting providers_v2_enabled after provider
  refresh; read the actual gateway setting via GetGatewayConfig at
  startup instead
- Fix SQLite json_extract with dotted label keys (e.g. example.com/env)
  by quoting the key in the JSON path
- Add authed_request wrappers to upstream OCI identity tests that were
  missing a principal after rebase
- Add test proving GetGatewayConfig is accessible without Platform Admin
- Add test for dotted/prefixed Kubernetes-style label key filtering

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address second gator re-review findings

- Loosen GetGatewayConfig from platform_admin to scope-only so workspace
  users can discover providers_v2_enabled during sandbox creation with
  inferred-provider commands; update proto descriptor, descriptor
  validation, and RFC 0011 access table
- Add validate_label_selector to handle_list_workspaces and escape
  single quotes in SQLite json_extract interpolation (CWE-89
  defense-in-depth)
- Re-fetch providers_v2_enabled after TUI gateway switch so the new
  gateway's capability is reflected
- Add e2e test for workspace user with inferred-provider command
- Add persistence test for adversarial label keys with SQL injection
  attempts
- Add handler test for invalid label selector rejection in
  ListWorkspaces

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): address third gator review findings

- Cap label selector pairs at 64 (CWE-400) to bound SQLite dynamic SQL
- Add SCOPE_ONLY_METHODS allowlist for scope-without-role RPCs (CWE-863)
- Normalize ID-based data-plane handlers to return NOT_FOUND for
  unauthorized sandboxes, closing the cross-workspace oracle (CWE-203)
- Fix TUI provider profile cache lookup key mismatch for legacy
  providers with empty profile_workspace
- Add whoami to CLI skill reference command tree
- Update TUI skill doc with workspace, provider, and settings coverage
- Document scope/workspace orthogonality on GetGatewayConfig proto

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): extend CWE-203 normalization to policy.rs sandbox handlers

GetSandboxConfig and GetSandboxLogs in policy.rs had the same
fetch-before-authorize pattern that leaked cross-workspace sandbox
existence. Promote fetch_and_authorize_sandbox to pub(super) and
use it from both sandbox.rs and policy.rs handlers.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(auth): update assertions for CWE-203 sandbox ID normalization

Cross-workspace sandbox access via ID-based handlers now returns
NOT_FOUND instead of PERMISSION_DENIED to prevent existence inference.
Update the unit test and OIDC e2e assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(auth): narrow CWE-203 error mapping and correct whoami output formats

Only remap PERMISSION_DENIED to NOT_FOUND in fetch_and_authorize_sandbox
and RevokeSshSession, letting INTERNAL and UNAUTHENTICATED propagate
as-is. Fix whoami --output format values in cli-reference.md to match
the actual CLI (table/json/yaml, not text/json).

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(ci): share network namespace with Keycloak in containerized CI

In GitHub Actions job containers, Docker port publishing lands on the
host, not inside the job container. Detect this environment and attach
Keycloak to the job container's network namespace instead, with
hardened defaults (cap-drop ALL, no-new-privileges, loopback-only
listener).

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-30 00:31:32 +00:00
Derek Carr 5952a5a23f feat(workspace): add workspace resource model with scoping, membershi… (#2243)
* feat(workspace): implement workspace model (Phase 1 of RFC 0011)

Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.

Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(cli): delegate sandbox upload command to existing upload function

The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): shorten sandbox names and fix test compatibility

Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).

Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(workspace): add test coverage for workspace CRUD and persistence isolation

Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(examples): update examples for workspace model compatibility

Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(sdk): add workspace-scoped client and workspace CRUD

Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings in workspace test assertions

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(docs): convert indented code blocks to fenced in RFC 0011

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings and apply cargo fmt across workspace

Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): address workspace scoping issues from review

- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
  to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
  Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
  returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog workspace-aware

Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.

Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(persistence): include workspace column in atomic policy revision INSERT

put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proxy): skip ancestor walk when socket owner is the entrypoint

collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog scope-aware

Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align podman e2e labels with centralized driver constants

The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align python profile isolation test with scope-aware catalog

Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): honor profile_workspace in runtime profile resolution

Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-07-21 00:59:18 +00:00
Russell Bryant 9377e0d5fe fix(providers): allow git clone/fetch via default GitHub provider (#2317)
* fix(providers): allow git clone/fetch via default GitHub provider

The github.com:443 git-transport endpoint used the read-only access
preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires
a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied
those operations and `gh repo clone` / `git clone https://...` failed.

Replace the preset with explicit rules that permit the read-only methods
plus POST */git-upload-pack, so clone/fetch work while push
(git-receive-pack) stays blocked. Enabling push still requires an
explicit policy proposal.

Why allowing this POST is still read-only: in git's smart HTTP protocol
POST is an RPC transport, not a write. A clone/fetch does GET
*/info/refs (ref discovery) followed by POST */git-upload-pack, whose
body is only the client's want/have negotiation; the server responds
with a packfile and nothing on the server is modified (data flows
server -> client). The service names are from the server's perspective:
git-upload-pack = the server uploads a pack to the client (a read/
download), while git-receive-pack = the server receives a pack from the
client (the actual write/push). The new rule is scoped to
*/git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs
to github.com remain denied.

Add a provider-profile regression test and a rego enforcement test
covering ref discovery, upload-pack (allowed), and receive-pack (denied).

Closes #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): strengthen git-transport regression and add clone e2e

Pin the exact allowed rule set for the built-in github git-transport
endpoint in both the provider-profile and composed-policy tests, so a
broader or additional POST rule (e.g. POST **) that could enable push
via git-receive-pack fails the test instead of passing a substring
check. Add an e2e test that attaches the built-in github provider and
clones a public repo over HTTPS, exercising provider attachment,
effective-policy composition, TLS interception, and real git behavior.

Update the Providers V2 docs so the github.com git-transport endpoint
shows explicit clone/fetch rules instead of the stale read-only preset.

Refs #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): isolate providers_v2 mutation in clone e2e

The clone e2e enables the gateway-global providers_v2_enabled setting.
Restore its exact prior value (or absence) captured via GetGatewayConfig
instead of unconditionally deleting it, and serialize the mutation
across xdist workers with an exclusive file lock on the run's shared
base temp dir, so a shared or pre-configured gateway is left untouched
and parallel workers cannot race the read-modify-restore.

Refs #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): serialize providers_v2 mutation with a suite-wide guard

The clone e2e's per-fixture lock only coordinated fixtures that acquired
it; other xdist workers hit the same gateway without it and could
observe the transiently-enabled providers_v2_enabled global during their
own sandbox creation (CWE-362).

Add an autouse readers-writer guard in conftest: every test holds a
shared lock on the gateway config, and a test marked
exclusive_gateway_config holds an exclusive lock. Mark the clone test
exclusive so no other worker is mid-test while it enables and restores
the gateway-global setting. Exact prior-value restoration is retained.

Refs #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-07-20 21:07:17 +00:00
emonq a2cd5f8eda fix(gateway): honor tty flag for interactive exec (#2315)
* fix(gateway): honor tty flag for interactive exec

Pass the requested TTY mode through the interactive SSH relay.
Skip PTY allocation and resize forwarding when TTY is disabled,
and add regression coverage for both modes.

Signed-off-by: emonq <emonq@outlook.com>

* test(gateway): improve `test_sandbox_interactive_exec_honors_tty` to test streamed stdin and stdout/stderr

Signed-off-by: emonq <emonq@outlook.com>

---------

Signed-off-by: emonq <emonq@outlook.com>
2026-07-20 17:47:37 +00:00
Kyle ZhengandJohn Myers ff9af8e320 fix(sandbox): acknowledge initial policy revision; expose SDK labels/selectors (#2170)
* fix(sandbox): acknowledge initial policy revision

The supervisor loaded and enforced a sandbox-scoped policy but never told
the gateway which revision it loaded. The policy poll loop seeded itself
with the initial revision's hash on its first poll, so `policy_changed`
was never true for that revision and `ReportPolicyStatus(LOADED)` — which
only ran in the hot-reload branch — was never called. The revision stayed
`Pending` and `current_policy_version` stayed 0 even though the sandbox was
`Ready` and the policy was effective. This was most visible with sparse
policies that get baseline-enriched into a new revision during startup.

After the OPA engine is constructed, report the exact sandbox revision the
supervisor loaded as LOADED, and seed the poll loop from that revision so
it is not re-reported. Report FAILED with the original construction error
if engine construction or conversion fails. Only sandbox-sourced revisions
(version > 0) whose canonical content matches the loaded policy are
acknowledged; global and local-file policies are untouched. Delivery uses
the shared bounded retry, is non-fatal on transient failure, and a pending
initial acknowledgement is delivered before any newer revision so policy
history is never reordered.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* feat(python): expose sandbox labels and selectors

The gateway protobuf and CLI already support request-level sandbox labels
(`CreateSandboxRequest.name`/`labels`) and selector-based listing
(`ListSandboxesRequest.label_selector`), but the public Python SDK dropped
them, so Python-created sandboxes could not be found via
`openshell sandbox list --selector ...`.

Add optional, source-compatible `name`/`labels` to `SandboxClient.create`,
`create_session`, and the high-level `Sandbox`, and `label_selector` to
`list`/`list_ids`. `SandboxRef` now carries the gateway labels as an
immutable mapping (default empty, so `SandboxRef(id, name, status)` still
works). Caller-provided label mappings are copied. Attaching the high-level
`Sandbox` to an existing sandbox rejects `name`/`labels` since creation
metadata cannot change on attach. Template labels remain a separate concept.
No protobuf changes are required.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(python): keep SandboxRef hashable and copy high-level labels

Excluding the new immutable `labels` field from SandboxRef equality/hash
(`compare=False`) preserves the original (id, name, status) identity and keeps
the frozen dataclass hashable — a MappingProxyType field would otherwise make
`hash(SandboxRef(...))` raise. Also defensively copy caller-provided labels in
the high-level `Sandbox` so later caller mutation cannot change what is sent.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): bound initial-policy-ack retries

The poll loop retried a pending initial acknowledgement before processing any
newer revision, but retried unconditionally forever. A permanently
undeliverable ack (e.g. the revision was superseded before it could be
reported) would then stall all later policy hot-reloads and provider-env
refreshes. Cap the retries; after the bound, give up and resume normal polling
so the loop cannot livelock on a stuck acknowledgement.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* test(sandbox): add sparse-policy revision-2 acknowledgement e2e

Regression for #2159: create a sandbox with the network-only policy-advisor
fixture, which the supervisor enriches with baseline filesystem paths during
startup (creating revision 2, superseding revision 1). Assert the effective
policy reaches revision 2 and no revision remains Pending once the supervisor
acknowledges the load. Adds SandboxGuard::create_keep_with_args to create a
kept sandbox with an initial --policy.

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(ci): correct sandbox checks

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* test(cli): serialize mTLS environment access

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): address policy review feedback

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): preserve exact policy acknowledgements

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>

* fix(sandbox): preserve local policy overrides

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: Kyle Zheng <kyzheng@nvidia.com>
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-07-08 17:24:26 -07:00
Evan Lezar f23c2c8e84 test(e2e): remove python gpu smoke test (#1948)
* fix(helm): build chart dependencies before lint

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* test(e2e): remove python gpu smoke test

Remove the Python GPU smoke test and its fixture. The e2e:k3s:gpu task only depended on e2e:python:gpu and did not have a separate k3s implementation, so remove that stale alias with the task it pointed at.

Signed-off-by: Evan Lezar <elezar@nvidia.com>
(cherry picked from commit 221a10378e188656c710560740cbc9463c002db6)

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-06-17 17:21:49 -05:00
Mesut Oezdil ec197a43ef fix(e2e): correct return type of _stub_with_token (#1897) 2026-06-13 14:20:59 -07:00
Drew Newberry 530aaf1360 feat(drivers): support docker and podman config mounts (#1785)
* feat(drivers): support docker and podman config mounts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(drivers): trim mount docs

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): cover local driver volume mounts

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(podman): satisfy linux clippy lint

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* feat(drivers): gate bind mounts behind gateway config

* docs(sandbox): simplify mount examples

* cleanup

* test(e2e): stabilize branch checks

* fix(drivers): tighten local mount validation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
2026-06-10 11:46:20 -07:00
Taylor Mutch 7a0c444445 refactor!(auth): drop SSH handshake secret (#1274)
* refactor!(auth): drop SSH handshake secret in favor of mTLS

The OPENSHELL_SSH_HANDSHAKE_SECRET / x-sandbox-secret mechanism was
misnamed: it does not authenticate SSH (which flows over the
RelayStream gRPC RPC and is gated by mTLS plus supervisor Unix-socket
permissions). It only gated a small set of sandbox-to-gateway
control-plane RPCs, and production deployments already enforce mTLS
on that channel — so the shared secret was redundant.

Replace the secret check with an mTLS-presence marker. Sandbox-class
methods (ReportPolicyStatus, PushSandboxLogs,
GetSandboxProviderEnvironment, SubmitPolicyAnalysis, GetSandboxConfig,
GetInferenceBundle) accept callers without a Bearer token; the gRPC
mTLS handshake is the trust boundary. Dual-auth methods treat
Bearer-present as full-scope CLI access and Bearer-absent as
sandbox-restricted scope via validate_sandbox_caller_update.

Drops the secret from all drivers (K8s, Podman, VM), the sandbox gRPC
interceptor, the Helm chart (values + pre-install hook + StatefulSet
env), the RPM bootstrap script, the man pages, and the
debug-openshell-cluster skill. Also removes the never-read
ssh_handshake_skew_secs flag and config field.

BREAKING CHANGE: --ssh-handshake-secret / OPENSHELL_SSH_HANDSHAKE_SECRET
and --ssh-handshake-skew-secs / OPENSHELL_SSH_HANDSHAKE_SKEW_SECS are
removed from the gateway, sandbox, and all driver binaries. The
openshell-ssh-handshake K8s Secret is no longer managed by the chart;
operators may delete the orphan. Deployments using
--disable-gateway-auth must enforce caller authentication at the
fronting proxy, since the gateway no longer validates a per-request
secret on sandbox-class methods.

Refs OS-174.

* docs(auth): scrub residual SSH handshake secret references

Sweep across docs, e2e scripts, the Podman driver README/NETWORKING
notes, the RPM/Helm/setup guides, the gateway man page, and RFC 0003
to remove instructions and examples that still referenced
OPENSHELL_SSH_HANDSHAKE_SECRET / --ssh-handshake-secret /
ssh_handshake_skew_secs. The mechanism is gone; nothing should still
suggest setting it. Negative-assertion regression tests are kept so
the env var cannot silently be re-introduced.

* test(drivers): drop SSH handshake secret negative-assertion tests

The supporting code, env vars, CLI flags, and config plumbing are
gone — these tests asserted absence of strings that no longer have
any path to being set. Remove the guards from the Docker, Podman,
Kubernetes, and VM driver test modules.
2026-05-14 13:14:30 -07:00
John T. Myers 1d3b741ee3 feat(providers): support sandbox provider attach lifecycle (#1242)
* feat(providers): support sandbox provider attach lifecycle

Closes #1171

Adds sandbox provider list, attach, and detach API/CLI support while keeping provider policy and credential resolution derived from current sandbox attachments.

* fix(providers): refresh sandbox provider credentials

Adds provider environment revisions and generation-scoped sandbox credential snapshots so future SSH and exec launches pick up provider attach, detach, and credential updates without mutating already-running processes.

Also blocks provider deletion while attached to prevent stale sandbox provider references.

* fix(providers): serialize sandbox object mutations

* test(providers): cover sandbox provider attach lifecycle

* test(providers): accept versioned credential placeholders
2026-05-08 11:14:44 -07:00
Drew Newberry e4b4e923ae test(e2e): run suites against docker gateway (#1153) 2026-05-05 13:28:08 -07:00
Drew Newberry a255ad9142 fix(e2e): stabilize wildcard host DNS test (#1144) 2026-05-04 07:59:56 -07:00
Mrunal Patel 084505425b feat(auth): add OIDC/Keycloak authentication with RBAC and scope-based permissions (#935)
* feat(auth): add OIDC/Keycloak authentication with RBAC

Add OAuth2/OIDC authentication to the gateway server with role-based
access control, CLI login flows, and full deployment plumbing.

Server: JWT validation against configurable OIDC issuer (oidc.rs),
JWKS key caching with TTL and rotation handling, method classification
(unauthenticated/sandbox-secret/dual-auth/bearer), identity extraction
with provider-agnostic Identity type, and RBAC enforcement via
AuthzPolicy with configurable admin/user roles and auth-only mode.

CLI: browser-based Authorization Code + PKCE flow, Client Credentials
flow for CI/automation, token storage with refresh, gateway add/login/
logout commands, OIDC bearer token injection over mTLS transport,
discovery endpoint for auto-configuration.

Security: sandbox-secret scope restriction on UpdateConfig (policy
sync only), anti-spoofing header stripping, dual-auth fallthrough
from sandbox-secret to Bearer token.

Deployment: OIDC config wired through DeployOptions, Docker env vars,
Helm values/templates, HelmChart manifest, cluster-entrypoint.sh, and
bootstrap scripts. Keycloak dev server script with pre-configured
realm (test users, roles, PKCE client, CI client).

Tested with Keycloak. The roles claim path and role names are
configurable to support other OIDC providers.

* feat(auth): add OAuth2 scope-based fine-grained permissions

Add opt-in scope enforcement on top of existing OIDC role-based access
control. When --oidc-scopes-claim is set, the server extracts scopes
from the JWT and checks them per-method against an exhaustive scope map.

Scopes: sandbox:read, sandbox:write, provider:read, provider:write,
config:read, config:write, inference:read, inference:write, and
openshell:all (wildcard). Methods not in the scope map require
openshell:all. Scopes layer on top of roles and cannot escalate
privilege. Auth-only mode (empty role names) still enforces scopes
when enabled.

Server: scopes_claim in OidcConfig, scope extraction from JWT
(space-delimited and JSON array formats), standard OIDC scope
filtering, scope check in AuthzPolicy after role check.

CLI: --oidc-scopes on gateway add/start stored in metadata and
consumed by gateway login, --oidc-scopes-claim on gateway start
forwarded to server, scopes parameter in browser and client
credentials OAuth2 flows with openid deduplication.

Deployment: oidc_scopes_claim wired through DeployOptions, docker.rs,
Helm, bootstrap scripts, and cluster entrypoint.

Keycloak: realm config updated with built-in OIDC scopes and 9
OpenShell client scopes as optional on openshell-cli and openshell:all
as default on openshell-ci.

* fix(auth): address branch review findings

Add GetInferenceBundle to sandbox-secret methods so sandbox inference
route refresh works under OIDC. Make GetSandboxConfig dual-auth so CLI
users can read sandbox settings with Bearer tokens.

Preserve OIDC gateway metadata on restart — a bare gateway start
without --oidc-* flags no longer erases the stored OIDC registration.

Document CI client ID requirement (openshell-ci vs openshell-cli) in
the testing guide. Add security note about auth-only mode blast radius
for GitHub Actions.

* fix(auth): complete review findings for OIDC auth boundary

Move OpenShell/GetSandboxConfig from sandbox-secret-only to dual-auth
so CLI users can read sandbox settings with Bearer tokens while sandbox
supervisors continue using the shared secret.

Add sandbox secret interceptor to the inference bundle fetch path so
GetInferenceBundle works under OIDC-enabled gateways. Extract shared
interceptor constructor to avoid duplication.

Add GetSandboxConfig to the config:read scope map so scope enforcement
applies consistently when scopes are enabled.

Refactor OIDC metadata preservation into apply_oidc_gateway_metadata()
with explicit resume semantics — only preserve existing OIDC metadata
on real resume paths, not on fresh deployments.

Update architecture docs and testing guide to reflect the corrected
method classifications and add new test coverage for interceptor
injection, scope requirements, metadata preservation, and dual-auth
classification.

* refactor(auth): use oauth2 crate for CLI OIDC flows

Replace hand-written PKCE generation, authorization URL construction,
token exchange, client credentials, and token refresh with the oauth2
crate's typed API.

Eliminates sha2, hex, and getrandom dependencies from the CLI. The
custom urlencoded() helper and manual form POST logic are replaced by
BasicClient methods with proper type-state safety.

Discovery and the callback server remain custom since the oauth2 crate
does not provide OIDC discovery or a localhost redirect listener.

* refactor(auth): move server auth modules into auth/ directory

Group oidc.rs, authz.rs, identity.rs, and the auth HTTP endpoints
under src/auth/ module directory. No behavioral changes.

  auth/mod.rs      — module root, re-exports HTTP router
  auth/oidc.rs     — JWT validation, JWKS caching, method classification
  auth/authz.rs    — role and scope authorization policy
  auth/identity.rs — provider-agnostic Identity type
  auth/http.rs     — /auth/connect and /auth/oidc-config endpoints

* fix(auth): use RequestBody auth type for client credentials flow

The oauth2 crate defaults to BasicAuth (HTTP Basic header) but Keycloak
and most OIDC providers expect client_secret_post (credentials in the
request body). Set AuthType::RequestBody explicitly to match the
pre-refactor behavior.

Also re-export Identity, IdentityProvider, and JwksCache from the auth
module so ServerState's public API remains nameable by external consumers.

* fix(auth): forward OPENSHELL_OIDC_SCOPES through cluster bootstrap

Pass --oidc-scopes to gateway start so the metadata includes requested
scopes after cluster bootstrap. Without this, users had to manually
edit metadata.json to set scopes for gateway login.

Usage: OPENSHELL_OIDC_SCOPES="openshell:all" mise run cluster

* test(auth): add OIDC e2e tests for RBAC, scopes, and client credentials

Add 10 end-to-end tests covering OIDC authentication against a live
K3s cluster with Keycloak:

RBAC (5 tests): admin can create providers, user cannot, user can list
sandboxes, unauthenticated requests rejected, health probe works
without auth.

Scopes (4 tests): sandbox-scoped token can list sandboxes but not
providers, openshell:all grants full access, no-scopes token denied.

Client credentials (1 test): CI token via client_credentials grant.

Tests are opt-in via OPENSHELL_E2E_OIDC=1 and OPENSHELL_E2E_OIDC_SCOPES=1
env vars. They derive the Keycloak URL from gateway metadata to match
the server's configured issuer.

Run with:

  OPENSHELL_E2E_OIDC=1 OPENSHELL_E2E_OIDC_SCOPES=1 \
  PYTHONPATH=python uv run pytest e2e/python/oidc/ -v

* fix(docs): fix markdown lint errors in OIDC architecture docs

Add blank lines before lists and fenced code blocks to satisfy
markdownlint MD031 and MD032 rules.
2026-04-30 10:37:23 -07:00
John T. Myers d414e69a20 refactor(server): unify policy persistence in objects table (#972)
* refactor(server): unify policy persistence in objects table

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(server): clean sandbox-owned records on reconcile delete

* refactor(server): use protos for stored policy records

* refactor(server): move policy persistence into policy_store

* fix(server): restore compute runtime merge compatibility

* fix(server): validate draft chunk sandbox ownership

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-04-28 16:15:14 -07:00
Derek Carr 5e28ea3a4b feat(server): add object meta convention to top-level objects (#919)
- adds filterable label selectors on resources

Closes #864

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-04-27 07:16:07 -07:00
Derek Carr 87f50f5e5d fix(e2e): add /dev/urandom to provider test sandbox policy (#948)
Python runtime requires /dev/urandom access during initialization
to seed the hash randomizer. The _default_policy() in provider tests
was missing this path, causing exec_python tests to fail with:
'Fatal Python error: _Py_HashRandomization_Init: failed to get
random numbers to initialize Python'.

Add /dev/urandom to read_only paths to match the policy used in
test_sandbox_policy.py, allowing Python to initialize successfully.

Signed-off-by: Derek Carr <decarr@redhat.com>
2026-04-24 07:51:04 -07:00
John T. Myers 29a3b1cacd fix(sandbox): two-phase Landlock to fix privilege ordering and add enforcement tests (#810)
* fix(sandbox): add parent-side Landlock availability probe logging

Landlock status was only logged from inside the pre_exec child process
where the tracing/OCSF pipeline is non-functional after fork. This made
Landlock failures completely invisible in sandbox logs.

Add a probe_availability() function that issues the raw
landlock_create_ruleset syscall to check kernel support, and call it
from the parent process before fork in all three spawn paths (entrypoint,
SSH PTY, SSH pipe). Uses std::sync::Once to emit exactly once per
sandbox lifetime.

WIP - addresses logging gap from #803.

* test(e2e): add Landlock filesystem enforcement tests

Verify Landlock availability logging and enforcement in e2e:
- OCSF probe event appears in sandbox logs
- Read-only paths block writes, allow reads
- Read-write paths allow both
- Paths outside policy are denied entirely
- User-owned paths outside policy are still blocked (proves
  Landlock enforces independently of Unix DAC permissions)

Requires Linux host with Landlock support (GitHub Actions runners,
Docker Desktop linuxkit). Related to #803.

* test(e2e): add xfail tests proving #803 privilege ordering bug

Two strict xfail tests that demonstrate the root cause of #803:
- PathFd::new() runs as uid 998 after drop_privileges, so root-only
  paths (mode 700) silently fail and Landlock degrades
- When ALL paths fail, best_effort silently drops Landlock entirely

These tests will pass after the two-phase Landlock fix (open PathFds
as root before drop_privileges, restrict_self after).

* fix(test): fix Landlock e2e tests based on test run results

- Remove log-reading tests: the Landlock probe logs to the
  supervisor's container stdout, not the in-sandbox file appender
  at /var/log/openshell*.log*
- Replace broken xfail test (checked for child-process log messages
  that never reach the file appender) with a stat()-based test that
  verifies /root is actually in the Landlock allowlist
- Keep enforcement tests (all passing) and the root-only policy
  xfail test (correctly proves #803 bug)

* fix(test): remove mixed-policy xfail test

Landlock doesn't restrict stat(), so the test passed unexpectedly.
In the mixed policy case where /root is silently skipped, Landlock
still applies (using other paths) and blocks /root even harder
(not in allowlist = denied). The observable security degradation
only occurs when ALL paths fail, which the existing xfail test
already covers.

* fix(sandbox): two-phase Landlock to fix privilege ordering (#803)

Split Landlock apply into prepare() and enforce():
- prepare() runs as root before drop_privileges: opens PathFds,
  creates ruleset, adds rules. Root-only paths (mode 700) now
  succeed instead of silently failing as uid 998.
- enforce() runs after drop_privileges: calls restrict_self()
  which does not require root.

This fixes the root cause of #803 where drop_privileges() ran
before sandbox::apply(), causing PathFd::new() to fail on
root-only paths. In best_effort mode this silently dropped all
Landlock restrictions.

The fix applies to all three spawn paths: entrypoint (process.rs),
SSH PTY (ssh.rs), and SSH pipe exec (ssh.rs).

Removes xfail marker from e2e test that now passes.

* fix(sandbox): address PR review feedback

- enforce() now respects best_effort: if restrict_self() fails and
  policy is best_effort, log and degrade instead of aborting startup
- log_sandbox_readiness distinguishes best_effort (degraded) from
  hard_requirement (will fail) in OCSF messages
2026-04-13 10:55:34 -07:00
John T. Myers 2ca553a4a0 fix(sandbox): validate always-blocked IPs at load time, enrich denial logs, and filter un-fixable proposals (#814) (#815)
Policies with allowed_ips entries targeting loopback, link-local, or
unspecified ranges now fail at connection time instead of being silently
blocked at runtime. The shorthand log format for DENIED events includes
a [reason:...] suffix so operators can distinguish 'allowlist miss' from
'structurally un-allowable'. The mechanistic mapper skips proposals for
always-blocked destinations, preventing the infinite TUI notification
loop. The gateway validates proposed rules on approval as defense-in-depth.

- Extract shared IP helpers (is_always_blocked_ip, is_always_blocked_net,
  is_internal_ip) to openshell_core::net
- Reject always-blocked entries in parse_allowed_ips with hard error
- Skip implicit allowed_ips synthesis for always-blocked literal IP hosts
- Add status_detail to HttpActivityBuilder for denial reason propagation
- Enrich NET and HTTP shorthand with [reason:...] for DENIED events
- Add engine: tag to HTTP shorthand (consistency with NET shorthand)
- Filter always-blocked proposals in mechanistic mapper generate_proposals
- Add validate_rule_not_always_blocked server-side defense-in-depth
- Update architecture docs, published docs, and E2E test assertions
2026-04-13 09:47:50 -07:00
John T. Myers b7779bdefa feat(sandbox): integrate OCSF structured logging for sandbox events (#720)
* feat(sandbox): integrate OCSF structured logging for all sandbox events

WIP: Replace ad-hoc tracing calls with OCSF event builders across all
sandbox subsystems (network, SSH, process, filesystem, config, lifecycle).

- Register ocsf_logging_enabled setting (defaults false)
- Replace stdout/file fmt layers with OcsfShorthandLayer
- Add conditional OcsfJsonlLayer for /var/log/openshell-ocsf.log
- Update LogPushLayer to extract OCSF shorthand for gRPC push
- Migrate ~106 log sites to OCSF builders (NetworkActivity, HttpActivity,
  SshActivity, ProcessActivity, DetectionFinding, ConfigStateChange,
  AppLifecycle)
- Add openshell-ocsf to all Docker build contexts

* fix(scripts): attach provider to all smoke test phases to avoid rate limits

GitHub's unauthenticated API rate limit (60/hour) causes flaky 403s for
Phases 1, 2, and 4. Fix by attaching the provider to all sandboxes and
upgrading the Phase 1 policy to L7 so credential injection works.

Phase 4 (tls:skip) cannot inject credentials by design, so relax the
assertion to accept either 200 or 403 from upstream -- both prove the
proxy forwarded the request.

* fix(ocsf): remove timestamp from shorthand format to avoid double-timestamp

The display layer (gateway logs, TUI, sandbox logs CLI) already prepends
a timestamp. Having one in the shorthand output too produces redundant
double-timestamps like:

  15:49:11 sandbox INFO  15:49:11.649 I NET:OPEN ALLOWED ...

Now the shorthand is just the severity + structured content:

  15:49:11 sandbox INFO  I NET:OPEN ALLOWED ...

* refactor(ocsf): replace single-char severity with bracketed labels

Replace cryptic single-character severity codes (I/L/M/H/C/F) with
readable bracketed labels: [LOW], [MED], [HIGH], [CRIT], [FATAL].

Informational severity (the happy-path default) is omitted entirely to
keep normal log output clean and avoid redundancy with the tracing-level
INFO that the display layer already provides.

Before: sandbox INFO  I NET:OPEN ALLOWED ...
After:  sandbox INFO  NET:OPEN ALLOWED ...

Before: sandbox INFO  M NET:OPEN DENIED ...
After:  sandbox INFO  [MED] NET:OPEN DENIED ...

* feat(sandbox): use OCSF level label for structured events in log push

Set the level field to 'OCSF' instead of 'INFO' for OCSF events in the
gRPC log push. This visually distinguishes structured OCSF events from
plain tracing output in the TUI and CLI sandbox logs:

  sandbox OCSF  NET:OPEN [INFO] ALLOWED python3(42) -> api.example.com:443
  sandbox OCSF  NET:OPEN [MED] DENIED python3(42) -> blocked.com:443
  sandbox INFO  Fetching sandbox policy via gRPC

* fix(sandbox): convert new Landlock path-skip warning to OCSF

PR #677 added a warn!() for inaccessible Landlock paths in best-effort
mode. Convert to ConfigStateChangeBuilder with degraded state so it
flows through the OCSF shorthand format consistently.

* fix(sandbox): use rolling appender for OCSF JSONL file

Match the main openshell.log rotation mechanics (daily, 3 files max)
instead of a single unbounded append-only file. Prevents disk exhaustion
when ocsf_logging_enabled is left on in long-running sandboxes.

* fix(sandbox): address reviewer warnings for OCSF integration

W1: Remove redundant 'OCSF' prefix from shorthand file layer — the
    class name (NET:OPEN, HTTP:GET) already identifies structured events
    and the LogPushLayer separately sets the level field.

W2: Log a debug message when OCSF_CTX.set() is called a second time
    instead of silently discarding via let _.

W3: Document the boundary between OCSF-migrated events and intentionally
    plain tracing calls (DEBUG/TRACE, transient, internal plumbing).

W4: Migrate remaining iptables LOG rule failure warnings in netns.rs
    (IPv4 TCP/UDP, IPv6 TCP/UDP) to ConfigStateChangeBuilder for
    consistency with the IPv4 bypass rule failure already migrated.

W5: Migrate malformed inference request warn to NetworkActivity with
    ActivityId::Refuse and SeverityId::Medium.

W6: Use Medium severity for L7 deny decisions (both CONNECT tunnel and
    FORWARD proxy paths) to match the CONNECT deny severity pattern.
    Allows and audits remain Informational.

* refactor(sandbox): rename ocsf_logging_enabled to ocsf_json_enabled

The shorthand logs are already OCSF-structured events. The setting
specifically controls the JSONL file export, so the name should reflect
that: ocsf_json_enabled.

* fix(ocsf): add timestamps to shorthand file layer output

The OcsfShorthandLayer writes directly to the log file with no outer
display layer to supply timestamps. Add a UTC timestamp prefix to every
line so the file output matches what tracing::fmt used to provide.

Before: CONFIG:VALIDATED [INFO] Validated 'sandbox' user exists in image
After:  2026-04-01T15:49:11.649Z CONFIG:VALIDATED [INFO] Validated ...

* fix(docker): touch openshell-ocsf source to invalidate cargo cache

The supervisor-workspace stage touches sandbox and core sources to force
recompilation over the rust-deps dummy stubs, but openshell-ocsf was
missing. This caused the Docker cargo cache to use stale ocsf objects
from the deps stage, preventing changes to the ocsf crate (like the
timestamp fix) from appearing in the final binary.

Also adds a shorthand layer test verifying timestamp output, and drafts
the observability docs section.

* fix(ocsf): add OCSF level prefix to file layer shorthand output

Without a level prefix, OCSF events in the log file have no visual
anchor at the position where standard tracing lines show INFO/WARN.
This makes scanning the file harder since the eye has nothing consistent
to lock onto after the timestamp.

Before: 2026-04-01T04:04:13.065Z CONFIG:DISCOVERY [INFO] ...
After:  2026-04-01T04:04:13.065Z OCSF CONFIG:DISCOVERY [INFO] ...

* fix(ocsf): clean up shorthand formatting for listen and SSH events

- Fix double space in NET:LISTEN, SSH:LISTEN, and other events where
  action is empty (e.g., 'NET:LISTEN [INFO]  10.200.0.1' -> 'NET:LISTEN [INFO] 10.200.0.1')
- Add listen address to SSH:LISTEN event (was empty)
- Downgrade SSH handshake intermediate steps (reading preface, verifying)
  from OCSF events to debug!() traces. Only the final verdict
  (accepted/denied) is an OCSF event now, reducing noise from 3 events
  to 1 per SSH connection.
- Apply same spacing fix to HTTP shorthand for consistency.

* docs(observability): update examples with OCSF prefix and formatting fixes

Align doc examples with the deployed output:
- Add OCSF level prefix to all shorthand examples in the log file
- Show mixed OCSF + standard tracing in the file format section
- Update listen events (no double space, SSH includes address)
- Show one SSH:OPEN per connection instead of three
- Update grep patterns to use 'OCSF NET:' etc.

* docs(agents): add OCSF logging guidance to AGENTS.md

Add a Sandbox Logging (OCSF) section to AGENTS.md so agents have
in-context guidance for deciding whether new log emissions should use
OCSF structured logging or plain tracing. Covers event class selection,
severity guidelines, builder API usage, dual-emit pattern for security
findings, and the no-secrets rule.

Also adds openshell-ocsf to the Architecture Overview table.

* fix: remove workflow files accidentally included during rebase

These files were already merged to main in separate PRs. They got
pulled into our branch during rebase conflict resolution for the
deleted docs-preview-pr.yml file.

* docs(observability): use sandbox connect instead of raw SSH

Users access sandboxes via 'openshell sandbox connect', not direct SSH.

* fix(docs): correct settings CLI syntax in OCSF JSON export page

The settings CLI requires --key and --value named flags, not positional
arguments. Also fix the per-sandbox form: the sandbox name is a
positional argument, not a --sandbox flag.

* fix(e2e): update log assertions for OCSF shorthand format

The E2E tests asserted on the old tracing::fmt key=value format
(action=allow, l7_decision=audit, FORWARD, L7_REQUEST, always-blocked).
Update to match the new OCSF shorthand (ALLOWED/DENIED, HTTP:, NET:,
engine:ssrf, policy:).

* feat(sandbox): convert WebSocket upgrade log calls to OCSF

PR #718 added two log calls for WebSocket upgrade handling:

- 101 Switching Protocols info → NetworkActivity with Upgrade activity.
  This is a significant state change (L7 enforcement drops to raw relay).

- Unsolicited 101 without client Upgrade header → DetectionFinding with
  High severity. A non-compliant upstream sending 101 without a client
  Upgrade request could be attempting to bypass L7 inspection.
2026-04-07 13:01:13 -07:00
John T. Myers 77e55ea989 test(e2e): replace flaky Python live policy update tests with Rust (#742)
Remove test_live_policy_update_and_logs and
test_live_policy_update_from_empty_network_policies from the Python e2e
suite. Both used a manual 90s poll loop against GetSandboxPolicyStatus
that flaked in CI with 'Policy v2 was not loaded within 90s'.

Add e2e/rust/tests/live_policy_update.rs with two replacement tests
that exercise the same policy lifecycle (version bumping, hash
idempotency, policy list history) through the CLI using the built-in
--wait flag for reliable synchronization.
2026-04-02 15:06:32 -07:00
Piotr Mlocek 1c659c1c12 fix(sandbox/bootstrap): GPU Landlock baseline paths and CDI spec missing diagnosis (#710)
* fix(sandbox): add GPU device nodes and nvidia-persistenced to landlock baseline

Landlock READ_FILE/WRITE_FILE restricts open(2) on character device files
even when DAC permissions would otherwise allow it. GPU sandboxes need
/dev/nvidiactl, /dev/nvidia-uvm, /dev/nvidia-uvm-tools, /dev/nvidia-modeset,
and per-GPU /dev/nvidiaX nodes in the policy to allow NVML initialization.

Additionally, CDI bind-mounts /run/nvidia-persistenced/socket into the
container. NVML tries to connect to this socket at init time; if the
directory is not in the landlock policy, it receives EACCES (not
ECONNREFUSED), which causes NVML to abort with NVML_ERROR_INSUFFICIENT_PERMISSIONS
even though nvidia-persistenced is optional.

Both classes of paths are auto-added to the baseline when /dev/nvidiactl is
present. Per-GPU device nodes are enumerated at runtime to handle multi-GPU
configurations.
2026-04-01 18:04:25 -07:00
Drew Newberry 151fca9dc5 fix(server): return already_exists for duplicate sandbox names (#695)
Check for existing sandbox name before persisting, matching the
provider-creation pattern. The CLI now surfaces a clear hint instead
of a raw UNIQUE constraint error.

Closes #691
2026-03-31 11:37:20 -07:00
John T. MyersandJohn Myers e8950e624c feat(sandbox): add L7 query parameter matchers (#617)
* feat(sandbox): add L7 query parameter matchers

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(sandbox): decode + as space in query params and validate glob syntax

Three improvements from PR #617 review:

1. Decode + as space in query string values per the
   application/x-www-form-urlencoded convention. This matches Python's
   urllib.parse, JavaScript's URLSearchParams, Go's url.ParseQuery, and
   most HTTP frameworks. Literal + should be sent as %2B.

2. Add glob pattern syntax validation (warnings) for query matchers.
   Checks for unclosed brackets and braces in glob/any patterns. These
   are warnings (not errors) because OPA's glob.match is forgiving,
   but they surface likely typos during policy loading.

3. Add missing test cases: empty query values, keys without values,
   unicode after percent-decoding, empty query strings, and literal +
   via %2B encoding.

* fix(sandbox): add missing query_params field in forward proxy L7 request info

* style(sandbox): fix formatting in proxy L7 query param parsing

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-03-30 13:53:12 -07:00
John T. Myers 256f7fc884 fix(sandbox,server): fix chunk merge duplicates and OPA variable collision with overlapping policies (#571)
* fix(sandbox,server): fix chunk merge duplicates and OPA variable collision with overlapping policies

Two related bugs triggered when a draft rule approval creates a second
policy entry for the same host:port:

1. merge_chunk_into_policy looked up existing rules by chunk.rule_name
   (auto-generated as allow_{host}_{port}), which never matched the
   user's original rule name.  Now scans all network_policies entries
   for a host:port endpoint match before falling back to insertion,
   and merges allowed_ips into the existing endpoint.

2. The Rego allow_request rule and _matching_endpoint_configs
   comprehension used 'some ep; ep := policy.endpoints[_]' which
   caused regorus to error with 'duplicated definition of local
   variable ep' when multiple policies covered the same host:port.
   Refactored to isolate endpoint iteration inside helper functions
   (_policy_allows_l7, _policy_endpoint_configs) so variables are
   scoped per-policy evaluation.

Refs: #567

* test(e2e): add overlapping policy tests and update FWD-2 for implicit allowed_ips

- Update FWD-2 (test_forward_proxy_denied_without_allowed_ips ->
  test_forward_proxy_allows_private_ip_host_without_allowed_ips):
  literal IP host no longer requires explicit allowed_ips, expects 200.

- Add OVL-1: overlapping L4 policies for same host:port must not crash
  OPA and should allow forward proxy connections.

- Add OVL-2: overlapping L7 policies for same host:port must not crash
  OPA and should allow CONNECT tunnel establishment.

Refs: #567

* style: apply cargo fmt formatting

* test(e2e): update SSRF-3 and SSRF-6 for implicit allowed_ips behavior

SSRF-6: Private IP with literal IP host now gets implicit allowed_ips
from PR #570, so CONNECT returns 200 instead of 403.

SSRF-3: Loopback is still blocked but via the always-blocked path
(implicit allowed_ips is synthesized, then resolve_and_check_allowed_ips
catches it). Log message says 'always-blocked' instead of 'internal
address'.

* fix(e2e): use negative assertion for SSRF-6 when nothing listens on target port

When the SSRF check passes but nothing listens on the target port,
recv() returns empty bytes. Use 'assert 403 not in' (matching SSRF-4
pattern) instead of 'assert 200 in'.

* fix(e2e): update provider tests for redacted credential values

PR #569 changed credential redaction from clearing the map to
replacing values with 'REDACTED'. Update e2e assertions to expect
credential keys with REDACTED values instead of an empty map.
2026-03-24 14:15:07 -07:00
Evan Lezar 1a9eea5351 feat(tasks): wire e2e:gpu to bootstrap cluster with GPU support (#547)
Pass CLUSTER_GPU=1 inline in e2e:python:gpu's depends so that the
cluster is bootstrapped with --gpu when GPU e2e tests are run.

Add --gpu flag handling to cluster-bootstrap.sh and default
OPENSHELL_E2E_GPU_IMAGE to an empty string so the server resolves
the default sandbox image when no override is provided.

Signed-off-by: Evan Lezar <elezar@nvidia.com>
2026-03-24 11:31:24 -07:00
John T. Myers 834f8aa184 fix: security hardening batch 1 (SEC-002 through SEC-010) (#548)
* fix(l7): reject ambiguous HTTP framing in REST proxy (SEC-009)

Harden the L7 REST proxy HTTP parser against request smuggling:

- Reject requests containing both Content-Length and Transfer-Encoding
  headers per RFC 7230 Section 3.3.3 (CL/TE ambiguity)
- Replace String::from_utf8_lossy with strict UTF-8 validation to
  prevent interpretation gaps with upstream servers
- Reject bare LF line endings (require CRLF per HTTP spec)
- Validate HTTP version string (HTTP/1.0 or HTTP/1.1 only)

* fix(server): harden shell_escape and command construction (SEC-002)

Harden the gRPC exec handler against command injection via the
structured-to-shell-string conversion:

- Reject null bytes and newlines/carriage returns in shell_escape()
- Add input validation in exec_sandbox: reject control characters in
  command args, env values, and workdir
- Enforce size limits: max 1024 args, 32 KiB per arg/value, 4 KiB
  workdir, 256 KiB total assembled command string
- Change shell_escape and build_remote_exec_command to return Result
  so callers must handle validation failures

* fix(server): add command validation at SSH transport boundary (SEC-003)

Add defense-in-depth validation in run_exec_with_russh before sending
the command to the sandbox SSH server:

- Reject null bytes in command string at transport boundary
- Enforce max command length (256 KiB) at transport boundary
- Enhance stream_exec_over_ssh logging with command length, stdin
  length, and truncated command preview for audit trail

* fix(sandbox): validate port range and extend loopback check in SSH (SEC-007)

Harden the SSH direct-tcpip channel handler:

- Validate port_to_connect <= 65535 before u32-to-u16 cast to prevent
  port truncation (e.g., 65537 becoming port 1)
- Replace string-literal loopback check with is_loopback_host() that
  covers the full 127.0.0.0/8 range, IPv4-mapped IPv6 (::ffff:127.x),
  bracketed IPv6, and case-insensitive localhost
- Remove #[allow(clippy::cast_possible_truncation)] since the cast is
  now proven safe by the preceding range check

* fix(sandbox): sanitize inference error messages returned to sandbox (SEC-008)

Replace verbatim internal error strings in router_error_to_http with
generic messages to prevent information leakage to sandboxed code.
Upstream URLs, internal hostnames, TLS details, and file paths are no
longer exposed. Full error context is still logged server-side at warn
level by the caller for debugging.

* fix(server): block internal IPs in SSH proxy target validation (SEC-006)

Add IP validation in start_single_use_ssh_proxy to prevent SSRF if a
sandbox status record were poisoned:

- Resolve DNS before connecting and validate the resolved IP
- Block loopback (127.0.0.0/8) and link-local (169.254.0.0/16, covers
  cloud metadata endpoint) addresses
- Block IPv4-mapped IPv6 variants of the same ranges
- Connect to the validated SocketAddr directly to prevent TOCTOU
- Add debug logging of resolved target IP for audit

* fix(sandbox): add resource limits to chunked body parser (SEC-010)

Harden parse_chunked_body in the inference interception path:

- Replace all unchecked +2 additions with checked_add for consistent
  overflow safety across all target architectures
- Add MAX_CHUNKED_BODY (10 MiB) to cap decoded body size
- Add MAX_CHUNK_COUNT (4096) to prevent CPU exhaustion via tiny chunks
- Early-reject chunk sizes larger than remaining buffer space

* fix(cli): double-escape command for SSH path, validate host and name (SEC-004)

Harden doctor_exec against command injection in the SSH remote path:

- Apply shell_escape to inner_cmd in the SSH path so it survives the
  double shell interpretation (SSH remote shell + sh -lc). This also
  fixes a correctness bug where multi-word commands were silently
  broken in the SSH path.
- Add validate_gateway_name to reject shell metacharacters in gateway
  names before use in container_name
- Add validate_ssh_host to reject metacharacters in remote_host loaded
  from metadata.json

* fix(sandbox): add CIDR breadth warning and control-plane port blocklist (SEC-005)

Defense-in-depth for the allowed_ips feature:

- Log a warning when a CIDR entry has a prefix length < /16, as overly
  broad ranges may unintentionally expose control-plane services
- Block K8s API (6443), etcd (2379/2380), and kubelet (10250/10255)
  ports unconditionally in resolve_and_check_allowed_ips, even when
  the resolved IP matches an allowed_ips entry

* test(e2e): update assertion for sanitized inference error message (SEC-008)

The SEC-008 fix changed the error message from 'no compatible route
for source protocol ...' to 'no compatible inference route available'.
Update the E2E assertion substring to match.
2026-03-23 14:41:03 -07:00
John T. Myers bbcaed2ea7 refactor(proto): rename UpdateSettings to UpdateConfig for consistency with read path (#515) 2026-03-20 16:45:59 -07:00
John T. Myers a831a8921b feat(settings): gateway-to-sandbox runtime settings channel (#474)
* feat(gateway/sandbox): add global and sandbox runtime settings flow
2026-03-20 14:08:57 -07:00