mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
pull-request/3947
22
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5c0c9e446e |
feat(providers): add OCI Generative AI example provider profile (#3904)
Add an example oci-genai inference profile for Oracle Cloud Infrastructure Generative AI through its OpenAI-compatible endpoint. The profile injects a compartment-scoped Generative AI API key as a bearer token only at the regional OCI inference hosts and only under /openai/v1 with GET, POST, and DELETE, so the sandbox never holds the key and the key cannot reach any other OCI surface. The header comments carry the OCI-side setup (create the IAM policy before the key, least-privilege statement, key creation and rotation), the operations verified through the sandbox proxy with a real key (chat completions with streaming, tool calling, vision input, embeddings, and the Responses API), the OCI error messages operators will meet, and the realm and signed-transport caveats. Scoped to the profile YAML per #3906; the only code change is the entry in the profile listing test, which enumerates providers/*.yaml. Signed-off-by: Federico Kamelhar <federico.kamelhar@oracle.com> |
||
|
|
73a181d32f |
docs: streamline README, add policy prover to architecture docs (#3718)
* docs(readme): streamline README and move reference detail to docs Restructure the README as a short path from overview to quickstart to further reading. Move prerelease install steps into the installation guide and telemetry build flags into a new observability page. Fix broken docs links and outdated runtime and credential descriptions. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): describe 0.1.0 as adding new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): add policy prover as a gateway component Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe OpenShell as a runtime for fleets of autonomous AI agents Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): describe policy prover as formal verification Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): name OpenShell Sandbox in component table Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): fold isolation backend into supervisor row Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(architecture): mention formal verification in overview Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): use the published OpenCode image in the first-agent guide Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(inference): correct provider examples and readiness Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(providers): correct Google binding and provider selection Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(readme): sharpen value prop, how it works, and explore further Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): use example Anthropic profile and add policy advisor step Import the example Anthropic profile, which now allows OpenCode, instead of editing it with sed. Add a step that shows how to review and approve mechanistic policy proposals as the agent needs more access. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(providers): add OpenRouter example for OpenCode Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(run-agent): run OpenCode against OpenRouter with a free model Add an example OpenRouter provider profile scoped to OpenCode and switch the first-agent guide to it, using a free Nemotron model so readers do not need OpenRouter credits. Revert the OpenCode binary added to the example Anthropic profile. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(readme): link first-agent guide and add agent skills section Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> Co-authored-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
854b2370b8 |
fix(policy): align quickstart, policy skills, and pypi profile with current behavior (#3695)
* fix(examples): add quickstart rule with policy update and show OCSF logs The quickstart applied policy.yaml with `openshell policy set`, which replaces the whole policy. The file omitted /bin from the restrictive default, so the live filesystem additivity check could reject it. Add the rule with `openshell policy update` instead, and keep policy.yaml as a complete policy for `sandbox create --policy` that covers the default read-only paths. The demo and README filtered logs with `--level warn`, but the server ranks OCSF events as INFO, which hid the policy decisions the demo shows. Query `--source sandbox` without a level filter and match the OCSF shorthand (DENIED/ALLOWED) instead of the retired key=value format. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): align generate-sandbox-policy with proxy behavior Remove the `protocol: sql` validation check and the SQL `command` matcher, which the published policy docs no longer describe. Correct the private IP guidance: exact user-declared hostnames may reach private addresses without allowed_ips. Wildcard, hostless, and advisor-proposed endpoints still need allowed_ips, and loopback, link-local, unspecified, and cloud metadata addresses stay blocked. Stop describing an omitted protocol as pure L4 or uninspected. The proxy still terminates TLS, parses HTTP strictly, and enforces request authority; it only skips method and path rules. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(providers): drop pip script paths from pypi profile binaries OpenShell identifies a process by /proc/<pid>/exe, so a pip script runs as its Python interpreter and the .venv/bin/pip entries could never match. The venv interpreters that run those scripts are already listed, so remove the script paths and explain in the header that users must list interpreters. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): fix openshell-cli policy iteration steps Monitoring denials with `--level warn` hides them, because the server ranks OCSF policy events as INFO. Drop the level filter and describe the OCSF shorthand DENIED lines instead of the retired `action: deny` format. `policy get --full > file` produced input that `policy set` cannot parse: the output starts with revision details before the `---` separator, and --full adds provider-composed rules. Export `--base` and keep only the YAML after the separator. Also stop recommending full replacement for filesystem, Landlock, or process changes, which require recreating the sandbox, and drop the SQL mention that generate-sandbox-policy no longer covers. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): qualify authority checks for omitted-protocol endpoints An endpoint without `protocol` only receives authority checks on HTTP requests the proxy parses after default TLS handling. Without an L7 route or required middleware, other CONNECT payloads such as HTTP/2 prior knowledge can use the raw relay, and `tls: skip` bypasses termination and parsing. Stop describing omitted-protocol endpoints as always authority-checked. Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
9244868056 |
docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe updated security architecture neutrally Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: highlight new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandboxes): clarify how to disconnect Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align architecture and guides with current navigation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(extensibility): streamline extension authentication guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
718dba3430 |
fix(policy)!: reject removed tls endpoint values (#3414)
Signed-off-by: Yuedong Wu <dwcn22@outlook.com> |
||
|
|
50230616d5 |
refactor(runtime): retire Community image dependencies (#3386)
* feat(sandbox): default to official Alpine sandbox image default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic version-qualified official image, so a fresh install no longer depends on the community sandbox image catalog. All compute drivers (docker, podman, kubernetes, vm) inherit this fallback. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(deploy): default deployment configs to the official Alpine sandbox image Update the shared gateway default_image, Helm chart values, the standalone Kubernetes manifest, and the dev gateway task scripts to use docker.io/library/alpine:3.22 instead of the community base image, consistent with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA needs a glibc base). Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * feat(driver): default to numeric non-root identity for USER-less images With the default sandbox image now Alpine, images that declare no OCI USER must start instead of being rejected. When the image declares no USER and the policy requests none, the Podman and Docker drivers now supply a numeric non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting, matching the numeric-identity behavior of the Kubernetes and VM drivers. The supervisor's resolved-identity path runs the sandbox as a synthesized non-root account without the account existing in the image. Images that declare a USER keep the OCI resolution path unchanged. Part of #3116. Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(conformance): use Alpine workload image Signed-off-by: Evan Lezar <elezar@nvidia.com> * refactor(policy): drop community image /app path from default policy The restrictive default policy granted read-only access to /app, a directory that only existed in the community base image. A generic Alpine default has no /app, so remove it. Landlock best-effort already ignores absent paths; this just stops advertising a community-specific layout in the default. Part of #3116. Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> * docs(config): document Alpine default images Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): report early sandbox termination Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(podman): initialize rootless workspace ownership Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(sandbox): qualify NVIDIA Ubuntu default Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): initialize rootful default workspace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sftp): add native sandbox adapter Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): gate runtime helper support to Linux Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): support standard OpenSSH file operations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sftp): harden rename and special file handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(runtime): remove community image dependencies Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): build provider readiness tool fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(e2e): use a dedicated Noble fixture for Docker tests Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Akram Signed-off-by: Akram <akram.benaissi@gmail.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
04146692d9 |
refactor(providers)!: make provider profiles import-only (#3383)
* chore(providers): remove dead provider plugin modules Twelve modules under crates/openshell-providers/src/providers/ were never declared in providers/mod.rs, so they have not been compiled since the plugin registry was narrowed to the two adapters it still registers. Four of them (generic, gitlab, opencode, outlook) key off provider type identifiers that normalize_provider_type already retired. Keep google_cloud and vertex, which are the only plugins ProviderRegistry::new registers. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(providers): load example profiles from providers/ at test time Add an example_profiles module that reads the YAML under providers/ from the source checkout at run time and parses it with the existing profile loader. It is gated behind a new non-default example-profiles feature so it is available to this crate's own tests and, once wired into dev-dependencies, to gateway and CLI tests, while never reaching a release binary. Nothing consumes it yet; later changes move the test fixtures off the compiled catalog and onto these files. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test: source profile fixtures from providers/ instead of the compiled catalog Unit and integration fixtures reached into builtin_profiles() to get a profile to work with, which ties the tests to the compiled catalog rather than to the files an operator would import. Point them at the example_profiles loader instead. The profiles.rs tests keep their coverage unchanged and become explicit golden tests over providers/*.yaml: they are what keeps those files valid once nothing compiles them. builtin_profiles_are_sorted_by_id widens into a test that the whole example set parses, sorts, and lints clean as one catalog. The CLI fake gateways serve the example profiles the way a real gateway serves what an operator imported, through a shared helper. No behavior change: the gateway still loads the built-in source by default and serves the same profiles. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document the example provider profiles Each file under providers/ now opens with a header naming its expected client binary identities, the image layout those paths assume, the credential scope, the endpoint access it grants, and a smoke test. Add a README covering the import commands and why a profile should be copied and edited rather than imported unchanged. Several of these profiles bind network access to paths that only exist in the OpenShell Community image — /sandbox/.venv, /app/.venv, /sandbox/.cursor-server, /usr/lib/node_modules. Imported unchanged into another image the profile matches nothing: the catalog still advertises it, but the credential is never injected and the traffic is denied. The headers say so where it applies. Comments only; the profile schema has no documentation fields and all fifteen files still lint clean. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(server): seed unit-test state with the example provider profiles Server unit tests inherited the builtin + user source default from ServerState, so around a hundred and forty assertions about github, openai and the rest resolved against the compiled catalog. Point the test state at the user source alone and import the example profiles from providers/ into its store first, the way an operator would. Every one of those assertions keeps passing unchanged, which is the point: it proves the gateway behaves identically with an imported catalog before the default moves. Four tests asserted builtin-source semantics specifically. Under import-only the only profiles a gateway cannot edit are the ones a non-user source vends, so they now exercise a source-managed profile composed with the user source; the read-only guard they cover is the one that still applies to interceptor catalogs. The list test asserts the imported profile's user/platform identity instead of builtin with an empty scope. test_server_state_with_user_only_github_profile is gone: the default test state now is a user-only gateway with github imported. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli): resolve credential suggestions from the gateway catalog The --env credential warning scanned a profile table compiled into the CLI, so its suggestions described the binary rather than the gateway the user is talking to: a profile the gateway does not serve was suggested anyway, and an imported custom profile never was. Fetch the catalog from the connected gateway instead. The warning moves out of argument parsing and into sandbox create and sandbox template create, where a client already exists. A catalog fetch failure is not fatal — the warning degrades to its generic form rather than blocking sandbox creation. Extract the ListProviderProfiles paging loop from provider list-profiles into a shared fetch_provider_profile_catalog, which the profile-driven paths now share. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: infer providers from the gateway catalog Command-to-provider inference went through a hardcoded alias table in openshell-providers, a second copy of the built-in catalog's identifiers. A custom profile could never be inferred no matter what binaries it declared, and the table drifted from the profiles it mirrored. Infer from the connected gateway's catalog instead: match the command's basename against each profile's ID and against the basenames of the binaries the profile authorizes. A profile that names /usr/bin/claude is the profile for running claude. The match must be unique — where several profiles claim a command, the user names one with --provider — and an empty catalog infers nothing. No command is special-cased. `binaries` is the operator's authorization statement, so a profile that declares a binary claims the command that runs it, whatever that binary is; narrowing that belongs in the profile rather than in a list compiled into the CLI, which could never cover an unbounded catalog anyway. Breaking: the retired aliases stop resolving, and commands the old table never listed can now infer. git, pip and uv are declared by the github and pypi example profiles, so they infer where they previously did not. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers,server)!: resolve provider profiles by exact ID normalize_provider_type was a hardcoded alias map — gh to github, claude to claude-code, vertex to google-vertex-ai — and a second copy of the built-in catalog's identifiers, independent of the YAML it mirrored. It let a provider type resolve to a profile the operator never named, and it made the built-in IDs behave as a reserved namespace. Remove it, along with detect_provider_from_command and the alias-normalizing ProviderRegistry::inject_env. A provider type now names a profile exactly: - the effective catalog resolves an ID or reports it absent, with no alias retry - plugins activate only for the ID of a profile the gateway resolved; a provider with no resolvable profile gets no plugin projection, instead of falling back to an alias guess - the CLI surfaces the gateway's not-found instead of retrying under an alias - telemetry buckets by profile ID, and the gitlab, opencode and outlook buckets go with the aliases that were their only source Breaking: `--type gh`, `--type claude` and the other aliases no longer resolve. Use the profile's own ID. Signed-off-by: Philippe Martin <phmartin@redhat.com> * feat(server): fail closed when a provider's profile is absent A sandbox composed from a provider whose profile the gateway cannot resolve started anyway: the credential and policy builders warned and skipped, so the sandbox came up carrying none of that provider's credentials or network policy. The operator learned about it later, as a denied connection or a missing environment variable, rather than as the configuration error it is. Check at the two composition boundaries — CreateSandbox and AttachSandboxProvider — right beside the catalog snapshot already taken there, and reject with a bounded diagnostic naming the provider, the profile it refers to, and the import command that supplies it. Scope-aware: a platform-scoped provider is told to import with --global. Read paths are untouched. ListProviders, GetProvider and profile export keep working so an operator can see and recover an affected provider, and the shared policy builders keep their warn-and-skip for the diagnostic paths that also reach them. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(server): scope the vendor base-URL pin to declared endpoints provider_profile_endpoints_are_active withheld a credential and the provider's policy layer when an openai or anthropic provider pointed its client somewhere other than the public vendor endpoint. The guard was keyed on profile.source == "builtin" and on those two profile IDs, so it protected only profiles OpenShell shipped. Once profiles are import-only no profile is ever builtin, and the guard would silently stop applying — including to an operator who imported providers/openai.yaml verbatim. Key it on the profile instead of on where the profile came from. A profile's endpoints are the boundary its credential is bound to, so if the provider configures a *_BASE_URL pointing at a host the profile does not declare, the profile no longer describes where that credential goes and is treated as endpointless. Host matching reuses the DNS-label-aware matcher in openshell-core, so wildcard endpoints such as Vertex's *-aiplatform.googleapis.com resolve correctly. A profile with no declared endpoints has no boundary to contradict, and a config value that names no host is not a redirect this can reason about. Both keep the profile active. The control now covers every endpoint-bearing profile, including an operator's own, and the last "builtin" string leaves the gateway. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): import example provider profiles during gateway bring-up The e2e suites create providers from github, openai, nvidia, claude-code and google-cloud, and the GitHub lane runs a real git clone through the profile's binary attribution. A gateway serves only the profiles an operator imported, so the lanes have to import them. Add e2e_import_example_provider_profiles to the shared bring-up helpers and call it from the Docker, Podman, Kubernetes and VM wrappers once the gateway is healthy and registered. It runs the same command the upgrade notes give operators, against the repository's own providers/ directory, so the lanes exercise the documented path rather than a test-only shortcut. Lands before the default changes: a same-ID user profile already shadows a built-in, so importing works today. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(providers)!: make provider profiles import-only OpenShell compiled fifteen provider profile YAML files into every release binary and selected them by default, so a fresh gateway published a catalog it was never configured with. Those profiles are not image-neutral: their binary selectors name paths that exist in the OpenShell Community image, so changing the sandbox image could make a profile inert while the catalog still advertised it. Every endpoint, credential name and binary path in providers/ was also effectively part of the 0.1.0 public contract. A gateway's catalog is now exactly what an operator imported: - providers/*.yaml is no longer include_str!'d, and builtin_profiles() is gone from the openshell-providers API - the default provider_profile_sources is [{ type = "user" }], and a gateway with nothing imported reaches ready and serves an empty catalog — an empty catalog is a valid state, not a startup failure - the builtin source type is removed from the configuration schema, and a gateway.toml that still names it is rejected at parse time with the import command rather than an unknown-variant error - nothing reserves the canonical identifiers any more, so github, pypi, anthropic and the rest import at their own IDs and the imported profile is the only definition for that ID; the static_fallback that kept a shipped definition resident behind an imported one is gone with the source that produced it A collision between an interceptor-vended profile and an imported one still fails closed, which is the behavior that shadowing quietly bypassed for built-ins. Operators upgrading should export the profiles their deployment relies on before upgrading, or copy them from the providers/ directory of the matching release tag, then import them at the scope their providers use. Signed-off-by: Philippe Martin <phmartin@redhat.com> * docs(providers): document import-only provider profiles Rewrite the provider profile documentation around a catalog the operator builds rather than one the gateway ships. - gateway-config: the default is [{ type = "user" }], the builtin source type is gone and rejected at startup, and an empty catalog is a valid ready state - profiles: replace the built-in profile table and its shadowing semantics with the import workflow, and say plainly that a profile whose binary paths do not match the image is inert - manage-providers: replace the two fixed provider-type tables with `provider list-profiles`, and describe catalog-driven command inference - inference-routing: drop the "still loads built-in profiles for compatibility" transition text; its migration walkthrough is now the normal path - quickstart and the Docker Compose, GitHub, AWS, Google Cloud and Vertex tutorials: import the profile before creating the provider, since nothing resolves without it - release notes: a 0.1.0 migration section covering export-before-upgrade, importing from the release tag's providers/ directory, what happens to a provider whose profile is missing, and the removal of the legacy type aliases - README, architecture, the openshell-cli and debug-openshell-cluster skills, and the governance interceptor example follow the same change; the example's smoke assertion now imports a profile to prove an authoritative interceptor hides it Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(cli,e2e): distinguish catalog failures and authenticate OIDC seeding Addresses two findings from review (GATOR-c1bd9867-02 and -03). The provider profile catalog lookup collapsed its error into an empty catalog, so an authorization, availability or transport failure on ListProviderProfiles was indistinguishable from a gateway that genuinely has no profiles. Command inference then resolved nothing and the sandbox was created without the provider it needed, deferring the failure to the workload. Keep the lookup's outcome instead of discarding it. The credential warning is advisory and still degrades to its generic form, but inference now consults the catalog only when there is a command to resolve and surfaces the lookup failure when there is, naming the gateway and pointing at explicit --provider selection. Two regression tests cover it through a fake gateway whose ListProviderProfiles returns UNAVAILABLE while every other RPC succeeds: sandbox creation fails without sending a provider-less create request, and a sandbox with no trailing command still succeeds because it needs no catalog. The e2e profile seeding also ran unauthenticated in the OIDC lanes. Those lanes deliberately skip gateway registration and start the gateway without a TLS client CA, so no mTLS identity exists and no token has been acquired when the import runs; the wrapper exited during setup. Skipping the import is not sufficient because the provider tests now require the claude-code profile. Add e2e_register_oidc_admin_session, which mints an administrator token with Keycloak's password grant — the same grant the OIDC test helpers use — and writes the gateway metadata and token bundle that an interactive login would have stored, so the import runs as an authenticated administrator. The mTLS lanes keep the direct import unchanged. Both affected wrappers are covered: with-podman-gateway.sh had the same defect as with-docker-gateway.sh. Signed-off-by: Philippe Martin <phmartin@redhat.com> * refactor(cli)!: remove command-derived provider attachment Addresses GATOR-c1bd9867-01. A profile's `binaries` list authorizes a binary to reach that profile's endpoints. It is not a statement that running the binary asks for the provider, and reading it as attachment intent let a command silently gain provider authority: the aws-s3 example declares /bin/bash, so `sandbox create -- bash` resolved an existing aws-s3 provider and attached its credential-backed capability and network policy to the shell and its descendants. Attachment of an already-created provider never prompted, so the confirmation flow did not guard it. The distinction that would make inference sound — whether a declared binary is a profile's client or merely a permitted runtime — cannot be expressed: NetworkBinary carries only a path, and the field that encoded it was removed in 0.1.0. Any substitute is a guess. Restricting the guess to a unique claimant does not help, because uniqueness measures how sparse the catalog is rather than what the user intended, and a compiled list of "generic" commands could never cover an unbounded operator catalog. Remove trailing-command inference rather than approximate it. A provider is attached only when named with --provider, which still creates a missing provider from local discovery when the name matches an imported profile ID. Sandboxes with no providers remain a normal, fully supported state. Removing inference also settles GATOR-c1bd9867-02: with no consumer deriving authority from the catalog, its only remaining use is the advisory credential warning, so a failed lookup degrades that warning instead of blocking creation. The regression test now asserts that an unreachable catalog still creates the sandbox and attaches nothing. While repurposing the deduplication test, a pre-existing defect surfaced: repeating a name in --provider auto-created it twice and the second attempt failed with "provider already exists", because the explicit pass never consulted the set of names it had already handled. Guard it. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): load the OIDC token and trust the gateway when seeding profiles Addresses the carried finding GATOR-c1bd9867-03. e2e_register_oidc_admin_session established a session the CLI never used. The CLI decides whether to load a stored bearer token by matching on the auth_mode field of the gateway metadata alone; the helper omitted that field, so the metadata fell through to the default arm and oidc_token.json stayed on disk unread. The profile import went out unauthenticated exactly as it had before the helper existed. The helper also installed no trust anchor, leaving the CLI unable to verify the gateway's self-signed serving certificate. Write auth_mode = "oidc" so the stored token is loaded, and install the CA at <gateway>/mtls/ca.crt. Only the CA is installed: with no client certificate or key on disk the CLI falls back to CA-only server verification and authenticates with the bearer token, which is what these lanes need because they start the gateway without --tls-client-ca. Certificate verification stays on; no insecure transport override is introduced. Assert the session before anything depends on it. ListProviderProfiles is annotated auth_mode: "bearer", so it cannot succeed unless the token was loaded and accepted. The helper now fails at that point, naming the gateway config directory and echoing the CLI output, rather than letting the defect surface later as an opaque profile import error. Both affected wrappers pass the PKI directory and CLI binary the helper needs; the mTLS lanes are untouched. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): request the OpenShell scopes when minting the admin token The OIDC lanes failed at the session assertion added in b25dd0298 with PERMISSION_DENIED and "scope 'provider:read' required". Authentication was working -- the gateway logged ListProviderProfiles at gRPC status 7, which it can only reach once the bearer token has been loaded and accepted. The token simply carried no OpenShell scope. sandbox:*, provider:*, config:*, workspace:* and openshell:all are optional client scopes on the openshell-cli client in scripts/keycloak-realm.json, so Keycloak mints them only when the request asks for them. The password grant here asked for nothing, leaving the realm defaults (openid, profile, email, roles, web-origins, acr) and an access token that authorizes no RPC. Request "openid openshell:all", as e2e/python/oidc/oidc_auth_test.py already does for its administrator tokens. openshell:all is SCOPE_ALL in crates/openshell-server/src/auth/authz.rs, so one scope covers the setup calls without enumerating them. Verified against the realm: the token goes from "email profile openid" to "email profile openshell:all openid". Record the same scopes in metadata.json. The helper writes the bundle openshell gateway login would have stored, and oidc_scopes is the field that login path reads back, so leaving it out would misdescribe the stored token. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): seed provider profiles once per VM lane gateway The VM lanes failed on the second test target with "custom provider profile 'anthropic' already exists" for all fifteen example profiles, and the import exited non-zero. e2e_import_example_provider_profiles was called from run_e2e_test, so it ran once per target -- four times against one long-lived gateway. That was harmless while import overwrote silently, but profiles are now import-only: ImportProviderProfiles is create-only and reports an existing id as an error-severity diagnostic, with no overwrite flag on the request. The first import therefore succeeds and every later one fails. Hoist the call to just after the conformance run, which is where the docker, podman and kube lanes already seed their catalogs. The profiles persist for the gateway's lifetime, so every target still finds them, and the E2E_TEST_OVERRIDE path is covered by the same single call. Signed-off-by: Philippe Martin <phmartin@redhat.com> * test(e2e): name the provider explicitly in the OIDC workspace-user test user_can_create_sandbox_with_inferred_provider_command reached the gateway once the admin session was fixed, and then failed: it asserted a missing-provider error, but the sandbox was created and died provisioning with "failed to spawn sandbox entrypoint process 'claude-code'". The test drove provider resolution by passing claude-code as the trailing command and relying on the CLI to infer the provider type from it. That inference is what this branch removed, so the trailing word is now nothing but an entrypoint, and the image has no such binary. The regression the test guards is not inference itself: it is that a workspace user resolving a provider is not gated behind Platform Admin. Name the provider with --provider, the only remaining way to attach one. The CLI still has to fetch the claude-code profile before it can auto-create the provider, so the lookup a workspace user must be allowed to make still happens, and auto-creation still stops at the non-interactive branch with "missing required provider". Both assertions therefore keep their meaning. Rename the test and rework its comments to describe what it now exercises; the old name would otherwise outlive the behavior it was named for. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(providers): reconcile import-only profiles with main Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
9505ca5ed1 |
chore: remove Bazel build support (#2840)
Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
0d708d6d51 |
fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints Signed-off-by: Adrien Langou <alangou@nvidia.com> * refactor(cli): extract allowed-ip option parsing Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(policy): gate endpointless credential bindings Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
1959ea19be |
build(bazel): establish RFC 0012 Rust reference graph (#2414)
Establish the Phase 1 reference implementation proposed by RFC 0012 while retaining the existing Cargo and Mise workflows during evaluation. - pin Bazel 9.1.1 and configure Bzlmod, rules_rs, LLVM, protobuf, and Rust 1.95 toolchains - import third-party crates from Cargo metadata and propagate the workspace version into Bazel targets - add library, binary, proc-macro, unit-test, and integration-test targets across the supported Rust workspace crates and drivers - generate protobuf Rust sources and descriptor sets under Bazel while preserving Cargo-compatible generated-code imports - annotate aws-lc-sys and zstd-sys native dependencies, build Z3 4.15.2 from source, and generate z3-sys bindings - make CLI and procfs test fixtures available as explicit Bazel inputs without relying on fixed host binary paths - define optimized release targets for Linux x86_64 and aarch64 CLI, sandbox, and gateway binaries, plus macOS aarch64 artifacts - add Bazel, buildifier, and lcov to the Nix development environment RFC: 0012 (rfc12 branch) Refs: #2491 Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
9377e0d5fe |
fix(providers): allow git clone/fetch via default GitHub provider (#2317)
* fix(providers): allow git clone/fetch via default GitHub provider The github.com:443 git-transport endpoint used the read-only access preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied those operations and `gh repo clone` / `git clone https://...` failed. Replace the preset with explicit rules that permit the read-only methods plus POST */git-upload-pack, so clone/fetch work while push (git-receive-pack) stays blocked. Enabling push still requires an explicit policy proposal. Why allowing this POST is still read-only: in git's smart HTTP protocol POST is an RPC transport, not a write. A clone/fetch does GET */info/refs (ref discovery) followed by POST */git-upload-pack, whose body is only the client's want/have negotiation; the server responds with a packfile and nothing on the server is modified (data flows server -> client). The service names are from the server's perspective: git-upload-pack = the server uploads a pack to the client (a read/ download), while git-receive-pack = the server receives a pack from the client (the actual write/push). The new rule is scoped to */git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs to github.com remain denied. Add a provider-profile regression test and a rego enforcement test covering ref discovery, upload-pack (allowed), and receive-pack (denied). Closes #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): strengthen git-transport regression and add clone e2e Pin the exact allowed rule set for the built-in github git-transport endpoint in both the provider-profile and composed-policy tests, so a broader or additional POST rule (e.g. POST **) that could enable push via git-receive-pack fails the test instead of passing a substring check. Add an e2e test that attaches the built-in github provider and clones a public repo over HTTPS, exercising provider attachment, effective-policy composition, TLS interception, and real git behavior. Update the Providers V2 docs so the github.com git-transport endpoint shows explicit clone/fetch rules instead of the stale read-only preset. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): isolate providers_v2 mutation in clone e2e The clone e2e enables the gateway-global providers_v2_enabled setting. Restore its exact prior value (or absence) captured via GetGatewayConfig instead of unconditionally deleting it, and serialize the mutation across xdist workers with an exclusive file lock on the run's shared base temp dir, so a shared or pre-configured gateway is left untouched and parallel workers cannot race the read-modify-restore. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): serialize providers_v2 mutation with a suite-wide guard The clone e2e's per-fixture lock only coordinated fixtures that acquired it; other xdist workers hit the same gateway without it and could observe the transiently-enabled providers_v2_enabled global during their own sandbox creation (CWE-362). Add an autouse readers-writer guard in conftest: every test holds a shared lock on the gateway config, and a test marked exclusive_gateway_config holds an exclusive lock. Mark the clone test exclusive so no other worker is mid-test while it enables and restores the gateway-global setting. Exact prior-value restoration is retained. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
aa483ecb9a |
feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782)
Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The gateway calls sts:AssumeRole and writes three short-lived credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the provider record; the proxy re-signs requests with SigV4. Adds the aws and aws-s3 provider profiles and a declarative multi-output refresh model (additional_outputs) so one AssumeRole co-mints all three credentials. Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
4ee27d9955 |
feat(sandbox,providers): add aws-bedrock as a recognized inference provider (#1704)
* feat(sandbox): allow AWS Bedrock InvokeModel paths through the L7 router
Adds two patterns to `default_patterns()` so the supervisor's L7
inference router recognizes the Bedrock InvokeModel URL shape and
forwards matched requests to the registered upstream:
- `POST /model/{modelId}/invoke` → aws_bedrock_invoke
- `POST /model/{modelId}/invoke-with-response-stream` → aws_bedrock_invoke_stream
The `{modelId}` segment is wildcarded by extending `detect_inference_pattern`
to handle one middle `/*/` segment in addition to the existing trailing
`/*`. The wildcard is constrained to a single non-empty path segment to
avoid path-traversal liabilities — `/model//invoke` and `/model/a/b/invoke`
both no-match.
Without this, sandboxes running Claude Code in its native Bedrock mode
(`CLAUDE_CODE_USE_BEDROCK=1`, `ANTHROPIC_BEDROCK_BASE_URL`, AWS-style
auth) hit the supervisor with `403 connection not allowed by policy`
because their URL doesn't match `/v1/*` shapes. The fix unblocks
operators wanting to register direct AWS Bedrock, an in-cluster
Bedrock-compatible bridge, or a Bedrock-emulating LiteLLM as
`--type aws-bedrock` providers.
Tests cover: positive matches for invoke + invoke-with-response-stream,
query-string handling, GET rejection, empty-segment rejection,
multi-segment rejection, and unknown-action rejection.
Companion changes (provider discovery spec + YAML profile) follow in
the next commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* feat(providers): add aws-bedrock provider profile + discovery spec
Adds `aws-bedrock` to the built-in provider catalog so operators can
run `openshell provider create --type aws-bedrock --credential ...`
and have the gateway treat it as a first-class inference provider
alongside `anthropic`, `openai`, etc.
- `providers/aws-bedrock.yaml`: YAML profile declaring four credentials
(AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, AWS_REGION).
Default endpoint is `bedrock-runtime.us-east-1.amazonaws.com:443`;
operators in other regions or running against a Bedrock-compatible
proxy override via the operator-supplied `BEDROCK_BASE_URL` config-key
(mirrors `ANTHROPIC_BASE_URL` for the `anthropic` provider).
- `crates/openshell-providers/src/providers/aws_bedrock.rs`: the
`ProviderDiscoverySpec` so `openshell provider create --auto-providers`
picks up AWS_* env vars from local credentials.
- `crates/openshell-providers/src/providers/mod.rs`: register the module.
- `crates/openshell-providers/src/lib.rs`: register the SPEC in the
default registry alongside the other providers.
- `crates/openshell-providers/src/profiles.rs`: include the new YAML in
`BUILT_IN_PROFILE_YAMLS`.
What this PR explicitly does NOT add (intentionally separated for
review-size reasons; will follow up):
- A SigV4 signer in `openshell-router`. The current change simply
declares the protocol; a follow-up PR adds outbound SigV4 signing
using the `aws-sigv4` crate and a new `auth_style: sigv4` validator
branch in profiles.rs. Operators who don't need SigV4 (e.g. an
in-cluster bridge that ignores it and authenticates separately to
the upstream) can use this PR today.
- Body translation between Bedrock InvokeModel shape and other
inference shapes. The router treats Bedrock requests as opaque
pass-through; if the operator's upstream is real AWS Bedrock it
speaks Bedrock natively, if it's a translating bridge the bridge
does any conversion server-side.
- `BEDROCK_BASE_URL` placeholder substitution in the YAML loader.
Today the YAML's `host` is a literal default; operators override
with the config-key the same way `ANTHROPIC_BASE_URL` works.
Tested: `cargo test -p openshell-providers` (35 tests green) and
`cargo test -p openshell-sandbox --lib l7::inference` (40 tests green
including the seven new aws_bedrock cases from the previous commit).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* revert(providers): drop legacy aws-bedrock SPEC, rely on v2 YAML profile
Addresses johntmyers's review on NVIDIA/OpenShell#1704: net-new
providers should land via the v2 YAML profile only and should NOT
require changes to the legacy `ProviderDiscoverySpec` registry.
- Delete `crates/openshell-providers/src/providers/aws_bedrock.rs`
(the legacy SPEC + `test_discovers_env_credential!` invocation).
- Drop `pub mod aws_bedrock;` from `crates/openshell-providers/src/providers/mod.rs`.
- Drop `registry.register(providers::aws_bedrock::SPEC)` from
`crates/openshell-providers/src/lib.rs`.
Kept:
- `providers/aws-bedrock.yaml` and the `include_str!` in
`BUILT_IN_PROFILE_YAMLS` (`profiles.rs`) — the v2 path.
`discover_from_profile()` (`crates/openshell-providers/src/discovery.rs`)
picks up AWS_* env vars via `discovery.credentials` in the YAML.
- L7 router patterns in `crates/openshell-sandbox/src/l7/inference.rs`
— orthogonal to the provider registry.
The discovery test in the deleted file goes with it; v2 doesn't have
an established per-provider env-var-pickup unit test pattern, and
other YAML-only registrations (none today, but this is the new
direction) won't carry one either.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(providers): include aws_session_token in discovery + update profile assertion
Two fixes from johntmyers's gator-agent re-check on NVIDIA/OpenShell#1704:
1. `providers/aws-bedrock.yaml`: add `aws_session_token` to
`discovery.credentials`. The credential is declared in the profile
but was missing from the discovery scan list, so Providers v2
`--from-existing` would silently drop temporary AWS credentials
(STS / IRSA scenarios).
2. `crates/openshell-server/src/grpc/provider.rs`: update the static
`list_provider_profiles_returns_built_in_profile_categories`
assertion to include `aws-bedrock` at alphabetical position 0.
Adding `providers/aws-bedrock.yaml` to BUILT_IN_PROFILE_YAMLS made
the prior `["claude-code", "github", "nvidia"]` expectation stale.
Remaining blockers from the same review (deferred to follow-up
commits): `inference::profile_for` registration for aws-bedrock,
user-facing provider + inference-routing docs, and an
`upsert_cluster_inference_route` integration test.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* feat(inference): register aws-bedrock profile (bridge-fronted) + docs
Addresses johntmyers's blocking review feedback on PR #1704:
"aws-bedrock still is not wired into the managed inference.local route
registry. profile_for only registers openai, anthropic, and nvidia, so
inference set --provider <aws-bedrock-provider> will reject this
provider before the new sandbox L7 patterns can be used."
Approach: register aws-bedrock as a *bridge-fronted* upstream — the
router does not inject any auth header on outbound requests; the
configured BEDROCK_BASE_URL is expected to point at a translating
bridge / Bedrock-compatible proxy that handles auth in its own pod.
This is the shape the L7 patterns commit (8b30211a) and the YAML
profile (6b51e1a6) were designed for. SigV4 signing for direct AWS
Bedrock is a separate follow-up; see PR thread.
Changes:
- core::inference::AuthHeader: add `None` variant for upstreams that
authenticate themselves.
- core::inference: add AWS_BEDROCK_PROFILE static + register in
profile_for. Default base URL is bedrock-runtime.us-east-1, override
via BEDROCK_BASE_URL config-key (mirrors ANTHROPIC_BASE_URL pattern).
Empty credential_key_names + auth: None means no router-side
credential lookup at route time.
- router::backend: handle AuthHeader::None as a no-op (skip auth
injection).
- server::inference::resolve_provider_route: gate find_provider_api_key
on auth != None. aws-bedrock providers with empty credentials now
resolve cleanly. Updated the unsupported-type error message to
include aws-bedrock in the supported list.
- server::inference tests: add positive
upsert_cluster_route_succeeds_for_aws_bedrock_without_api_key test
covering the new code path end-to-end (provider with empty creds +
BEDROCK_BASE_URL config → upsert succeeds → resolved route has
empty api_key + provider_type aws-bedrock + bridge URL).
- core::inference tests: profile_for_known_types covers aws-bedrock,
case-insensitive lookup, plus three new aws-bedrock-specific tests
(auth: None, no credential keys, bedrock-specific protocols).
- docs/sandboxes/inference-routing.mdx: header forwarding row
mentions aws-bedrock has no passthrough headers; new tabs in
Supported API Patterns (InvokeModel + InvokeModelWithResponseStream)
and Create a Provider (with the bridge-fronted shape note + SigV4
deferral).
- docs/sandboxes/manage-providers.mdx: new row in Supported Provider
Types table; new row in Supported Inference Providers table.
Verification (in dev container):
- cargo check -p openshell-core -p openshell-router -p openshell-server: clean
- cargo test -p openshell-core --lib inference: 14/14 pass (incl. 3 new)
- cargo test -p openshell-server --lib inference::tests::upsert: 6/6 pass
(incl. new aws-bedrock test)
- cargo fmt --check: clean
- cargo clippy --all-targets -D warnings: clean
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(aws-bedrock): bridge-only YAML, doc actual cmd shape, neg test
Addresses four findings from gator-agent's #1704 re-check on 4ab587f1:
- **Item 5** (YAML collects unused AWS creds): mark all four AWS
credentials `required: false` and clear `discovery.credentials`.
Bridge-fronted routing intentionally does not consume AWS
credentials, so `--from-existing` no longer scans for them. The
credentials remain in the schema (not deleted) so the SigV4
follow-up can flip them back without a schema migration. Added a
multi-line description that names the bridge-fronted shape and the
SigV4 deferral so readers don't have to cross-reference the PR
thread.
- **Item 3** (docs show command that the CLI rejects): rewrite the
Create-a-Provider example for AWS Bedrock to use the actual
required shape — placeholder `--credential AWS_ACCESS_KEY_ID=
unused-bridge-fronted-shape` plus the `--config BEDROCK_BASE_URL`.
The placeholder satisfies the gRPC handler's
`provider.credentials.is_empty()` rejection without expanding
server-side validation; the router ignores it on the outbound path
because `auth: AuthHeader::None` skips header injection. Operators
see a clearly-labeled placeholder in `provider get` output.
- **Item 1** (validator probe): document `--no-verify` as required
for `openshell inference set --provider <aws-bedrock>` since the
default validation probe doesn't recognize the
`aws_bedrock_invoke` / `aws_bedrock_invoke_stream` protocols. Doc
now shows the full `provider create` + `inference set --no-verify`
flow with rationale for both decisions inline.
- **Item 6** (docs polish): `inference-routing.mdx` summary row now
lists AWS Bedrock alongside NVIDIA, Anthropic, Vertex AI, and
OpenAI-compatible providers, with the bridge-fronted caveat
inline.
Test additions in `crates/openshell-server/src/inference.rs`:
- Renamed the existing aws-bedrock test from
`..._without_api_key` to `..._with_bridge_url` and updated it to
use a placeholder credential (mirroring the doc-recommended
pattern operators will copy-paste). The `auth: None` path still
produces an empty `api_key` on the resolved route — the test now
documents that the credential is *stored* but not *used*.
- Added `upsert_cluster_route_rejects_aws_bedrock_without_bedrock_base_url`:
the negative half of johntmyers' "successfully used by
upsert_cluster_inference_route or intentionally rejected with a
clear documented error" ask. With
`default_base_url: ""` and no `BEDROCK_BASE_URL` config, route
resolution returns `InvalidArgument` naming the missing base_url
rather than silently forwarding prompts to AWS Bedrock with no
usable auth.
Verification (in dev container):
- cargo test -p openshell-core --lib inference: 18/18 (incl. 3 new)
- cargo test -p openshell-server --lib inference::tests::upsert: 8/8
(incl. 2 new aws-bedrock cases — positive + negative)
- cargo fmt --check: clean
- cargo clippy --all-targets -D warnings: clean
Item 2 (router-side enforcement of operator-configured Bedrock model
path, replacing the current verbatim path forwarding + body-only
model rewrite) is the remaining blocker and is genuinely separable —
it touches the L7 router with streaming-aware test coverage.
Deferring to its own commit so the security-critical change gets the
review attention it deserves.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(router): enforce operator-configured Bedrock model in request path
Closes the security-blocking item from gator-agent's #1704 re-check
on 4ab587f1: "Bedrock carries the model id in /model/{modelId}/invoke,
but the router currently forwards the caller's original path and only
rewrites JSON body model. That lets sandbox code choose a different
upstream model than the operator-configured route model, and may also
mutate native Bedrock request bodies incorrectly."
Two changes in `prepare_backend_request`:
1. **Path rewrite for Bedrock routes.** Before computing the upstream
URL, parse the inbound path's `/model/<id>/invoke[-with-response-stream]`
shape and substitute the operator-configured `route.model` for the
caller-supplied model segment. Sandbox code that hardcodes a
different model still works (we don't reject on mismatch), but the
operator's configured model is what reaches the upstream / bridge.
If the inbound path is somehow not a recognized Bedrock shape on a
Bedrock route (the L7 pattern detector upstream of the router
should never produce this combination), reject with
RouterError::Internal naming the offending path rather than
forwarding verbatim.
2. **Skip body-model injection for Bedrock routes.** The existing body
rewriter unconditionally inserts `route.model` into the JSON body
for non-Vertex routes. AWS Bedrock InvokeModel encodes the model
in the URL path; the body is the raw provider-specific payload
(Anthropic Messages for Claude, Mistral payload for Mistral, etc.)
and must not be mutated. The branch ordering is now:
needs_vertex_anthropic_version → strip body model + inject
anthropic_version; route_is_bedrock → leave body alone; else →
inject route.model (existing default).
New helpers, all in `crates/openshell-router/src/backend.rs`:
- `route_is_bedrock(route)` — true when route.protocols contains
aws_bedrock_invoke or aws_bedrock_invoke_stream.
- `parse_bedrock_invocation_path(path)` — returns
Some((model_id, "/invoke" | "/invoke-with-response-stream")) for
paths matching the recognized Bedrock shapes. Strips query strings.
Rejects empty model ids and multi-segment ids (defense-in-depth
matching the L7 pattern detector's existing guards).
- `rewrite_bedrock_path(route, path)` — returns the path with the
caller's model segment replaced by route.model.
Test coverage in the same file (9 new tests):
- parse_bedrock_invocation_path: positive cases for both invoke
variants, query-string stripping; negative cases for empty model id,
multi-segment id, unknown action, wrong prefix, missing slash.
- route_is_bedrock: matches both protocol variants singly and
combined; rejects openai_chat_completions.
- rewrite_bedrock_path: substitutes operator model on both invoke
variants; returns None for non-Bedrock paths.
- bedrock_route_rewrites_model_in_path_and_preserves_body
(wiremock end-to-end): caller sends /model/some-other-model/invoke
with a body containing model: "caller-supplied-model-name". Mock
asserts the upstream receives /model/<operator-model>/invoke and the
body's model field is the caller's value (NOT route.model) — proves
both the path rewrite and the body preservation.
- bedrock_route_streaming_rewrites_model_in_path: same contract for
invoke-with-response-stream.
- bedrock_route_rejects_non_bedrock_path: defense-in-depth coverage of
the Internal-error path when a Bedrock route receives a path that
doesn't match Bedrock shape.
Verification (in dev container):
- cargo test -p openshell-router --lib: 53/53 (incl. 9 new)
- cargo fmt --check: clean
- cargo clippy -p openshell-core -p openshell-router -p openshell-server
--all-targets -- -D warnings: clean
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(sandbox/l7): declare framing for Bedrock patterns
Closes the buffered-vs-streaming framing warning from gator-agent's
re-check on 4ab587f1: "Bedrock InvokeModel should be buffered while
InvokeModelWithResponseStream is streaming. Please add framing/coverage
so /model/{id}/invoke cannot be corrupted by the streaming proxy's
truncation/error-frame behavior."
The InferenceApiPattern struct gained a `framing: ResponseFraming`
field upstream after the original Bedrock-patterns commit (#22b78cff)
landed; the cherry-pick onto current upstream/main left the two
Bedrock entries without the new field. Fixed here:
- aws_bedrock_invoke (POST /model/{id}/invoke):
framing = ResponseFraming::Buffered
InvokeModel returns one JSON object the caller decodes whole. Sending
it through the streaming proxy would risk a mid-body size-cap
truncation or idle-timeout failure appending an SSE error event onto
bytes the caller decodes as one JSON body — the same corruption mode
that drove the existing embeddings + model-discovery to Buffered.
- aws_bedrock_invoke_stream (POST /model/{id}/invoke-with-response-stream):
framing = ResponseFraming::Streaming
InvokeModelWithResponseStream returns an AWS event-stream of binary
chunks; the caller wants chunks incrementally, so the streaming proxy
path is correct.
Two new tests in `crates/openshell-sandbox/src/l7/inference.rs` pin
down the contract:
- aws_bedrock_invoke_is_buffered — detect_inference_pattern returns a
Buffered pattern for /model/<id>/invoke, with explanatory message
naming the corruption mode being prevented.
- aws_bedrock_invoke_stream_is_streaming — same shape, asserting
Streaming for /model/<id>/invoke-with-response-stream.
Verification (in dev container):
- cargo check -p openshell-sandbox: clean (was failing on missing
`framing` field before this commit)
- cargo test -p openshell-sandbox --lib l7::inference::tests::aws_bedrock:
7/7 (incl. 2 new framing tests)
- cargo fmt --check: clean
- cargo clippy --all-targets -- -D warnings: clean
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(inference): drop aws_bedrock_invoke_stream until protocol-aware errors land
Per PR #1704 review (johntmyers): defer Bedrock streaming to a follow-up
that also wires protocol-aware error framing for the AWS event-stream
shape. Until then, surfacing `/model/{id}/invoke-with-response-stream`
risks shipping responses the sandbox cannot interpret on failure.
- AWS_BEDROCK_PROTOCOLS no longer advertises `aws_bedrock_invoke_stream`.
- The L7 inference pattern table drops the streaming entry; only
`aws_bedrock_invoke` (buffered) is recognized.
- Test `aws_bedrock_invoke_stream_pattern_is_deferred` asserts no
pattern claims that protocol so the gap is visible.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(inference): validate AWS Bedrock model_id at route + rewrite, preserve query
Per PR #1704 review (johntmyers, gator agent): the operator-configured
Bedrock model_id flows verbatim into the upstream URL path; the previous
plumbing left no enforcement that the value was a single benign path
segment, opening a path-injection vector.
Defense in depth, both layers:
* `openshell-server::inference`: new `validate_aws_bedrock_model_id`
rejects empty, leading/trailing whitespace, `/`, `\\`, `?`, `#`, `%`,
`..`, control/whitespace characters. Wired into `resolve_provider_route`
ahead of base_url resolution so the route store cannot persist a
malformed model_id. Mirrors `validate_vertex_model_id` exactly.
* `openshell-router::backend`: `rewrite_bedrock_path` now refuses to
construct the upstream URL unless `route.model` passes
`is_valid_bedrock_model_id`, so even a stale or hand-edited route
cannot reach the wire. The parser also drops the
`/invoke-with-response-stream` arm to match the protocol catalog.
* `parse_bedrock_invocation_path` returns the `?`-prefixed query tail
as a third element; `rewrite_bedrock_path` re-attaches it so any
caller-supplied query string is preserved through the model rewrite.
Tests: 5 unit tests for the validator, 1 integration test that
exercises every unsafe-model_id reject path through
`upsert_cluster_inference_route`, plus a router-side rewrite-rejects
test covering 11 unsafe `route.model` values. All 72 server inference
tests + router tests pass.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(aws-bedrock): make built-in profile non-egress-granting; tighten docs
Per PR #1704 review (johntmyers): a single profile that is "usable as
is" — not a profile that auto-grants direct AWS Bedrock egress to any
sandbox that selects it. The bridge IS the egress point and is
operator-managed; the profile must not implicitly punch a hole through
the cluster's network policy on its behalf.
* `providers/aws-bedrock.yaml`: clear `endpoints` and `binaries` to
empty arrays. Rewrite the description to spell out that the profile
is intentionally non-egress-granting, that operators are responsible
for declaring their bridge's egress endpoint and binary attribution,
and that the SigV4 follow-up will repopulate these fields once
router-side signing exists.
* `docs/sandboxes/inference-routing.mdx`:
- Drop the `InvokeModelWithResponseStream` row from the supported
patterns table (matches the protocol-catalog + L7-pattern drop).
- Update the `--no-verify` paragraph to reference only
`aws_bedrock_invoke`.
- Generalise the placeholder-credential rationale: any
standalone-router profile registering `AuthHeader::None` will hit
the same non-empty-credentials structural requirement.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* style(router): satisfy cargo fmt for Bedrock query-tail extraction
CI's rustfmt rejected the one-line `let (path_only, query_tail) = path.find('?').map_or(...)`
shape from commit c51160e4 and required the chained-method layout
instead. Functional behaviour is unchanged; tests still pass.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* style(router): drop unnecessary `crate::` prefix on RouterError matcher
`RouterError` is already imported at the top of the file, so the
two test-side `matches!(result, Err(crate::RouterError::UpstreamProtocol(_)))`
uses trip clippy's `unused_qualifications` under `-D warnings`. Drop
the prefix on both sites; functional behaviour is unchanged.
These warnings predate this PR (originated in
|
||
|
|
48545cfb59 |
feat(sandbox): add GCE metadata emulator for Google Cloud (#1763)
* feat(core): add shared GCP constants module Single source of truth for GCP naming: env var aliases, provider config keys, token search order, and Vertex-specific env vars. Consumed by openshell-server, openshell-providers, and openshell-sandbox. - Add google_cloud.rs with metadata emulator host and loopback address - Define PROJECT_ID, REGION, and SERVICE_ACCOUNT_EMAIL env var aliases - Add provider config key constants for gcp provider implementations - Define TOKEN_ENV_KEYS search order (SA token takes priority over ADC) - Add Vertex-specific env vars for Goose and Claude Code SDK integration - Add STATIC_CONFIG_KEYS as union of all alias arrays for env resolution - Export module via openshell-core lib.rs Signed-off-by: Robert Sturla <rsturla@redhat.com> * feat(providers): add google-cloud and vertex provider plugins Add GoogleCloudProvider and VertexProvider implementing inject_env to project GCP config (project ID, region, SA email, metadata host) into sandbox environment variables. Replace the inline Vertex AI env injection in the server with the registry-based inject_env dispatch. Also adds the google-cloud.yaml provider profile with SA JWT and ADC OAuth2 credential refresh flows. Signed-off-by: Robert Sturla <rsturla@redhat.com> * feat(sandbox): add GCE metadata emulator for GCP Add a loopback HTTP server on 127.0.0.1:8174 inside the sandbox network namespace that emulates the GCE instance metadata API. GCP client SDKs discover it via GCE_METADATA_HOST and obtain credential placeholders that the proxy resolves to real tokens at egress. Add metadata_server module with MetadataHandler trait and netns-aware TCP binding via std::thread (not spawn_blocking) to avoid tokio pool namespace contamination Add google_cloud_metadata module implementing the GCE metadata API subset (token, project-id, email, scopes, service-accounts) Add child_env_resolved() and gcp_token_response() to ProviderCredentialState for GCP-aware credential projection Wire metadata server into sandbox lifecycle before SSH handler Collapse multi-line HTTP response format string into single line Signed-off-by: Robert Sturla <rsturla@redhat.com> * docs(sandbox): add GCP credentials documentation Document the google-cloud provider setup for ADC and service account flows, injected environment variables, metadata emulator behavior, and network policy configuration for GCP APIs. Signed-off-by: Robert Sturla <rsturla@redhat.com> * feat(cli): support --from-gcloud-adc for google-cloud providers Widen --from-gcloud-adc to accept google-cloud providers. The ADC credential key is derived from the provider profile rather than hardcoded per type, so future GCP provider types get ADC support by declaring the right refresh metadata in their profile YAML. Add ProviderTypeProfile::adc_credential() to find the ADC-compatible credential from a profile's refresh metadata. Remove unused VERTEX_AI_ADC_TOKEN_KEY and GCP_ADC_TOKEN_KEY constants. Signed-off-by: Robert Sturla <rsturla@redhat.com> --------- Signed-off-by: Robert Sturla <rsturla@redhat.com> |
||
|
|
36bb9e3ec2 |
feat(providers): add DeepInfra as a built-in inference provider (#1902)
* feat(providers): add DeepInfra as a built-in inference provider (v2 only) - Adds `deepinfra` as a built-in Providers v2 profile (`providers/deepinfra.yaml`) with inference category, Bearer auth, and `DEEPINFRA_API_KEY` discovery - Adds `DEEPINFRA_PROFILE` to inference routing so `inference.local` works with the `deepinfra` provider type - Fixes `build_backend_url` to strip `/v1` from request paths when the base URL contains `/v1/` as an internal segment (e.g. `api.deepinfra.com/v1/openai`), preventing double-versioned paths like `.../v1/openai/v1/chat/completions` - Updates `docs/sandboxes/providers-v2.mdx` and `docs/sandboxes/manage-providers.mdx` with DeepInfra entries; removes the old v1 workaround row that used `openai` type with `OPENAI_API_KEY` Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * fix(providers): address gator review findings for DeepInfra provider - Narrow build_backend_url /v1 dedupe to URLs whose path component is exactly /v1 or starts with /v1/ — prevents regression on proxy endpoints where /v1 is buried deeper (e.g. /api/v1/openai); add regression test for the nested proxy path case - Add deepinfra provider plugin with DEEPINFRA_API_KEY discovery, registered in ProviderRegistry so known_types() and TUI include it - Add deepinfra to unsupported-inference-provider error message in openshell-server for accurate user-facing debugging guidance - Add deepinfra to openai_compatible_profiles_include_embeddings test to lock in the OpenAI-compatible protocol contract Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * fix(router): handle /v1 as final path segment in build_backend_url dedup Extends the /v1 deduplication logic to also strip /v1 from request paths when the base URL's path ends with /v1 (e.g. https://api.groq.com/openai/v1). The previous fix only matched paths starting with /v1/, which regressed providers like Groq whose base path has /v1 as the last segment rather than the first. The nested-proxy exclusion (e.g. /api/v1/openai) is preserved since /v1 appears in the middle — neither first nor last segment. Adds a regression test for the Groq-style base URL. Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * fix(providers): add deepinfra telemetry bucket and update profile list test - Add DeepInfra variant to ProviderProfile telemetry enum and from_raw() mapping so deepinfra providers are tracked in their own bucket rather than falling through to Custom - Map deepinfra in telemetry_provider_profile() in openshell-server - Add deepinfra to list_provider_profiles_returns_built_in_profile_categories test (sorted between cursor and github) - Update architecture/gateway.md inference provider list to include deepinfra Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * style(router): apply cargo fmt to backend.rs Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> --------- Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> |
||
|
|
62c421b21e |
feat(providers): add profile-backed policy visibility (#1640)
* chore: wip providers v2 tui and codex profile * chore: wip effective policy get and codex profile * chore: wip provider profiles and tui detail views * feat(tui): annotate policy proposal review status |
||
|
|
f061b1d923 |
feat(providers): add Google Vertex AI inference provider (#1568)
* feat(providers): add Google Vertex AI provider Adds Vertex AI provider profiles, routing, credential refresh plumbing, CLI support, docs, and regression coverage. Keeps the related NETLINK_ROUTE seccomp allowance needed by Vertex client tooling that calls getifaddrs. * docs: add Vertex AI sandbox usage for Claude Code and OpenCode Cover the full end-to-end setup for running Claude Code and OpenCode inside an OpenShell sandbox via inference.local with a Vertex AI backend: - google-vertex-ai.mdx: add 'Use from a Sandbox' section with tabbed examples for Claude Code (--bare flag, no /v1 suffix) and OpenCode (/v1 suffix required). Add providers_v2_enabled prerequisite and --no-verify note for global region. Document policy proposals table covering metadata.google.internal (always blocked), downloads.claude.ai, and storage.googleapis.com. - inference-routing.mdx: expand 'Use the Local Endpoint' section with tabbed examples for Claude Code, OpenCode, Python OpenAI SDK, and Python Anthropic SDK. Add notes explaining the /v1 path suffix difference between clients. - supported-agents.mdx: update Claude Code and OpenCode rows to mention inference.local support and correct base URL requirements. * fix: address vertex review findings * test(sandbox): retry on spurious Ok in fork-exec ambiguity test On arm64 under heavy CI load, the /proc fd scan in find_socket_inode_owners can transiently miss the parent process's socket fd entry, returning only the child as an owner. This causes resolve_process_identity to return Ok (single owner, no ambiguity check fires) instead of the expected ambiguous-ownership Err. Extend the retry loop to also handle unexpected Ok results, mirroring the existing retry for transient Err results. 10 retries at 50ms gives a 500ms settling window, which is sufficient for procfs to stabilize on loaded arm64 runners. * fix: address vertex review regressions * docs(router): clarify stream_response semantics for Vertex rawPredict routing Document the three call sites of prepare_backend_request and their stream_response values in a caller table: - send_backend_request: false → :rawPredict (unary endpoint) - send_backend_request_streaming: true → :streamRawPredict - verify_backend_endpoint: explicitly false to probe the unary endpoint Cross-reference the table from build_provider_url and is_vertex_anthropic_rawpredict_route so the stream_response=true guard in the suffix upgrade branch is understood in full context. Also note that is_vertex_anthropic_rawpredict_route is a structural predicate (model_in_path + anthropic_messages + :rawPredict suffix), not a named-provider check, so any future provider with the same route shape inherits the transforms automatically. |
||
|
|
e98ea3ee93 | feat(policy): add agentic approval loop (#1528) | ||
|
|
0cef26521a |
feat(providers): derive discovery from profiles (#1503)
* feat(providers): derive discovery from profiles Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): keep v2 discovery profile-only Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(providers): update providers v2 behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): make github profile read-only Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
d255cdd9c9 |
feat(providers): add credential refresh foundation (#1349)
* feat(providers): add credential refresh foundation * feat(providers): mint oauth refresh token credentials * fix(providers): allow delegated refresh bootstrap * fix(providers): guard refresh credential modes * fix(providers): refine refresh lifecycle UX * fix(cli): restore provider test gateway response import * fix(server): clean auth endpoint test qualifications * fix(providers): tighten refresh authorization and collisions * chore(providers): trim bundled v2 profiles * fix(providers): resolve refresh rebase fallout * fix(cli): accept rfc3339 credential expiry * test(providers): update inferred claude provider type * test(providers): avoid removed outlook default profile * test(providers): isolate attach limit fixtures |
||
|
|
043bde279a |
feat(providers): add profile-backed policy composition (#1037)
Foundation for providers v2. Add provider profiles and provider profile composition with user policies. |