Commit Graph
15 Commits
Author SHA1 Message Date
John T. MyersandJohn Myers 07df822090 feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative

Closes #1988

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): move profiles into provider navigation

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* docs(providers): clarify provider attachment lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(tui): scroll provider profile picker

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): honor profile credential semantics

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): prefer exact profile IDs

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): harden authoritative profile adoption

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* test(oidc): align provider fixtures with profiles

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

* fix(providers): preserve authoritative profile lifecycle

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
2026-09-01 19:22:07 +00:00
Simon Scatton 9505ca5ed1 chore: remove Bazel build support (#2840)
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-20 15:08:09 +00:00
alangou 0d708d6d51 fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* refactor(cli): extract allowed-ip option parsing

Signed-off-by: Adrien Langou <alangou@nvidia.com>

* fix(policy): gate endpointless credential bindings

Signed-off-by: Adrien Langou <alangou@nvidia.com>

---------

Signed-off-by: Adrien Langou <alangou@nvidia.com>
2026-08-19 15:18:02 +00:00
Simon Scatton 1959ea19be build(bazel): establish RFC 0012 Rust reference graph (#2414)
Establish the Phase 1 reference implementation proposed by RFC 0012
while retaining the existing Cargo and Mise workflows during evaluation.

- pin Bazel 9.1.1 and configure Bzlmod, rules_rs, LLVM, protobuf, and
  Rust 1.95 toolchains
- import third-party crates from Cargo metadata and propagate the
  workspace version into Bazel targets
- add library, binary, proc-macro, unit-test, and integration-test
  targets across the supported Rust workspace crates and drivers
- generate protobuf Rust sources and descriptor sets under Bazel while
  preserving Cargo-compatible generated-code imports
- annotate aws-lc-sys and zstd-sys native dependencies, build Z3 4.15.2
  from source, and generate z3-sys bindings
- make CLI and procfs test fixtures available as explicit Bazel inputs
  without relying on fixed host binary paths
- define optimized release targets for Linux x86_64 and aarch64 CLI,
  sandbox, and gateway binaries, plus macOS aarch64 artifacts
- add Bazel, buildifier, and lcov to the Nix development environment

RFC: 0012 (rfc12 branch)
Refs: #2491

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
2026-08-03 09:05:40 +00:00
Russell Bryant 9377e0d5fe fix(providers): allow git clone/fetch via default GitHub provider (#2317)
* fix(providers): allow git clone/fetch via default GitHub provider

The github.com:443 git-transport endpoint used the read-only access
preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires
a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied
those operations and `gh repo clone` / `git clone https://...` failed.

Replace the preset with explicit rules that permit the read-only methods
plus POST */git-upload-pack, so clone/fetch work while push
(git-receive-pack) stays blocked. Enabling push still requires an
explicit policy proposal.

Why allowing this POST is still read-only: in git's smart HTTP protocol
POST is an RPC transport, not a write. A clone/fetch does GET
*/info/refs (ref discovery) followed by POST */git-upload-pack, whose
body is only the client's want/have negotiation; the server responds
with a packfile and nothing on the server is modified (data flows
server -> client). The service names are from the server's perspective:
git-upload-pack = the server uploads a pack to the client (a read/
download), while git-receive-pack = the server receives a pack from the
client (the actual write/push). The new rule is scoped to
*/git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs
to github.com remain denied.

Add a provider-profile regression test and a rego enforcement test
covering ref discovery, upload-pack (allowed), and receive-pack (denied).

Closes #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): strengthen git-transport regression and add clone e2e

Pin the exact allowed rule set for the built-in github git-transport
endpoint in both the provider-profile and composed-policy tests, so a
broader or additional POST rule (e.g. POST **) that could enable push
via git-receive-pack fails the test instead of passing a substring
check. Add an e2e test that attaches the built-in github provider and
clones a public repo over HTTPS, exercising provider attachment,
effective-policy composition, TLS interception, and real git behavior.

Update the Providers V2 docs so the github.com git-transport endpoint
shows explicit clone/fetch rules instead of the stale read-only preset.

Refs #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): isolate providers_v2 mutation in clone e2e

The clone e2e enables the gateway-global providers_v2_enabled setting.
Restore its exact prior value (or absence) captured via GetGatewayConfig
instead of unconditionally deleting it, and serialize the mutation
across xdist workers with an exclusive file lock on the run's shared
base temp dir, so a shared or pre-configured gateway is left untouched
and parallel workers cannot race the read-modify-restore.

Refs #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): serialize providers_v2 mutation with a suite-wide guard

The clone e2e's per-fixture lock only coordinated fixtures that acquired
it; other xdist workers hit the same gateway without it and could
observe the transiently-enabled providers_v2_enabled global during their
own sandbox creation (CWE-362).

Add an autouse readers-writer guard in conftest: every test holds a
shared lock on the gateway config, and a test marked
exclusive_gateway_config holds an exclusive lock. Mark the clone test
exclusive so no other worker is mid-test while it enables and restores
the gateway-global setting. Exact prior-value restoration is retained.

Refs #1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-07-20 21:07:17 +00:00
Russell Bryant aa483ecb9a feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782)
Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The
gateway calls sts:AssumeRole and writes three short-lived credentials
(AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the
provider record; the proxy re-signs requests with SigV4. Adds the aws and
aws-s3 provider profiles and a declarative multi-output refresh model
(additional_outputs) so one AssumeRole co-mints all three credentials.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
2026-07-16 13:28:01 -07:00
st-grandClaude Opus 4.7 4ee27d9955 feat(sandbox,providers): add aws-bedrock as a recognized inference provider (#1704)
* feat(sandbox): allow AWS Bedrock InvokeModel paths through the L7 router

Adds two patterns to `default_patterns()` so the supervisor's L7
inference router recognizes the Bedrock InvokeModel URL shape and
forwards matched requests to the registered upstream:

- `POST /model/{modelId}/invoke`                       → aws_bedrock_invoke
- `POST /model/{modelId}/invoke-with-response-stream`  → aws_bedrock_invoke_stream

The `{modelId}` segment is wildcarded by extending `detect_inference_pattern`
to handle one middle `/*/` segment in addition to the existing trailing
`/*`. The wildcard is constrained to a single non-empty path segment to
avoid path-traversal liabilities — `/model//invoke` and `/model/a/b/invoke`
both no-match.

Without this, sandboxes running Claude Code in its native Bedrock mode
(`CLAUDE_CODE_USE_BEDROCK=1`, `ANTHROPIC_BEDROCK_BASE_URL`, AWS-style
auth) hit the supervisor with `403 connection not allowed by policy`
because their URL doesn't match `/v1/*` shapes. The fix unblocks
operators wanting to register direct AWS Bedrock, an in-cluster
Bedrock-compatible bridge, or a Bedrock-emulating LiteLLM as
`--type aws-bedrock` providers.

Tests cover: positive matches for invoke + invoke-with-response-stream,
query-string handling, GET rejection, empty-segment rejection,
multi-segment rejection, and unknown-action rejection.

Companion changes (provider discovery spec + YAML profile) follow in
the next commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* feat(providers): add aws-bedrock provider profile + discovery spec

Adds `aws-bedrock` to the built-in provider catalog so operators can
run `openshell provider create --type aws-bedrock --credential ...`
and have the gateway treat it as a first-class inference provider
alongside `anthropic`, `openai`, etc.

- `providers/aws-bedrock.yaml`: YAML profile declaring four credentials
  (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, AWS_REGION).
  Default endpoint is `bedrock-runtime.us-east-1.amazonaws.com:443`;
  operators in other regions or running against a Bedrock-compatible
  proxy override via the operator-supplied `BEDROCK_BASE_URL` config-key
  (mirrors `ANTHROPIC_BASE_URL` for the `anthropic` provider).

- `crates/openshell-providers/src/providers/aws_bedrock.rs`: the
  `ProviderDiscoverySpec` so `openshell provider create --auto-providers`
  picks up AWS_* env vars from local credentials.

- `crates/openshell-providers/src/providers/mod.rs`: register the module.

- `crates/openshell-providers/src/lib.rs`: register the SPEC in the
  default registry alongside the other providers.

- `crates/openshell-providers/src/profiles.rs`: include the new YAML in
  `BUILT_IN_PROFILE_YAMLS`.

What this PR explicitly does NOT add (intentionally separated for
review-size reasons; will follow up):

- A SigV4 signer in `openshell-router`. The current change simply
  declares the protocol; a follow-up PR adds outbound SigV4 signing
  using the `aws-sigv4` crate and a new `auth_style: sigv4` validator
  branch in profiles.rs. Operators who don't need SigV4 (e.g. an
  in-cluster bridge that ignores it and authenticates separately to
  the upstream) can use this PR today.

- Body translation between Bedrock InvokeModel shape and other
  inference shapes. The router treats Bedrock requests as opaque
  pass-through; if the operator's upstream is real AWS Bedrock it
  speaks Bedrock natively, if it's a translating bridge the bridge
  does any conversion server-side.

- `BEDROCK_BASE_URL` placeholder substitution in the YAML loader.
  Today the YAML's `host` is a literal default; operators override
  with the config-key the same way `ANTHROPIC_BASE_URL` works.

Tested: `cargo test -p openshell-providers` (35 tests green) and
`cargo test -p openshell-sandbox --lib l7::inference` (40 tests green
including the seven new aws_bedrock cases from the previous commit).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* revert(providers): drop legacy aws-bedrock SPEC, rely on v2 YAML profile

Addresses johntmyers's review on NVIDIA/OpenShell#1704: net-new
providers should land via the v2 YAML profile only and should NOT
require changes to the legacy `ProviderDiscoverySpec` registry.

- Delete `crates/openshell-providers/src/providers/aws_bedrock.rs`
  (the legacy SPEC + `test_discovers_env_credential!` invocation).
- Drop `pub mod aws_bedrock;` from `crates/openshell-providers/src/providers/mod.rs`.
- Drop `registry.register(providers::aws_bedrock::SPEC)` from
  `crates/openshell-providers/src/lib.rs`.

Kept:

- `providers/aws-bedrock.yaml` and the `include_str!` in
  `BUILT_IN_PROFILE_YAMLS` (`profiles.rs`) — the v2 path.
  `discover_from_profile()` (`crates/openshell-providers/src/discovery.rs`)
  picks up AWS_* env vars via `discovery.credentials` in the YAML.
- L7 router patterns in `crates/openshell-sandbox/src/l7/inference.rs`
  — orthogonal to the provider registry.

The discovery test in the deleted file goes with it; v2 doesn't have
an established per-provider env-var-pickup unit test pattern, and
other YAML-only registrations (none today, but this is the new
direction) won't carry one either.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(providers): include aws_session_token in discovery + update profile assertion

Two fixes from johntmyers's gator-agent re-check on NVIDIA/OpenShell#1704:

1. `providers/aws-bedrock.yaml`: add `aws_session_token` to
   `discovery.credentials`. The credential is declared in the profile
   but was missing from the discovery scan list, so Providers v2
   `--from-existing` would silently drop temporary AWS credentials
   (STS / IRSA scenarios).

2. `crates/openshell-server/src/grpc/provider.rs`: update the static
   `list_provider_profiles_returns_built_in_profile_categories`
   assertion to include `aws-bedrock` at alphabetical position 0.
   Adding `providers/aws-bedrock.yaml` to BUILT_IN_PROFILE_YAMLS made
   the prior `["claude-code", "github", "nvidia"]` expectation stale.

Remaining blockers from the same review (deferred to follow-up
commits): `inference::profile_for` registration for aws-bedrock,
user-facing provider + inference-routing docs, and an
`upsert_cluster_inference_route` integration test.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* feat(inference): register aws-bedrock profile (bridge-fronted) + docs

Addresses johntmyers's blocking review feedback on PR #1704:
"aws-bedrock still is not wired into the managed inference.local route
registry. profile_for only registers openai, anthropic, and nvidia, so
inference set --provider <aws-bedrock-provider> will reject this
provider before the new sandbox L7 patterns can be used."

Approach: register aws-bedrock as a *bridge-fronted* upstream — the
router does not inject any auth header on outbound requests; the
configured BEDROCK_BASE_URL is expected to point at a translating
bridge / Bedrock-compatible proxy that handles auth in its own pod.
This is the shape the L7 patterns commit (8b30211a) and the YAML
profile (6b51e1a6) were designed for. SigV4 signing for direct AWS
Bedrock is a separate follow-up; see PR thread.

Changes:

- core::inference::AuthHeader: add `None` variant for upstreams that
  authenticate themselves.
- core::inference: add AWS_BEDROCK_PROFILE static + register in
  profile_for. Default base URL is bedrock-runtime.us-east-1, override
  via BEDROCK_BASE_URL config-key (mirrors ANTHROPIC_BASE_URL pattern).
  Empty credential_key_names + auth: None means no router-side
  credential lookup at route time.
- router::backend: handle AuthHeader::None as a no-op (skip auth
  injection).
- server::inference::resolve_provider_route: gate find_provider_api_key
  on auth != None. aws-bedrock providers with empty credentials now
  resolve cleanly. Updated the unsupported-type error message to
  include aws-bedrock in the supported list.
- server::inference tests: add positive
  upsert_cluster_route_succeeds_for_aws_bedrock_without_api_key test
  covering the new code path end-to-end (provider with empty creds +
  BEDROCK_BASE_URL config → upsert succeeds → resolved route has
  empty api_key + provider_type aws-bedrock + bridge URL).
- core::inference tests: profile_for_known_types covers aws-bedrock,
  case-insensitive lookup, plus three new aws-bedrock-specific tests
  (auth: None, no credential keys, bedrock-specific protocols).
- docs/sandboxes/inference-routing.mdx: header forwarding row
  mentions aws-bedrock has no passthrough headers; new tabs in
  Supported API Patterns (InvokeModel + InvokeModelWithResponseStream)
  and Create a Provider (with the bridge-fronted shape note + SigV4
  deferral).
- docs/sandboxes/manage-providers.mdx: new row in Supported Provider
  Types table; new row in Supported Inference Providers table.

Verification (in dev container):
- cargo check -p openshell-core -p openshell-router -p openshell-server: clean
- cargo test -p openshell-core --lib inference: 14/14 pass (incl. 3 new)
- cargo test -p openshell-server --lib inference::tests::upsert: 6/6 pass
  (incl. new aws-bedrock test)
- cargo fmt --check: clean
- cargo clippy --all-targets -D warnings: clean

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(aws-bedrock): bridge-only YAML, doc actual cmd shape, neg test

Addresses four findings from gator-agent's #1704 re-check on 4ab587f1:

- **Item 5** (YAML collects unused AWS creds): mark all four AWS
  credentials `required: false` and clear `discovery.credentials`.
  Bridge-fronted routing intentionally does not consume AWS
  credentials, so `--from-existing` no longer scans for them. The
  credentials remain in the schema (not deleted) so the SigV4
  follow-up can flip them back without a schema migration. Added a
  multi-line description that names the bridge-fronted shape and the
  SigV4 deferral so readers don't have to cross-reference the PR
  thread.

- **Item 3** (docs show command that the CLI rejects): rewrite the
  Create-a-Provider example for AWS Bedrock to use the actual
  required shape — placeholder `--credential AWS_ACCESS_KEY_ID=
  unused-bridge-fronted-shape` plus the `--config BEDROCK_BASE_URL`.
  The placeholder satisfies the gRPC handler's
  `provider.credentials.is_empty()` rejection without expanding
  server-side validation; the router ignores it on the outbound path
  because `auth: AuthHeader::None` skips header injection. Operators
  see a clearly-labeled placeholder in `provider get` output.

- **Item 1** (validator probe): document `--no-verify` as required
  for `openshell inference set --provider <aws-bedrock>` since the
  default validation probe doesn't recognize the
  `aws_bedrock_invoke` / `aws_bedrock_invoke_stream` protocols. Doc
  now shows the full `provider create` + `inference set --no-verify`
  flow with rationale for both decisions inline.

- **Item 6** (docs polish): `inference-routing.mdx` summary row now
  lists AWS Bedrock alongside NVIDIA, Anthropic, Vertex AI, and
  OpenAI-compatible providers, with the bridge-fronted caveat
  inline.

Test additions in `crates/openshell-server/src/inference.rs`:

- Renamed the existing aws-bedrock test from
  `..._without_api_key` to `..._with_bridge_url` and updated it to
  use a placeholder credential (mirroring the doc-recommended
  pattern operators will copy-paste). The `auth: None` path still
  produces an empty `api_key` on the resolved route — the test now
  documents that the credential is *stored* but not *used*.
- Added `upsert_cluster_route_rejects_aws_bedrock_without_bedrock_base_url`:
  the negative half of johntmyers' "successfully used by
  upsert_cluster_inference_route or intentionally rejected with a
  clear documented error" ask. With
  `default_base_url: ""` and no `BEDROCK_BASE_URL` config, route
  resolution returns `InvalidArgument` naming the missing base_url
  rather than silently forwarding prompts to AWS Bedrock with no
  usable auth.

Verification (in dev container):
- cargo test -p openshell-core --lib inference: 18/18 (incl. 3 new)
- cargo test -p openshell-server --lib inference::tests::upsert: 8/8
  (incl. 2 new aws-bedrock cases — positive + negative)
- cargo fmt --check: clean
- cargo clippy --all-targets -D warnings: clean

Item 2 (router-side enforcement of operator-configured Bedrock model
path, replacing the current verbatim path forwarding + body-only
model rewrite) is the remaining blocker and is genuinely separable —
it touches the L7 router with streaming-aware test coverage.
Deferring to its own commit so the security-critical change gets the
review attention it deserves.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(router): enforce operator-configured Bedrock model in request path

Closes the security-blocking item from gator-agent's #1704 re-check
on 4ab587f1: "Bedrock carries the model id in /model/{modelId}/invoke,
but the router currently forwards the caller's original path and only
rewrites JSON body model. That lets sandbox code choose a different
upstream model than the operator-configured route model, and may also
mutate native Bedrock request bodies incorrectly."

Two changes in `prepare_backend_request`:

1. **Path rewrite for Bedrock routes.** Before computing the upstream
   URL, parse the inbound path's `/model/<id>/invoke[-with-response-stream]`
   shape and substitute the operator-configured `route.model` for the
   caller-supplied model segment. Sandbox code that hardcodes a
   different model still works (we don't reject on mismatch), but the
   operator's configured model is what reaches the upstream / bridge.
   If the inbound path is somehow not a recognized Bedrock shape on a
   Bedrock route (the L7 pattern detector upstream of the router
   should never produce this combination), reject with
   RouterError::Internal naming the offending path rather than
   forwarding verbatim.

2. **Skip body-model injection for Bedrock routes.** The existing body
   rewriter unconditionally inserts `route.model` into the JSON body
   for non-Vertex routes. AWS Bedrock InvokeModel encodes the model
   in the URL path; the body is the raw provider-specific payload
   (Anthropic Messages for Claude, Mistral payload for Mistral, etc.)
   and must not be mutated. The branch ordering is now:
   needs_vertex_anthropic_version → strip body model + inject
   anthropic_version; route_is_bedrock → leave body alone; else →
   inject route.model (existing default).

New helpers, all in `crates/openshell-router/src/backend.rs`:

- `route_is_bedrock(route)` — true when route.protocols contains
  aws_bedrock_invoke or aws_bedrock_invoke_stream.
- `parse_bedrock_invocation_path(path)` — returns
  Some((model_id, "/invoke" | "/invoke-with-response-stream")) for
  paths matching the recognized Bedrock shapes. Strips query strings.
  Rejects empty model ids and multi-segment ids (defense-in-depth
  matching the L7 pattern detector's existing guards).
- `rewrite_bedrock_path(route, path)` — returns the path with the
  caller's model segment replaced by route.model.

Test coverage in the same file (9 new tests):

- parse_bedrock_invocation_path: positive cases for both invoke
  variants, query-string stripping; negative cases for empty model id,
  multi-segment id, unknown action, wrong prefix, missing slash.
- route_is_bedrock: matches both protocol variants singly and
  combined; rejects openai_chat_completions.
- rewrite_bedrock_path: substitutes operator model on both invoke
  variants; returns None for non-Bedrock paths.
- bedrock_route_rewrites_model_in_path_and_preserves_body
  (wiremock end-to-end): caller sends /model/some-other-model/invoke
  with a body containing model: "caller-supplied-model-name". Mock
  asserts the upstream receives /model/<operator-model>/invoke and the
  body's model field is the caller's value (NOT route.model) — proves
  both the path rewrite and the body preservation.
- bedrock_route_streaming_rewrites_model_in_path: same contract for
  invoke-with-response-stream.
- bedrock_route_rejects_non_bedrock_path: defense-in-depth coverage of
  the Internal-error path when a Bedrock route receives a path that
  doesn't match Bedrock shape.

Verification (in dev container):
- cargo test -p openshell-router --lib: 53/53 (incl. 9 new)
- cargo fmt --check: clean
- cargo clippy -p openshell-core -p openshell-router -p openshell-server
  --all-targets -- -D warnings: clean

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(sandbox/l7): declare framing for Bedrock patterns

Closes the buffered-vs-streaming framing warning from gator-agent's
re-check on 4ab587f1: "Bedrock InvokeModel should be buffered while
InvokeModelWithResponseStream is streaming. Please add framing/coverage
so /model/{id}/invoke cannot be corrupted by the streaming proxy's
truncation/error-frame behavior."

The InferenceApiPattern struct gained a `framing: ResponseFraming`
field upstream after the original Bedrock-patterns commit (#22b78cff)
landed; the cherry-pick onto current upstream/main left the two
Bedrock entries without the new field. Fixed here:

- aws_bedrock_invoke (POST /model/{id}/invoke):
    framing = ResponseFraming::Buffered
  InvokeModel returns one JSON object the caller decodes whole. Sending
  it through the streaming proxy would risk a mid-body size-cap
  truncation or idle-timeout failure appending an SSE error event onto
  bytes the caller decodes as one JSON body — the same corruption mode
  that drove the existing embeddings + model-discovery to Buffered.
- aws_bedrock_invoke_stream (POST /model/{id}/invoke-with-response-stream):
    framing = ResponseFraming::Streaming
  InvokeModelWithResponseStream returns an AWS event-stream of binary
  chunks; the caller wants chunks incrementally, so the streaming proxy
  path is correct.

Two new tests in `crates/openshell-sandbox/src/l7/inference.rs` pin
down the contract:

- aws_bedrock_invoke_is_buffered — detect_inference_pattern returns a
  Buffered pattern for /model/<id>/invoke, with explanatory message
  naming the corruption mode being prevented.
- aws_bedrock_invoke_stream_is_streaming — same shape, asserting
  Streaming for /model/<id>/invoke-with-response-stream.

Verification (in dev container):
- cargo check -p openshell-sandbox: clean (was failing on missing
  `framing` field before this commit)
- cargo test -p openshell-sandbox --lib l7::inference::tests::aws_bedrock:
  7/7 (incl. 2 new framing tests)
- cargo fmt --check: clean
- cargo clippy --all-targets -- -D warnings: clean

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(inference): drop aws_bedrock_invoke_stream until protocol-aware errors land

Per PR #1704 review (johntmyers): defer Bedrock streaming to a follow-up
that also wires protocol-aware error framing for the AWS event-stream
shape. Until then, surfacing `/model/{id}/invoke-with-response-stream`
risks shipping responses the sandbox cannot interpret on failure.

- AWS_BEDROCK_PROTOCOLS no longer advertises `aws_bedrock_invoke_stream`.
- The L7 inference pattern table drops the streaming entry; only
  `aws_bedrock_invoke` (buffered) is recognized.
- Test `aws_bedrock_invoke_stream_pattern_is_deferred` asserts no
  pattern claims that protocol so the gap is visible.

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(inference): validate AWS Bedrock model_id at route + rewrite, preserve query

Per PR #1704 review (johntmyers, gator agent): the operator-configured
Bedrock model_id flows verbatim into the upstream URL path; the previous
plumbing left no enforcement that the value was a single benign path
segment, opening a path-injection vector.

Defense in depth, both layers:

* `openshell-server::inference`: new `validate_aws_bedrock_model_id`
  rejects empty, leading/trailing whitespace, `/`, `\\`, `?`, `#`, `%`,
  `..`, control/whitespace characters. Wired into `resolve_provider_route`
  ahead of base_url resolution so the route store cannot persist a
  malformed model_id. Mirrors `validate_vertex_model_id` exactly.

* `openshell-router::backend`: `rewrite_bedrock_path` now refuses to
  construct the upstream URL unless `route.model` passes
  `is_valid_bedrock_model_id`, so even a stale or hand-edited route
  cannot reach the wire. The parser also drops the
  `/invoke-with-response-stream` arm to match the protocol catalog.

* `parse_bedrock_invocation_path` returns the `?`-prefixed query tail
  as a third element; `rewrite_bedrock_path` re-attaches it so any
  caller-supplied query string is preserved through the model rewrite.

Tests: 5 unit tests for the validator, 1 integration test that
exercises every unsafe-model_id reject path through
`upsert_cluster_inference_route`, plus a router-side rewrite-rejects
test covering 11 unsafe `route.model` values. All 72 server inference
tests + router tests pass.

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* fix(aws-bedrock): make built-in profile non-egress-granting; tighten docs

Per PR #1704 review (johntmyers): a single profile that is "usable as
is" — not a profile that auto-grants direct AWS Bedrock egress to any
sandbox that selects it. The bridge IS the egress point and is
operator-managed; the profile must not implicitly punch a hole through
the cluster's network policy on its behalf.

* `providers/aws-bedrock.yaml`: clear `endpoints` and `binaries` to
  empty arrays. Rewrite the description to spell out that the profile
  is intentionally non-egress-granting, that operators are responsible
  for declaring their bridge's egress endpoint and binary attribution,
  and that the SigV4 follow-up will repopulate these fields once
  router-side signing exists.

* `docs/sandboxes/inference-routing.mdx`:
  - Drop the `InvokeModelWithResponseStream` row from the supported
    patterns table (matches the protocol-catalog + L7-pattern drop).
  - Update the `--no-verify` paragraph to reference only
    `aws_bedrock_invoke`.
  - Generalise the placeholder-credential rationale: any
    standalone-router profile registering `AuthHeader::None` will hit
    the same non-empty-credentials structural requirement.

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* style(router): satisfy cargo fmt for Bedrock query-tail extraction

CI's rustfmt rejected the one-line `let (path_only, query_tail) = path.find('?').map_or(...)`
shape from commit c51160e4 and required the chained-method layout
instead. Functional behaviour is unchanged; tests still pass.

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* style(router): drop unnecessary `crate::` prefix on RouterError matcher

`RouterError` is already imported at the top of the file, so the
two test-side `matches!(result, Err(crate::RouterError::UpstreamProtocol(_)))`
uses trip clippy's `unused_qualifications` under `-D warnings`. Drop
the prefix on both sites; functional behaviour is unchanged.

These warnings predate this PR (originated in 25abc9e3c on 2026-06-07)
but NVIDIA's `rust:lint` re-runs because this PR touches
`backend.rs`, so the lint regression surfaces here.

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

* style(server): silence unused-variable lint on first shutdown_tx

`crates/openshell-server/src/lib.rs` declares `(shutdown_tx, shutdown_rx)`
twice in `run_server` — the first pair (introduced by upstream #1577
reconciler-lease work) only consumes `shutdown_rx`, leaving the first
`shutdown_tx` unused. Clippy's `-D warnings` flags it on workspace
lint. Rebasing this Bedrock-scoped PR onto current upstream/main
exposes the regression because we re-touch the file and the lint
re-runs.

Prefix the unused half with an underscore so the workspace lint is
clean. Behaviour is unchanged — only the second shutdown_tx (line
~425) ever sends a shutdown signal today.

This is a drive-by upstream fix unrelated to the Bedrock provider
work; keeping it separate so it can be cherry-picked or reverted
independently of the Bedrock commits.

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>

---------

Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-23 13:10:37 -07:00
Robert Sturla 48545cfb59 feat(sandbox): add GCE metadata emulator for Google Cloud (#1763)
* feat(core): add shared GCP constants module

Single source of truth for GCP naming: env var aliases, provider config
keys, token search order, and Vertex-specific env vars. Consumed by
openshell-server, openshell-providers, and openshell-sandbox.

- Add google_cloud.rs with metadata emulator host and loopback address
- Define PROJECT_ID, REGION, and SERVICE_ACCOUNT_EMAIL env var aliases
- Add provider config key constants for gcp provider implementations
- Define TOKEN_ENV_KEYS search order (SA token takes priority over ADC)
- Add Vertex-specific env vars for Goose and Claude Code SDK integration
- Add STATIC_CONFIG_KEYS as union of all alias arrays for env resolution
- Export module via openshell-core lib.rs

Signed-off-by: Robert Sturla <rsturla@redhat.com>

* feat(providers): add google-cloud and vertex provider plugins

Add GoogleCloudProvider and VertexProvider implementing inject_env to
project GCP config (project ID, region, SA email, metadata host) into
sandbox environment variables. Replace the inline Vertex AI env
injection in the server with the registry-based inject_env dispatch.

Also adds the google-cloud.yaml provider profile with SA JWT and ADC
OAuth2 credential refresh flows.

Signed-off-by: Robert Sturla <rsturla@redhat.com>

* feat(sandbox): add GCE metadata emulator for GCP

Add a loopback HTTP server on 127.0.0.1:8174 inside the sandbox
network namespace that emulates the GCE instance metadata API.
GCP client SDKs discover it via GCE_METADATA_HOST and obtain
credential placeholders that the proxy resolves to real tokens
at egress.

Add metadata_server module with MetadataHandler trait and
netns-aware TCP binding via std::thread (not spawn_blocking)
to avoid tokio pool namespace contamination
Add google_cloud_metadata module implementing the GCE metadata
API subset (token, project-id, email, scopes, service-accounts)
Add child_env_resolved() and gcp_token_response() to
ProviderCredentialState for GCP-aware credential projection
Wire metadata server into sandbox lifecycle before SSH handler
Collapse multi-line HTTP response format string into single line

Signed-off-by: Robert Sturla <rsturla@redhat.com>

* docs(sandbox): add GCP credentials documentation

Document the google-cloud provider setup for ADC and service account
flows, injected environment variables, metadata emulator behavior, and
network policy configuration for GCP APIs.

Signed-off-by: Robert Sturla <rsturla@redhat.com>

* feat(cli): support --from-gcloud-adc for google-cloud providers

Widen --from-gcloud-adc to accept google-cloud providers. The ADC
credential key is derived from the provider profile rather than
hardcoded per type, so future GCP provider types get ADC support by
declaring the right refresh metadata in their profile YAML.

Add ProviderTypeProfile::adc_credential() to find the ADC-compatible
credential from a profile's refresh metadata. Remove unused
VERTEX_AI_ADC_TOKEN_KEY and GCP_ADC_TOKEN_KEY constants.

Signed-off-by: Robert Sturla <rsturla@redhat.com>

---------

Signed-off-by: Robert Sturla <rsturla@redhat.com>
2026-06-23 15:05:07 -05:00
mmilutinovic371 36bb9e3ec2 feat(providers): add DeepInfra as a built-in inference provider (#1902)
* feat(providers): add DeepInfra as a built-in inference provider (v2 only)

- Adds `deepinfra` as a built-in Providers v2 profile (`providers/deepinfra.yaml`)
  with inference category, Bearer auth, and `DEEPINFRA_API_KEY` discovery
- Adds `DEEPINFRA_PROFILE` to inference routing so `inference.local` works
  with the `deepinfra` provider type
- Fixes `build_backend_url` to strip `/v1` from request paths when the base
  URL contains `/v1/` as an internal segment (e.g. `api.deepinfra.com/v1/openai`),
  preventing double-versioned paths like `.../v1/openai/v1/chat/completions`
- Updates `docs/sandboxes/providers-v2.mdx` and `docs/sandboxes/manage-providers.mdx`
  with DeepInfra entries; removes the old v1 workaround row that used `openai`
  type with `OPENAI_API_KEY`

Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com>

* fix(providers): address gator review findings for DeepInfra provider

- Narrow build_backend_url /v1 dedupe to URLs whose path component is
  exactly /v1 or starts with /v1/ — prevents regression on proxy
  endpoints where /v1 is buried deeper (e.g. /api/v1/openai); add
  regression test for the nested proxy path case
- Add deepinfra provider plugin with DEEPINFRA_API_KEY discovery,
  registered in ProviderRegistry so known_types() and TUI include it
- Add deepinfra to unsupported-inference-provider error message in
  openshell-server for accurate user-facing debugging guidance
- Add deepinfra to openai_compatible_profiles_include_embeddings test
  to lock in the OpenAI-compatible protocol contract

Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com>

* fix(router): handle /v1 as final path segment in build_backend_url dedup

Extends the /v1 deduplication logic to also strip /v1 from request paths
when the base URL's path ends with /v1 (e.g. https://api.groq.com/openai/v1).
The previous fix only matched paths starting with /v1/, which regressed
providers like Groq whose base path has /v1 as the last segment rather than
the first. The nested-proxy exclusion (e.g. /api/v1/openai) is preserved
since /v1 appears in the middle — neither first nor last segment. Adds a
regression test for the Groq-style base URL.

Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com>

* fix(providers): add deepinfra telemetry bucket and update profile list test

- Add DeepInfra variant to ProviderProfile telemetry enum and from_raw()
  mapping so deepinfra providers are tracked in their own bucket rather
  than falling through to Custom
- Map deepinfra in telemetry_provider_profile() in openshell-server
- Add deepinfra to list_provider_profiles_returns_built_in_profile_categories
  test (sorted between cursor and github)
- Update architecture/gateway.md inference provider list to include deepinfra

Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com>

* style(router): apply cargo fmt to backend.rs

Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com>

---------

Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com>
2026-06-16 16:39:03 -07:00
John T. Myers 62c421b21e feat(providers): add profile-backed policy visibility (#1640)
* chore: wip providers v2 tui and codex profile

* chore: wip effective policy get and codex profile

* chore: wip provider profiles and tui detail views

* feat(tui): annotate policy proposal review status
2026-06-02 17:39:01 -07:00
Adam Miller f061b1d923 feat(providers): add Google Vertex AI inference provider (#1568)
* feat(providers): add Google Vertex AI provider

Adds Vertex AI provider profiles, routing, credential refresh plumbing, CLI support, docs, and regression coverage. Keeps the related NETLINK_ROUTE seccomp allowance needed by Vertex client tooling that calls getifaddrs.

* docs: add Vertex AI sandbox usage for Claude Code and OpenCode

Cover the full end-to-end setup for running Claude Code and OpenCode
inside an OpenShell sandbox via inference.local with a Vertex AI backend:

- google-vertex-ai.mdx: add 'Use from a Sandbox' section with tabbed
  examples for Claude Code (--bare flag, no /v1 suffix) and OpenCode
  (/v1 suffix required). Add providers_v2_enabled prerequisite and
  --no-verify note for global region. Document policy proposals table
  covering metadata.google.internal (always blocked), downloads.claude.ai,
  and storage.googleapis.com.

- inference-routing.mdx: expand 'Use the Local Endpoint' section with
  tabbed examples for Claude Code, OpenCode, Python OpenAI SDK, and
  Python Anthropic SDK. Add notes explaining the /v1 path suffix
  difference between clients.

- supported-agents.mdx: update Claude Code and OpenCode rows to mention
  inference.local support and correct base URL requirements.

* fix: address vertex review findings

* test(sandbox): retry on spurious Ok in fork-exec ambiguity test

On arm64 under heavy CI load, the /proc fd scan in
find_socket_inode_owners can transiently miss the parent process's
socket fd entry, returning only the child as an owner. This causes
resolve_process_identity to return Ok (single owner, no ambiguity
check fires) instead of the expected ambiguous-ownership Err.

Extend the retry loop to also handle unexpected Ok results, mirroring
the existing retry for transient Err results. 10 retries at 50ms gives
a 500ms settling window, which is sufficient for procfs to stabilize
on loaded arm64 runners.

* fix: address vertex review regressions

* docs(router): clarify stream_response semantics for Vertex rawPredict routing

Document the three call sites of prepare_backend_request and their
stream_response values in a caller table:

- send_backend_request: false → :rawPredict (unary endpoint)
- send_backend_request_streaming: true → :streamRawPredict
- verify_backend_endpoint: explicitly false to probe the unary endpoint

Cross-reference the table from build_provider_url and
is_vertex_anthropic_rawpredict_route so the stream_response=true guard
in the suffix upgrade branch is understood in full context.

Also note that is_vertex_anthropic_rawpredict_route is a structural
predicate (model_in_path + anthropic_messages + :rawPredict suffix),
not a named-provider check, so any future provider with the same route
shape inherits the transforms automatically.
2026-06-02 10:45:50 -05:00
Alexander Watson e98ea3ee93 feat(policy): add agentic approval loop (#1528) 2026-05-29 21:11:31 -07:00
John T. Myers 0cef26521a feat(providers): derive discovery from profiles (#1503)
* feat(providers): derive discovery from profiles

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(providers): keep v2 discovery profile-only

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* docs(providers): update providers v2 behavior

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

* fix(providers): make github profile read-only

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>

---------

Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
2026-05-22 08:54:15 -07:00
John T. Myers d255cdd9c9 feat(providers): add credential refresh foundation (#1349)
* feat(providers): add credential refresh foundation

* feat(providers): mint oauth refresh token credentials

* fix(providers): allow delegated refresh bootstrap

* fix(providers): guard refresh credential modes

* fix(providers): refine refresh lifecycle UX

* fix(cli): restore provider test gateway response import

* fix(server): clean auth endpoint test qualifications

* fix(providers): tighten refresh authorization and collisions

* chore(providers): trim bundled v2 profiles

* fix(providers): resolve refresh rebase fallout

* fix(cli): accept rfc3339 credential expiry

* test(providers): update inferred claude provider type

* test(providers): avoid removed outlook default profile

* test(providers): isolate attach limit fixtures
2026-05-19 07:42:10 -07:00
John T. Myers 043bde279a feat(providers): add profile-backed policy composition (#1037)
Foundation for providers v2. Add provider profiles and provider profile composition with user policies.
2026-05-04 18:34:33 -07:00