mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 08:28:19 +08:00
codex/mxc-https-test-socket-owner
15
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
07df822090 |
feat(providers): make profiles authoritative (#2962)
* feat(providers): make profiles authoritative Closes #1988 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): move profiles into provider navigation Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(providers): clarify provider attachment lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(tui): scroll provider profile picker Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): honor profile credential semantics Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): prefer exact profile IDs Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): harden authoritative profile adoption Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(oidc): align provider fixtures with profiles Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(providers): preserve authoritative profile lifecycle Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
9505ca5ed1 |
chore: remove Bazel build support (#2840)
Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
0d708d6d51 |
fix(policy): gate uninspected credentialed endpoints (#2493)
* fix(policy): gate uninspected credentialed endpoints Signed-off-by: Adrien Langou <alangou@nvidia.com> * refactor(cli): extract allowed-ip option parsing Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(policy): gate endpointless credential bindings Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
1959ea19be |
build(bazel): establish RFC 0012 Rust reference graph (#2414)
Establish the Phase 1 reference implementation proposed by RFC 0012 while retaining the existing Cargo and Mise workflows during evaluation. - pin Bazel 9.1.1 and configure Bzlmod, rules_rs, LLVM, protobuf, and Rust 1.95 toolchains - import third-party crates from Cargo metadata and propagate the workspace version into Bazel targets - add library, binary, proc-macro, unit-test, and integration-test targets across the supported Rust workspace crates and drivers - generate protobuf Rust sources and descriptor sets under Bazel while preserving Cargo-compatible generated-code imports - annotate aws-lc-sys and zstd-sys native dependencies, build Z3 4.15.2 from source, and generate z3-sys bindings - make CLI and procfs test fixtures available as explicit Bazel inputs without relying on fixed host binary paths - define optimized release targets for Linux x86_64 and aarch64 CLI, sandbox, and gateway binaries, plus macOS aarch64 artifacts - add Bazel, buildifier, and lcov to the Nix development environment RFC: 0012 (rfc12 branch) Refs: #2491 Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
9377e0d5fe |
fix(providers): allow git clone/fetch via default GitHub provider (#2317)
* fix(providers): allow git clone/fetch via default GitHub provider The github.com:443 git-transport endpoint used the read-only access preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied those operations and `gh repo clone` / `git clone https://...` failed. Replace the preset with explicit rules that permit the read-only methods plus POST */git-upload-pack, so clone/fetch work while push (git-receive-pack) stays blocked. Enabling push still requires an explicit policy proposal. Why allowing this POST is still read-only: in git's smart HTTP protocol POST is an RPC transport, not a write. A clone/fetch does GET */info/refs (ref discovery) followed by POST */git-upload-pack, whose body is only the client's want/have negotiation; the server responds with a packfile and nothing on the server is modified (data flows server -> client). The service names are from the server's perspective: git-upload-pack = the server uploads a pack to the client (a read/ download), while git-receive-pack = the server receives a pack from the client (the actual write/push). The new rule is scoped to */git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs to github.com remain denied. Add a provider-profile regression test and a rego enforcement test covering ref discovery, upload-pack (allowed), and receive-pack (denied). Closes #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): strengthen git-transport regression and add clone e2e Pin the exact allowed rule set for the built-in github git-transport endpoint in both the provider-profile and composed-policy tests, so a broader or additional POST rule (e.g. POST **) that could enable push via git-receive-pack fails the test instead of passing a substring check. Add an e2e test that attaches the built-in github provider and clones a public repo over HTTPS, exercising provider attachment, effective-policy composition, TLS interception, and real git behavior. Update the Providers V2 docs so the github.com git-transport endpoint shows explicit clone/fetch rules instead of the stale read-only preset. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): isolate providers_v2 mutation in clone e2e The clone e2e enables the gateway-global providers_v2_enabled setting. Restore its exact prior value (or absence) captured via GetGatewayConfig instead of unconditionally deleting it, and serialize the mutation across xdist workers with an exclusive file lock on the run's shared base temp dir, so a shared or pre-configured gateway is left untouched and parallel workers cannot race the read-modify-restore. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> * test(providers): serialize providers_v2 mutation with a suite-wide guard The clone e2e's per-fixture lock only coordinated fixtures that acquired it; other xdist workers hit the same gateway without it and could observe the transiently-enabled providers_v2_enabled global during their own sandbox creation (CWE-362). Add an autouse readers-writer guard in conftest: every test holds a shared lock on the gateway config, and a test marked exclusive_gateway_config holds an exclusive lock. Mark the clone test exclusive so no other worker is mid-test while it enables and restores the gateway-global setting. Exact prior-value restoration is retained. Refs #1769 Signed-off-by: Russell Bryant <rbryant@redhat.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
aa483ecb9a |
feat(providers): AWS STS AssumeRole refresh strategy and aws-s3 profile (#1782)
Add gateway-managed AWS STS credential refresh (provider-v2, #1576). The gateway calls sts:AssumeRole and writes three short-lived credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the provider record; the proxy re-signs requests with SigV4. Adds the aws and aws-s3 provider profiles and a declarative multi-output refresh model (additional_outputs) so one AssumeRole co-mints all three credentials. Signed-off-by: Russell Bryant <rbryant@redhat.com> |
||
|
|
4ee27d9955 |
feat(sandbox,providers): add aws-bedrock as a recognized inference provider (#1704)
* feat(sandbox): allow AWS Bedrock InvokeModel paths through the L7 router
Adds two patterns to `default_patterns()` so the supervisor's L7
inference router recognizes the Bedrock InvokeModel URL shape and
forwards matched requests to the registered upstream:
- `POST /model/{modelId}/invoke` → aws_bedrock_invoke
- `POST /model/{modelId}/invoke-with-response-stream` → aws_bedrock_invoke_stream
The `{modelId}` segment is wildcarded by extending `detect_inference_pattern`
to handle one middle `/*/` segment in addition to the existing trailing
`/*`. The wildcard is constrained to a single non-empty path segment to
avoid path-traversal liabilities — `/model//invoke` and `/model/a/b/invoke`
both no-match.
Without this, sandboxes running Claude Code in its native Bedrock mode
(`CLAUDE_CODE_USE_BEDROCK=1`, `ANTHROPIC_BEDROCK_BASE_URL`, AWS-style
auth) hit the supervisor with `403 connection not allowed by policy`
because their URL doesn't match `/v1/*` shapes. The fix unblocks
operators wanting to register direct AWS Bedrock, an in-cluster
Bedrock-compatible bridge, or a Bedrock-emulating LiteLLM as
`--type aws-bedrock` providers.
Tests cover: positive matches for invoke + invoke-with-response-stream,
query-string handling, GET rejection, empty-segment rejection,
multi-segment rejection, and unknown-action rejection.
Companion changes (provider discovery spec + YAML profile) follow in
the next commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* feat(providers): add aws-bedrock provider profile + discovery spec
Adds `aws-bedrock` to the built-in provider catalog so operators can
run `openshell provider create --type aws-bedrock --credential ...`
and have the gateway treat it as a first-class inference provider
alongside `anthropic`, `openai`, etc.
- `providers/aws-bedrock.yaml`: YAML profile declaring four credentials
(AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, AWS_REGION).
Default endpoint is `bedrock-runtime.us-east-1.amazonaws.com:443`;
operators in other regions or running against a Bedrock-compatible
proxy override via the operator-supplied `BEDROCK_BASE_URL` config-key
(mirrors `ANTHROPIC_BASE_URL` for the `anthropic` provider).
- `crates/openshell-providers/src/providers/aws_bedrock.rs`: the
`ProviderDiscoverySpec` so `openshell provider create --auto-providers`
picks up AWS_* env vars from local credentials.
- `crates/openshell-providers/src/providers/mod.rs`: register the module.
- `crates/openshell-providers/src/lib.rs`: register the SPEC in the
default registry alongside the other providers.
- `crates/openshell-providers/src/profiles.rs`: include the new YAML in
`BUILT_IN_PROFILE_YAMLS`.
What this PR explicitly does NOT add (intentionally separated for
review-size reasons; will follow up):
- A SigV4 signer in `openshell-router`. The current change simply
declares the protocol; a follow-up PR adds outbound SigV4 signing
using the `aws-sigv4` crate and a new `auth_style: sigv4` validator
branch in profiles.rs. Operators who don't need SigV4 (e.g. an
in-cluster bridge that ignores it and authenticates separately to
the upstream) can use this PR today.
- Body translation between Bedrock InvokeModel shape and other
inference shapes. The router treats Bedrock requests as opaque
pass-through; if the operator's upstream is real AWS Bedrock it
speaks Bedrock natively, if it's a translating bridge the bridge
does any conversion server-side.
- `BEDROCK_BASE_URL` placeholder substitution in the YAML loader.
Today the YAML's `host` is a literal default; operators override
with the config-key the same way `ANTHROPIC_BASE_URL` works.
Tested: `cargo test -p openshell-providers` (35 tests green) and
`cargo test -p openshell-sandbox --lib l7::inference` (40 tests green
including the seven new aws_bedrock cases from the previous commit).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* revert(providers): drop legacy aws-bedrock SPEC, rely on v2 YAML profile
Addresses johntmyers's review on NVIDIA/OpenShell#1704: net-new
providers should land via the v2 YAML profile only and should NOT
require changes to the legacy `ProviderDiscoverySpec` registry.
- Delete `crates/openshell-providers/src/providers/aws_bedrock.rs`
(the legacy SPEC + `test_discovers_env_credential!` invocation).
- Drop `pub mod aws_bedrock;` from `crates/openshell-providers/src/providers/mod.rs`.
- Drop `registry.register(providers::aws_bedrock::SPEC)` from
`crates/openshell-providers/src/lib.rs`.
Kept:
- `providers/aws-bedrock.yaml` and the `include_str!` in
`BUILT_IN_PROFILE_YAMLS` (`profiles.rs`) — the v2 path.
`discover_from_profile()` (`crates/openshell-providers/src/discovery.rs`)
picks up AWS_* env vars via `discovery.credentials` in the YAML.
- L7 router patterns in `crates/openshell-sandbox/src/l7/inference.rs`
— orthogonal to the provider registry.
The discovery test in the deleted file goes with it; v2 doesn't have
an established per-provider env-var-pickup unit test pattern, and
other YAML-only registrations (none today, but this is the new
direction) won't carry one either.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(providers): include aws_session_token in discovery + update profile assertion
Two fixes from johntmyers's gator-agent re-check on NVIDIA/OpenShell#1704:
1. `providers/aws-bedrock.yaml`: add `aws_session_token` to
`discovery.credentials`. The credential is declared in the profile
but was missing from the discovery scan list, so Providers v2
`--from-existing` would silently drop temporary AWS credentials
(STS / IRSA scenarios).
2. `crates/openshell-server/src/grpc/provider.rs`: update the static
`list_provider_profiles_returns_built_in_profile_categories`
assertion to include `aws-bedrock` at alphabetical position 0.
Adding `providers/aws-bedrock.yaml` to BUILT_IN_PROFILE_YAMLS made
the prior `["claude-code", "github", "nvidia"]` expectation stale.
Remaining blockers from the same review (deferred to follow-up
commits): `inference::profile_for` registration for aws-bedrock,
user-facing provider + inference-routing docs, and an
`upsert_cluster_inference_route` integration test.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* feat(inference): register aws-bedrock profile (bridge-fronted) + docs
Addresses johntmyers's blocking review feedback on PR #1704:
"aws-bedrock still is not wired into the managed inference.local route
registry. profile_for only registers openai, anthropic, and nvidia, so
inference set --provider <aws-bedrock-provider> will reject this
provider before the new sandbox L7 patterns can be used."
Approach: register aws-bedrock as a *bridge-fronted* upstream — the
router does not inject any auth header on outbound requests; the
configured BEDROCK_BASE_URL is expected to point at a translating
bridge / Bedrock-compatible proxy that handles auth in its own pod.
This is the shape the L7 patterns commit (8b30211a) and the YAML
profile (6b51e1a6) were designed for. SigV4 signing for direct AWS
Bedrock is a separate follow-up; see PR thread.
Changes:
- core::inference::AuthHeader: add `None` variant for upstreams that
authenticate themselves.
- core::inference: add AWS_BEDROCK_PROFILE static + register in
profile_for. Default base URL is bedrock-runtime.us-east-1, override
via BEDROCK_BASE_URL config-key (mirrors ANTHROPIC_BASE_URL pattern).
Empty credential_key_names + auth: None means no router-side
credential lookup at route time.
- router::backend: handle AuthHeader::None as a no-op (skip auth
injection).
- server::inference::resolve_provider_route: gate find_provider_api_key
on auth != None. aws-bedrock providers with empty credentials now
resolve cleanly. Updated the unsupported-type error message to
include aws-bedrock in the supported list.
- server::inference tests: add positive
upsert_cluster_route_succeeds_for_aws_bedrock_without_api_key test
covering the new code path end-to-end (provider with empty creds +
BEDROCK_BASE_URL config → upsert succeeds → resolved route has
empty api_key + provider_type aws-bedrock + bridge URL).
- core::inference tests: profile_for_known_types covers aws-bedrock,
case-insensitive lookup, plus three new aws-bedrock-specific tests
(auth: None, no credential keys, bedrock-specific protocols).
- docs/sandboxes/inference-routing.mdx: header forwarding row
mentions aws-bedrock has no passthrough headers; new tabs in
Supported API Patterns (InvokeModel + InvokeModelWithResponseStream)
and Create a Provider (with the bridge-fronted shape note + SigV4
deferral).
- docs/sandboxes/manage-providers.mdx: new row in Supported Provider
Types table; new row in Supported Inference Providers table.
Verification (in dev container):
- cargo check -p openshell-core -p openshell-router -p openshell-server: clean
- cargo test -p openshell-core --lib inference: 14/14 pass (incl. 3 new)
- cargo test -p openshell-server --lib inference::tests::upsert: 6/6 pass
(incl. new aws-bedrock test)
- cargo fmt --check: clean
- cargo clippy --all-targets -D warnings: clean
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(aws-bedrock): bridge-only YAML, doc actual cmd shape, neg test
Addresses four findings from gator-agent's #1704 re-check on 4ab587f1:
- **Item 5** (YAML collects unused AWS creds): mark all four AWS
credentials `required: false` and clear `discovery.credentials`.
Bridge-fronted routing intentionally does not consume AWS
credentials, so `--from-existing` no longer scans for them. The
credentials remain in the schema (not deleted) so the SigV4
follow-up can flip them back without a schema migration. Added a
multi-line description that names the bridge-fronted shape and the
SigV4 deferral so readers don't have to cross-reference the PR
thread.
- **Item 3** (docs show command that the CLI rejects): rewrite the
Create-a-Provider example for AWS Bedrock to use the actual
required shape — placeholder `--credential AWS_ACCESS_KEY_ID=
unused-bridge-fronted-shape` plus the `--config BEDROCK_BASE_URL`.
The placeholder satisfies the gRPC handler's
`provider.credentials.is_empty()` rejection without expanding
server-side validation; the router ignores it on the outbound path
because `auth: AuthHeader::None` skips header injection. Operators
see a clearly-labeled placeholder in `provider get` output.
- **Item 1** (validator probe): document `--no-verify` as required
for `openshell inference set --provider <aws-bedrock>` since the
default validation probe doesn't recognize the
`aws_bedrock_invoke` / `aws_bedrock_invoke_stream` protocols. Doc
now shows the full `provider create` + `inference set --no-verify`
flow with rationale for both decisions inline.
- **Item 6** (docs polish): `inference-routing.mdx` summary row now
lists AWS Bedrock alongside NVIDIA, Anthropic, Vertex AI, and
OpenAI-compatible providers, with the bridge-fronted caveat
inline.
Test additions in `crates/openshell-server/src/inference.rs`:
- Renamed the existing aws-bedrock test from
`..._without_api_key` to `..._with_bridge_url` and updated it to
use a placeholder credential (mirroring the doc-recommended
pattern operators will copy-paste). The `auth: None` path still
produces an empty `api_key` on the resolved route — the test now
documents that the credential is *stored* but not *used*.
- Added `upsert_cluster_route_rejects_aws_bedrock_without_bedrock_base_url`:
the negative half of johntmyers' "successfully used by
upsert_cluster_inference_route or intentionally rejected with a
clear documented error" ask. With
`default_base_url: ""` and no `BEDROCK_BASE_URL` config, route
resolution returns `InvalidArgument` naming the missing base_url
rather than silently forwarding prompts to AWS Bedrock with no
usable auth.
Verification (in dev container):
- cargo test -p openshell-core --lib inference: 18/18 (incl. 3 new)
- cargo test -p openshell-server --lib inference::tests::upsert: 8/8
(incl. 2 new aws-bedrock cases — positive + negative)
- cargo fmt --check: clean
- cargo clippy --all-targets -D warnings: clean
Item 2 (router-side enforcement of operator-configured Bedrock model
path, replacing the current verbatim path forwarding + body-only
model rewrite) is the remaining blocker and is genuinely separable —
it touches the L7 router with streaming-aware test coverage.
Deferring to its own commit so the security-critical change gets the
review attention it deserves.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(router): enforce operator-configured Bedrock model in request path
Closes the security-blocking item from gator-agent's #1704 re-check
on 4ab587f1: "Bedrock carries the model id in /model/{modelId}/invoke,
but the router currently forwards the caller's original path and only
rewrites JSON body model. That lets sandbox code choose a different
upstream model than the operator-configured route model, and may also
mutate native Bedrock request bodies incorrectly."
Two changes in `prepare_backend_request`:
1. **Path rewrite for Bedrock routes.** Before computing the upstream
URL, parse the inbound path's `/model/<id>/invoke[-with-response-stream]`
shape and substitute the operator-configured `route.model` for the
caller-supplied model segment. Sandbox code that hardcodes a
different model still works (we don't reject on mismatch), but the
operator's configured model is what reaches the upstream / bridge.
If the inbound path is somehow not a recognized Bedrock shape on a
Bedrock route (the L7 pattern detector upstream of the router
should never produce this combination), reject with
RouterError::Internal naming the offending path rather than
forwarding verbatim.
2. **Skip body-model injection for Bedrock routes.** The existing body
rewriter unconditionally inserts `route.model` into the JSON body
for non-Vertex routes. AWS Bedrock InvokeModel encodes the model
in the URL path; the body is the raw provider-specific payload
(Anthropic Messages for Claude, Mistral payload for Mistral, etc.)
and must not be mutated. The branch ordering is now:
needs_vertex_anthropic_version → strip body model + inject
anthropic_version; route_is_bedrock → leave body alone; else →
inject route.model (existing default).
New helpers, all in `crates/openshell-router/src/backend.rs`:
- `route_is_bedrock(route)` — true when route.protocols contains
aws_bedrock_invoke or aws_bedrock_invoke_stream.
- `parse_bedrock_invocation_path(path)` — returns
Some((model_id, "/invoke" | "/invoke-with-response-stream")) for
paths matching the recognized Bedrock shapes. Strips query strings.
Rejects empty model ids and multi-segment ids (defense-in-depth
matching the L7 pattern detector's existing guards).
- `rewrite_bedrock_path(route, path)` — returns the path with the
caller's model segment replaced by route.model.
Test coverage in the same file (9 new tests):
- parse_bedrock_invocation_path: positive cases for both invoke
variants, query-string stripping; negative cases for empty model id,
multi-segment id, unknown action, wrong prefix, missing slash.
- route_is_bedrock: matches both protocol variants singly and
combined; rejects openai_chat_completions.
- rewrite_bedrock_path: substitutes operator model on both invoke
variants; returns None for non-Bedrock paths.
- bedrock_route_rewrites_model_in_path_and_preserves_body
(wiremock end-to-end): caller sends /model/some-other-model/invoke
with a body containing model: "caller-supplied-model-name". Mock
asserts the upstream receives /model/<operator-model>/invoke and the
body's model field is the caller's value (NOT route.model) — proves
both the path rewrite and the body preservation.
- bedrock_route_streaming_rewrites_model_in_path: same contract for
invoke-with-response-stream.
- bedrock_route_rejects_non_bedrock_path: defense-in-depth coverage of
the Internal-error path when a Bedrock route receives a path that
doesn't match Bedrock shape.
Verification (in dev container):
- cargo test -p openshell-router --lib: 53/53 (incl. 9 new)
- cargo fmt --check: clean
- cargo clippy -p openshell-core -p openshell-router -p openshell-server
--all-targets -- -D warnings: clean
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(sandbox/l7): declare framing for Bedrock patterns
Closes the buffered-vs-streaming framing warning from gator-agent's
re-check on 4ab587f1: "Bedrock InvokeModel should be buffered while
InvokeModelWithResponseStream is streaming. Please add framing/coverage
so /model/{id}/invoke cannot be corrupted by the streaming proxy's
truncation/error-frame behavior."
The InferenceApiPattern struct gained a `framing: ResponseFraming`
field upstream after the original Bedrock-patterns commit (#22b78cff)
landed; the cherry-pick onto current upstream/main left the two
Bedrock entries without the new field. Fixed here:
- aws_bedrock_invoke (POST /model/{id}/invoke):
framing = ResponseFraming::Buffered
InvokeModel returns one JSON object the caller decodes whole. Sending
it through the streaming proxy would risk a mid-body size-cap
truncation or idle-timeout failure appending an SSE error event onto
bytes the caller decodes as one JSON body — the same corruption mode
that drove the existing embeddings + model-discovery to Buffered.
- aws_bedrock_invoke_stream (POST /model/{id}/invoke-with-response-stream):
framing = ResponseFraming::Streaming
InvokeModelWithResponseStream returns an AWS event-stream of binary
chunks; the caller wants chunks incrementally, so the streaming proxy
path is correct.
Two new tests in `crates/openshell-sandbox/src/l7/inference.rs` pin
down the contract:
- aws_bedrock_invoke_is_buffered — detect_inference_pattern returns a
Buffered pattern for /model/<id>/invoke, with explanatory message
naming the corruption mode being prevented.
- aws_bedrock_invoke_stream_is_streaming — same shape, asserting
Streaming for /model/<id>/invoke-with-response-stream.
Verification (in dev container):
- cargo check -p openshell-sandbox: clean (was failing on missing
`framing` field before this commit)
- cargo test -p openshell-sandbox --lib l7::inference::tests::aws_bedrock:
7/7 (incl. 2 new framing tests)
- cargo fmt --check: clean
- cargo clippy --all-targets -- -D warnings: clean
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(inference): drop aws_bedrock_invoke_stream until protocol-aware errors land
Per PR #1704 review (johntmyers): defer Bedrock streaming to a follow-up
that also wires protocol-aware error framing for the AWS event-stream
shape. Until then, surfacing `/model/{id}/invoke-with-response-stream`
risks shipping responses the sandbox cannot interpret on failure.
- AWS_BEDROCK_PROTOCOLS no longer advertises `aws_bedrock_invoke_stream`.
- The L7 inference pattern table drops the streaming entry; only
`aws_bedrock_invoke` (buffered) is recognized.
- Test `aws_bedrock_invoke_stream_pattern_is_deferred` asserts no
pattern claims that protocol so the gap is visible.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(inference): validate AWS Bedrock model_id at route + rewrite, preserve query
Per PR #1704 review (johntmyers, gator agent): the operator-configured
Bedrock model_id flows verbatim into the upstream URL path; the previous
plumbing left no enforcement that the value was a single benign path
segment, opening a path-injection vector.
Defense in depth, both layers:
* `openshell-server::inference`: new `validate_aws_bedrock_model_id`
rejects empty, leading/trailing whitespace, `/`, `\\`, `?`, `#`, `%`,
`..`, control/whitespace characters. Wired into `resolve_provider_route`
ahead of base_url resolution so the route store cannot persist a
malformed model_id. Mirrors `validate_vertex_model_id` exactly.
* `openshell-router::backend`: `rewrite_bedrock_path` now refuses to
construct the upstream URL unless `route.model` passes
`is_valid_bedrock_model_id`, so even a stale or hand-edited route
cannot reach the wire. The parser also drops the
`/invoke-with-response-stream` arm to match the protocol catalog.
* `parse_bedrock_invocation_path` returns the `?`-prefixed query tail
as a third element; `rewrite_bedrock_path` re-attaches it so any
caller-supplied query string is preserved through the model rewrite.
Tests: 5 unit tests for the validator, 1 integration test that
exercises every unsafe-model_id reject path through
`upsert_cluster_inference_route`, plus a router-side rewrite-rejects
test covering 11 unsafe `route.model` values. All 72 server inference
tests + router tests pass.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* fix(aws-bedrock): make built-in profile non-egress-granting; tighten docs
Per PR #1704 review (johntmyers): a single profile that is "usable as
is" — not a profile that auto-grants direct AWS Bedrock egress to any
sandbox that selects it. The bridge IS the egress point and is
operator-managed; the profile must not implicitly punch a hole through
the cluster's network policy on its behalf.
* `providers/aws-bedrock.yaml`: clear `endpoints` and `binaries` to
empty arrays. Rewrite the description to spell out that the profile
is intentionally non-egress-granting, that operators are responsible
for declaring their bridge's egress endpoint and binary attribution,
and that the SigV4 follow-up will repopulate these fields once
router-side signing exists.
* `docs/sandboxes/inference-routing.mdx`:
- Drop the `InvokeModelWithResponseStream` row from the supported
patterns table (matches the protocol-catalog + L7-pattern drop).
- Update the `--no-verify` paragraph to reference only
`aws_bedrock_invoke`.
- Generalise the placeholder-credential rationale: any
standalone-router profile registering `AuthHeader::None` will hit
the same non-empty-credentials structural requirement.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* style(router): satisfy cargo fmt for Bedrock query-tail extraction
CI's rustfmt rejected the one-line `let (path_only, query_tail) = path.find('?').map_or(...)`
shape from commit c51160e4 and required the chained-method layout
instead. Functional behaviour is unchanged; tests still pass.
Signed-off-by: st-gr <38470677+st-gr@users.noreply.github.com>
* style(router): drop unnecessary `crate::` prefix on RouterError matcher
`RouterError` is already imported at the top of the file, so the
two test-side `matches!(result, Err(crate::RouterError::UpstreamProtocol(_)))`
uses trip clippy's `unused_qualifications` under `-D warnings`. Drop
the prefix on both sites; functional behaviour is unchanged.
These warnings predate this PR (originated in
|
||
|
|
48545cfb59 |
feat(sandbox): add GCE metadata emulator for Google Cloud (#1763)
* feat(core): add shared GCP constants module Single source of truth for GCP naming: env var aliases, provider config keys, token search order, and Vertex-specific env vars. Consumed by openshell-server, openshell-providers, and openshell-sandbox. - Add google_cloud.rs with metadata emulator host and loopback address - Define PROJECT_ID, REGION, and SERVICE_ACCOUNT_EMAIL env var aliases - Add provider config key constants for gcp provider implementations - Define TOKEN_ENV_KEYS search order (SA token takes priority over ADC) - Add Vertex-specific env vars for Goose and Claude Code SDK integration - Add STATIC_CONFIG_KEYS as union of all alias arrays for env resolution - Export module via openshell-core lib.rs Signed-off-by: Robert Sturla <rsturla@redhat.com> * feat(providers): add google-cloud and vertex provider plugins Add GoogleCloudProvider and VertexProvider implementing inject_env to project GCP config (project ID, region, SA email, metadata host) into sandbox environment variables. Replace the inline Vertex AI env injection in the server with the registry-based inject_env dispatch. Also adds the google-cloud.yaml provider profile with SA JWT and ADC OAuth2 credential refresh flows. Signed-off-by: Robert Sturla <rsturla@redhat.com> * feat(sandbox): add GCE metadata emulator for GCP Add a loopback HTTP server on 127.0.0.1:8174 inside the sandbox network namespace that emulates the GCE instance metadata API. GCP client SDKs discover it via GCE_METADATA_HOST and obtain credential placeholders that the proxy resolves to real tokens at egress. Add metadata_server module with MetadataHandler trait and netns-aware TCP binding via std::thread (not spawn_blocking) to avoid tokio pool namespace contamination Add google_cloud_metadata module implementing the GCE metadata API subset (token, project-id, email, scopes, service-accounts) Add child_env_resolved() and gcp_token_response() to ProviderCredentialState for GCP-aware credential projection Wire metadata server into sandbox lifecycle before SSH handler Collapse multi-line HTTP response format string into single line Signed-off-by: Robert Sturla <rsturla@redhat.com> * docs(sandbox): add GCP credentials documentation Document the google-cloud provider setup for ADC and service account flows, injected environment variables, metadata emulator behavior, and network policy configuration for GCP APIs. Signed-off-by: Robert Sturla <rsturla@redhat.com> * feat(cli): support --from-gcloud-adc for google-cloud providers Widen --from-gcloud-adc to accept google-cloud providers. The ADC credential key is derived from the provider profile rather than hardcoded per type, so future GCP provider types get ADC support by declaring the right refresh metadata in their profile YAML. Add ProviderTypeProfile::adc_credential() to find the ADC-compatible credential from a profile's refresh metadata. Remove unused VERTEX_AI_ADC_TOKEN_KEY and GCP_ADC_TOKEN_KEY constants. Signed-off-by: Robert Sturla <rsturla@redhat.com> --------- Signed-off-by: Robert Sturla <rsturla@redhat.com> |
||
|
|
36bb9e3ec2 |
feat(providers): add DeepInfra as a built-in inference provider (#1902)
* feat(providers): add DeepInfra as a built-in inference provider (v2 only) - Adds `deepinfra` as a built-in Providers v2 profile (`providers/deepinfra.yaml`) with inference category, Bearer auth, and `DEEPINFRA_API_KEY` discovery - Adds `DEEPINFRA_PROFILE` to inference routing so `inference.local` works with the `deepinfra` provider type - Fixes `build_backend_url` to strip `/v1` from request paths when the base URL contains `/v1/` as an internal segment (e.g. `api.deepinfra.com/v1/openai`), preventing double-versioned paths like `.../v1/openai/v1/chat/completions` - Updates `docs/sandboxes/providers-v2.mdx` and `docs/sandboxes/manage-providers.mdx` with DeepInfra entries; removes the old v1 workaround row that used `openai` type with `OPENAI_API_KEY` Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * fix(providers): address gator review findings for DeepInfra provider - Narrow build_backend_url /v1 dedupe to URLs whose path component is exactly /v1 or starts with /v1/ — prevents regression on proxy endpoints where /v1 is buried deeper (e.g. /api/v1/openai); add regression test for the nested proxy path case - Add deepinfra provider plugin with DEEPINFRA_API_KEY discovery, registered in ProviderRegistry so known_types() and TUI include it - Add deepinfra to unsupported-inference-provider error message in openshell-server for accurate user-facing debugging guidance - Add deepinfra to openai_compatible_profiles_include_embeddings test to lock in the OpenAI-compatible protocol contract Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * fix(router): handle /v1 as final path segment in build_backend_url dedup Extends the /v1 deduplication logic to also strip /v1 from request paths when the base URL's path ends with /v1 (e.g. https://api.groq.com/openai/v1). The previous fix only matched paths starting with /v1/, which regressed providers like Groq whose base path has /v1 as the last segment rather than the first. The nested-proxy exclusion (e.g. /api/v1/openai) is preserved since /v1 appears in the middle — neither first nor last segment. Adds a regression test for the Groq-style base URL. Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * fix(providers): add deepinfra telemetry bucket and update profile list test - Add DeepInfra variant to ProviderProfile telemetry enum and from_raw() mapping so deepinfra providers are tracked in their own bucket rather than falling through to Custom - Map deepinfra in telemetry_provider_profile() in openshell-server - Add deepinfra to list_provider_profiles_returns_built_in_profile_categories test (sorted between cursor and github) - Update architecture/gateway.md inference provider list to include deepinfra Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> * style(router): apply cargo fmt to backend.rs Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> --------- Signed-off-by: Milos Milutinovic <codemastermilos@gmail.com> |
||
|
|
62c421b21e |
feat(providers): add profile-backed policy visibility (#1640)
* chore: wip providers v2 tui and codex profile * chore: wip effective policy get and codex profile * chore: wip provider profiles and tui detail views * feat(tui): annotate policy proposal review status |
||
|
|
f061b1d923 |
feat(providers): add Google Vertex AI inference provider (#1568)
* feat(providers): add Google Vertex AI provider Adds Vertex AI provider profiles, routing, credential refresh plumbing, CLI support, docs, and regression coverage. Keeps the related NETLINK_ROUTE seccomp allowance needed by Vertex client tooling that calls getifaddrs. * docs: add Vertex AI sandbox usage for Claude Code and OpenCode Cover the full end-to-end setup for running Claude Code and OpenCode inside an OpenShell sandbox via inference.local with a Vertex AI backend: - google-vertex-ai.mdx: add 'Use from a Sandbox' section with tabbed examples for Claude Code (--bare flag, no /v1 suffix) and OpenCode (/v1 suffix required). Add providers_v2_enabled prerequisite and --no-verify note for global region. Document policy proposals table covering metadata.google.internal (always blocked), downloads.claude.ai, and storage.googleapis.com. - inference-routing.mdx: expand 'Use the Local Endpoint' section with tabbed examples for Claude Code, OpenCode, Python OpenAI SDK, and Python Anthropic SDK. Add notes explaining the /v1 path suffix difference between clients. - supported-agents.mdx: update Claude Code and OpenCode rows to mention inference.local support and correct base URL requirements. * fix: address vertex review findings * test(sandbox): retry on spurious Ok in fork-exec ambiguity test On arm64 under heavy CI load, the /proc fd scan in find_socket_inode_owners can transiently miss the parent process's socket fd entry, returning only the child as an owner. This causes resolve_process_identity to return Ok (single owner, no ambiguity check fires) instead of the expected ambiguous-ownership Err. Extend the retry loop to also handle unexpected Ok results, mirroring the existing retry for transient Err results. 10 retries at 50ms gives a 500ms settling window, which is sufficient for procfs to stabilize on loaded arm64 runners. * fix: address vertex review regressions * docs(router): clarify stream_response semantics for Vertex rawPredict routing Document the three call sites of prepare_backend_request and their stream_response values in a caller table: - send_backend_request: false → :rawPredict (unary endpoint) - send_backend_request_streaming: true → :streamRawPredict - verify_backend_endpoint: explicitly false to probe the unary endpoint Cross-reference the table from build_provider_url and is_vertex_anthropic_rawpredict_route so the stream_response=true guard in the suffix upgrade branch is understood in full context. Also note that is_vertex_anthropic_rawpredict_route is a structural predicate (model_in_path + anthropic_messages + :rawPredict suffix), not a named-provider check, so any future provider with the same route shape inherits the transforms automatically. |
||
|
|
e98ea3ee93 | feat(policy): add agentic approval loop (#1528) | ||
|
|
0cef26521a |
feat(providers): derive discovery from profiles (#1503)
* feat(providers): derive discovery from profiles Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): keep v2 discovery profile-only Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs(providers): update providers v2 behavior Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(providers): make github profile read-only Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
d255cdd9c9 |
feat(providers): add credential refresh foundation (#1349)
* feat(providers): add credential refresh foundation * feat(providers): mint oauth refresh token credentials * fix(providers): allow delegated refresh bootstrap * fix(providers): guard refresh credential modes * fix(providers): refine refresh lifecycle UX * fix(cli): restore provider test gateway response import * fix(server): clean auth endpoint test qualifications * fix(providers): tighten refresh authorization and collisions * chore(providers): trim bundled v2 profiles * fix(providers): resolve refresh rebase fallout * fix(cli): accept rfc3339 credential expiry * test(providers): update inferred claude provider type * test(providers): avoid removed outlook default profile * test(providers): isolate attach limit fixtures |
||
|
|
043bde279a |
feat(providers): add profile-backed policy composition (#1037)
Foundation for providers v2. Add provider profiles and provider profile composition with user policies. |