mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 00:23:53 +08:00
Closes #67 ## Summary Implements transparent inference interception and routing for sandboxed AI agents. The sandbox proxy intercepts outbound AI SDK calls (OpenAI, Anthropic) and reroutes them through the gateway to policy-controlled backends — enabling organizations to redirect inference traffic to local or self-hosted models without modifying agent code. **Decision model** — a tri-state OPA evaluation for every CONNECT request: 1. Binary + endpoint explicitly allowed in `network_policies` → **allow** (pass through) 2. Not explicitly allowed + `inference.allowed_routes` configured → **inspect for inference** (TLS intercept, detect API patterns, route through gateway) 3. Otherwise → **deny** No endpoint declarations or binary lists needed for inference routing. Just configure `inference.allowed_routes`. ## Key Changes ### Sandbox (interception) - **OPA policy**: New `network_action` Rego rule with three outcomes (`allow`, `inspect_for_inference`, `deny`). New `NetworkAction` enum replaces `PolicyDecision.allowed` bool for the proxy's main decision path. - **Proxy**: New `InspectForInference` path — TLS-terminates client, parses HTTP, detects inference API patterns (`POST /v1/chat/completions`, `/v1/completions`, `/v1/messages`), strips auth headers, forwards via gRPC. - **New module**: `l7/inference.rs` — `InferenceApiPattern`, `detect_inference_pattern()`, HTTP request/response parsing. - **gRPC client**: New `proxy_inference()` for sandbox→gateway forwarding. - **Sandbox init**: Creates OPA engine when inference is configured, even without `network_policies`. ### Gateway (dispatch) - **InferenceService**: `ProxyInference` RPC loads sandbox policy, resolves allowed routes, dispatches to router. Full CRUD for inference routes. - **Proto**: `InferenceRoute`, `InferenceRouteSpec`, `ProxyInferenceRequest/Response`, Inference gRPC service. ### Router (backend proxying) - **New crate**: `navigator-router` with `Router`, `proxy_with_candidates()`, protocol-based route selection, backend HTTP proxying with auth header rewriting. - **Mock support** for testing (`mock://` scheme). ### CLI - `nav inference create/update/delete/list` commands for route management. ### Python SDK - Updated protobuf bindings. Removed old `inference.py` client (replaced by transparent interception). ### Documentation - New `architecture/inference-routing.md` — end-to-end system documentation. - Updated `architecture/sandbox.md` — proxy, OPA, and source index sections. - Updated `architecture/README.md` — new subsystem overview and diagram. ## Addendum: Chunked Transfer Compatibility This branch now also fixes intercepted SDK requests that send chunked request bodies: - `inspect_for_inference` now accepts `Transfer-Encoding: chunked` and decodes chunked request bodies before forwarding to the gateway - Removed the prior `411 Length Required` response for chunked intercepted requests - Added request/response header sanitization for framing and hop-by-hop headers (`content-length`, `transfer-encoding`, `connection`, etc.) to keep forwarded requests and returned responses valid - Added unit tests for chunked parsing and header sanitization Note: this improves compatibility for streaming-style SDK request patterns; true token-by-token passthrough response streaming is still a separate follow-up. ## Minimal Policy for Inference Routing ```yaml inference: allowed_routes: - local ``` Any outgoing connection from a binary not explicitly allowed in `network_policies` will be intercepted and checked for inference API patterns. ## Test Plan - [x] `cargo test --workspace` — all tests pass - [x] `mise run pre-commit` — all checks pass - [x] E2E: OpenAI chat completions routed through gateway - [x] E2E: Anthropic messages routed through gateway - [x] E2E: Python OpenAI SDK from sandbox (`examples/inference/inference.py`) - [x] E2E test: `e2e/python/test_inference_routing.py`