Files
OpenShell/python
Piotr Mlocek 34fd3cbf64 feat(inference): inference interception and routing (!38)
Closes #67

## Summary

Implements transparent inference interception and routing for sandboxed AI agents. The sandbox proxy intercepts outbound AI SDK calls (OpenAI, Anthropic) and reroutes them through the gateway to policy-controlled backends — enabling organizations to redirect inference traffic to local or self-hosted models without modifying agent code.

**Decision model** — a tri-state OPA evaluation for every CONNECT request:
1. Binary + endpoint explicitly allowed in `network_policies` → **allow** (pass through)
2. Not explicitly allowed + `inference.allowed_routes` configured → **inspect for inference** (TLS intercept, detect API patterns, route through gateway)
3. Otherwise → **deny**

No endpoint declarations or binary lists needed for inference routing. Just configure `inference.allowed_routes`.

## Key Changes

### Sandbox (interception)
- **OPA policy**: New `network_action` Rego rule with three outcomes (`allow`, `inspect_for_inference`, `deny`). New `NetworkAction` enum replaces `PolicyDecision.allowed` bool for the proxy's main decision path.
- **Proxy**: New `InspectForInference` path — TLS-terminates client, parses HTTP, detects inference API patterns (`POST /v1/chat/completions`, `/v1/completions`, `/v1/messages`), strips auth headers, forwards via gRPC.
- **New module**: `l7/inference.rs` — `InferenceApiPattern`, `detect_inference_pattern()`, HTTP request/response parsing.
- **gRPC client**: New `proxy_inference()` for sandbox→gateway forwarding.
- **Sandbox init**: Creates OPA engine when inference is configured, even without `network_policies`.

### Gateway (dispatch)
- **InferenceService**: `ProxyInference` RPC loads sandbox policy, resolves allowed routes, dispatches to router. Full CRUD for inference routes.
- **Proto**: `InferenceRoute`, `InferenceRouteSpec`, `ProxyInferenceRequest/Response`, Inference gRPC service.

### Router (backend proxying)
- **New crate**: `navigator-router` with `Router`, `proxy_with_candidates()`, protocol-based route selection, backend HTTP proxying with auth header rewriting.
- **Mock support** for testing (`mock://` scheme).

### CLI
- `nav inference create/update/delete/list` commands for route management.

### Python SDK
- Updated protobuf bindings. Removed old `inference.py` client (replaced by transparent interception).

### Documentation
- New `architecture/inference-routing.md` — end-to-end system documentation.
- Updated `architecture/sandbox.md` — proxy, OPA, and source index sections.
- Updated `architecture/README.md` — new subsystem overview and diagram.

## Addendum: Chunked Transfer Compatibility

This branch now also fixes intercepted SDK requests that send chunked request bodies:

- `inspect_for_inference` now accepts `Transfer-Encoding: chunked` and decodes chunked request bodies before forwarding to the gateway
- Removed the prior `411 Length Required` response for chunked intercepted requests
- Added request/response header sanitization for framing and hop-by-hop headers (`content-length`, `transfer-encoding`, `connection`, etc.) to keep forwarded requests and returned responses valid
- Added unit tests for chunked parsing and header sanitization

Note: this improves compatibility for streaming-style SDK request patterns; true token-by-token passthrough response streaming is still a separate follow-up.

## Minimal Policy for Inference Routing

```yaml
inference:
  allowed_routes:
    - local
```

Any outgoing connection from a binary not explicitly allowed in `network_policies` will be intercepted and checked for inference API patterns.

## Test Plan

- [x] `cargo test --workspace` — all tests pass
- [x] `mise run pre-commit` — all checks pass
- [x] E2E: OpenAI chat completions routed through gateway
- [x] E2E: Anthropic messages routed through gateway
- [x] E2E: Python OpenAI SDK from sandbox (`examples/inference/inference.py`)
- [x] E2E test: `e2e/python/test_inference_routing.py`
2026-02-23 11:48:48 -08:00
..