Closes#273
Verify inference endpoints synchronously on the server during set/update, expose a --no-verify escape hatch in the CLI and Python helper, and return actionable failures when validation does not pass.
Closes#67
## Summary
Implements transparent inference interception and routing for sandboxed AI agents. The sandbox proxy intercepts outbound AI SDK calls (OpenAI, Anthropic) and reroutes them through the gateway to policy-controlled backends — enabling organizations to redirect inference traffic to local or self-hosted models without modifying agent code.
**Decision model** — a tri-state OPA evaluation for every CONNECT request:
1. Binary + endpoint explicitly allowed in `network_policies` → **allow** (pass through)
2. Not explicitly allowed + `inference.allowed_routes` configured → **inspect for inference** (TLS intercept, detect API patterns, route through gateway)
3. Otherwise → **deny**
No endpoint declarations or binary lists needed for inference routing. Just configure `inference.allowed_routes`.
## Key Changes
### Sandbox (interception)
- **OPA policy**: New `network_action` Rego rule with three outcomes (`allow`, `inspect_for_inference`, `deny`). New `NetworkAction` enum replaces `PolicyDecision.allowed` bool for the proxy's main decision path.
- **Proxy**: New `InspectForInference` path — TLS-terminates client, parses HTTP, detects inference API patterns (`POST /v1/chat/completions`, `/v1/completions`, `/v1/messages`), strips auth headers, forwards via gRPC.
- **New module**: `l7/inference.rs` — `InferenceApiPattern`, `detect_inference_pattern()`, HTTP request/response parsing.
- **gRPC client**: New `proxy_inference()` for sandbox→gateway forwarding.
- **Sandbox init**: Creates OPA engine when inference is configured, even without `network_policies`.
### Gateway (dispatch)
- **InferenceService**: `ProxyInference` RPC loads sandbox policy, resolves allowed routes, dispatches to router. Full CRUD for inference routes.
- **Proto**: `InferenceRoute`, `InferenceRouteSpec`, `ProxyInferenceRequest/Response`, Inference gRPC service.
### Router (backend proxying)
- **New crate**: `navigator-router` with `Router`, `proxy_with_candidates()`, protocol-based route selection, backend HTTP proxying with auth header rewriting.
- **Mock support** for testing (`mock://` scheme).
### CLI
- `nav inference create/update/delete/list` commands for route management.
### Python SDK
- Updated protobuf bindings. Removed old `inference.py` client (replaced by transparent interception).
### Documentation
- New `architecture/inference-routing.md` — end-to-end system documentation.
- Updated `architecture/sandbox.md` — proxy, OPA, and source index sections.
- Updated `architecture/README.md` — new subsystem overview and diagram.
## Addendum: Chunked Transfer Compatibility
This branch now also fixes intercepted SDK requests that send chunked request bodies:
- `inspect_for_inference` now accepts `Transfer-Encoding: chunked` and decodes chunked request bodies before forwarding to the gateway
- Removed the prior `411 Length Required` response for chunked intercepted requests
- Added request/response header sanitization for framing and hop-by-hop headers (`content-length`, `transfer-encoding`, `connection`, etc.) to keep forwarded requests and returned responses valid
- Added unit tests for chunked parsing and header sanitization
Note: this improves compatibility for streaming-style SDK request patterns; true token-by-token passthrough response streaming is still a separate follow-up.
## Minimal Policy for Inference Routing
```yaml
inference:
allowed_routes:
- local
```
Any outgoing connection from a binary not explicitly allowed in `network_policies` will be intercepted and checked for inference API patterns.
## Test Plan
- [x] `cargo test --workspace` — all tests pass
- [x] `mise run pre-commit` — all checks pass
- [x] E2E: OpenAI chat completions routed through gateway
- [x] E2E: Anthropic messages routed through gateway
- [x] E2E: Python OpenAI SDK from sandbox (`examples/inference/inference.py`)
- [x] E2E test: `e2e/python/test_inference_routing.py`
## Summary
- Add `Provider` entity for managing 3p deps from a sandbox
- Add provider CRUD API/server persistence and new CLI workflows (`nav provider create/get/list/update/delete`), including `--from-existing` laptop discovery.
- Integrate providers into sandbox create flow: infer from command (`-- claude`), support repeatable `--provider <type>`, prompt before auto-create, and allow manual in-sandbox setup.
- Add a dedicated `navigator-providers` crate with per-provider modules and mockable discovery test helpers.
## Key UX Changes
- `nav sandbox create --provider gitlab -- claude`
- Missing provider prompt now asks before creating from local state.
- `nav provider list --names` for scripting/cleanup.
## Test Plan
- `mise run cluster:deploy`
- `mise run test:e2e:sandbox`
- `mise run pre-commit`
Closes#19Closes#22Closes#11
Closes#13
## Summary
- add Python sandbox execution APIs for command and callable workflows
- consolidate sandbox policy fixtures and expand e2e test coverage for policy and Python exec paths
- update CI/build config and images for sandbox e2e execution dependencies
## Test Plan
- mise run pre-commit
Closes#24
## Problem
SSH sessions into the sandbox were **not entering the network namespace** or receiving proxy environment variables. This meant every command run via SSH (the only user-facing path) had unrestricted internet access, completely bypassing OPA network policy enforcement.
The root cause: the SSH server was started **before** the network namespace and proxy were created in `lib.rs`, so it never received the netns fd or proxy URL.
Additionally, **gRPC inference from within the sandbox did not work at all** — even after fixing the netns, multiple issues prevented the Python SDK from reaching the navigator server through the CONNECT proxy.
## Changes
### Sandbox binary
**Core fix — reorder startup + thread netns through SSH:**
- `lib.rs`: Move netns + proxy creation before SSH server start. Compute `ssh_netns_fd` and `ssh_proxy_url`, pass them to `run_ssh_server()`.
- `ssh.rs`: Thread `netns_fd` and `proxy_url` through the full SSH call chain into `spawn_pty_shell()`. Set proxy env vars on the shell command. Call `setns(fd, CLONE_NEWNET)` in `install_pre_exec()`.
**Proxy — control plane allowlist + IPv6 socket lookup:**
- `proxy.rs`: Connections to the navigator endpoint (derived from `NAVIGATOR_ENDPOINT`) are always allowed without OPA evaluation, logged with `engine=control_plane`. This is infrastructure the sandbox needs to function, not a user-configurable policy.
- `procfs.rs`: Extended `parse_proc_net_tcp` to also check `/proc/<pid>/net/tcp6`. gRPC C-core uses `AF_INET6` sockets even for IPv4 connections, so its TCP entries were invisible to the proxy's identity resolver. Also fixed port parsing to use `rsplit_once(':')` for correct IPv6 address handling.
**Proxy env vars — lowercase variants for gRPC C-core:**
- `process.rs` + `ssh.rs`: Added lowercase `http_proxy`, `https_proxy`, `grpc_proxy` alongside uppercase. gRPC C-core (libgrpc) checks lowercase first and was ignoring the uppercase-only vars.
### Server
- `sandbox/mod.rs`: Added `CAP_SYS_PTRACE` to sandbox pod security context. Required for the proxy (root) to read `/proc/<pid>/fd/` of sandbox-user processes for binary identity resolution.
### Python SDK
- `inference.py`: Strip `http://`/`https://` scheme from endpoint before passing to `grpc.insecure_channel()`, which expects `host:port`. Default `endpoint` and `sandbox_id` from `NAVIGATOR_ENDPOINT` / `NAVIGATOR_SANDBOX_ID` env vars so `Inference()` works inside sandboxes with no arguments.
### Build / infra
- `ci.toml`: Added `--cap-add=SYS_PTRACE` to `mise run sandbox` to mirror the k8s pod capabilities.
### Documentation
- `architecture/sandbox.md`: Documented `CAP_SYS_PTRACE` requirement and the full set of proxy env vars (uppercase + lowercase).
## Testing
All 29 sandbox unit tests pass. `mise run pre-commit` passes (fmt, clippy, all workspace tests, python lint).
E2E verified on a live cluster:
| Test | Result |
|------|--------|
| Proxy env vars (6 vars, upper+lowercase) | PASS |
| Blocked endpoints (google, anthropic via curl) | PASS — all denied |
| **gRPC inference from SSH session** (`Inference()` with no args, env var defaults) | **PASS** |
| Proxy log: navigator requests show `engine=control_plane` | PASS |
| Proxy log: blocked requests show `engine=opa` with correct deny reasons | PASS |