* refactor(inference): remove managed inference routes
Closes#3172
Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(policy): preserve alternate upstream isolation
Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints.
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
---------
Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com>
* fix(proxy): stream inference responses instead of buffering entire body
The inference.local proxy path called response.bytes().await which
buffered the entire upstream response before sending anything to the
client. For streaming SSE responses this inflated TTFB from sub-second
to the full generation time, causing clients with TTFB timeouts to abort.
Add a streaming proxy variant that returns response headers immediately
and forwards body chunks incrementally using HTTP chunked transfer
encoding. Non-streaming responses and mock routes continue to work
through the existing buffered path.
Closes#260
* docs: update architecture docs and example for inference streaming
* test(example): add direct NVIDIA endpoint tests via L7 TLS intercept
Expand inference example to 4 test cases: inference.local and direct
endpoint, each streaming and non-streaming. The direct path exercises
the L7 REST relay (relay_chunked) to verify it already streams
correctly. NVIDIA_API_KEY is picked up from the sandbox env when
started with --provider nvidia.