mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
This is part of #1266 , opening now for discussion. Stacked on https://github.com/agent-substrate/substrate/pull/1682 which was slightly orthogonal. This is loosely based on the mini proposal by @howardjohn as discussed in the community meeting, and feedback from @bowei @EItanya @LiorLieberman @aojea. https://docs.google.com/document/d/1TycfQ3iiEpbI3rveMIj0S2PpPuLecb8I5R--yTpt9Ig/edit?resourcekey=0-kJbtEZ-KGzuL5eCjHDvBhg&tab=t.0#heading=h.ga9bfaf55ptk Roughly: 1. Actor sandboxes each get their own netns. - gVisor grabs all interfaces in the netns, and tap currently requires running something like slipr, so for now we do two netns + a veth when gVisor. 3. In the netns we directly intercept TCP => atunnel for general traffic. 4. In the netns we serve a trivial TCP+UDP DNS relay to the pod resolution. - In the future we can insert policy here. 6. We consistently inject a modified resolv.conf instead of bind-mounting it (gVisor) across both runtimes. 7. Readiness probe dialing happens in the actor netns. Every actor gets the same fixed guest IP as before, which is only visible to the actor. All inbound/outbound traffic comes from atunnel / the DNS relay. The actor no longer has any direct use of the pod interface, so we can begin to consider ateom using the network itself. When we add the rest of multi-actor changes, this greatly simplifies thing. Full multi-actor requires further changes, but this diff is already large (suggest reading commit by commit) and can stand-alone. I'll file more stacked changes when we've got consensus on this one.