35 Commits
Author SHA1 Message Date
Lior Lieberman 944096fa21 Passthrough dial address resolved by the gateway (#2045)
The passthrough chain from #2019 dialed the ORIGINAL_DST filter state,
which the CONNECT leg filled with the address the actor connected to.
The SNI picked the chain, the actor's own resolution picked the
destination, so a passthrough rule for one name let an actor reach any
IP by claiming that name in the ClientHello.

We fix it by:

- The passthrough chain is now `sni_dynamic_forward_proxy` then
`tcp_proxy` to a new raw forward-proxy cluster,
`egress_forward_proxy_passthrough`, on the shared egress_dns_cache. The
gateway resolves the SNI itself and sends the bytes to it.

- The CONNECT leg answers the dialed port under
dev.ate.egress:dialed_port; the outer chain copies it into
envoy.upstream.dynamic_port, shared with the inner listener, which both
the SNI filter and the cluster read before their configured port.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-10-01 18:50:20 +00:00
Lior Lieberman 7526b9cc75 remove ssl_cert_dir and add clearer guidance for actor trust store (#2044)
We should not override the whole actor trust store with SSL_CERT_DIR,
removing that to defer using SSL_CERT_FILE, which is additive.

Additionally, this does not cover all languages and runtimes, adding
guidance for other env vars required and other runtimes that need union
of the trust stores in a single file.

In a followup, we are going to remove all "sdsmint" and "plain gateway"
language and just have one gateway thats capable of minting or not based
on the sni.

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-10-01 04:59:44 +00:00
Lior Lieberman 13e8efff25 envoy dyn module followup (#1995)
do not merge, not a draft cause i do want ci running

implements: 
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching


tls_passhtrough is a followup. 



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-30 16:57:00 +00:00
Lior Lieberman 1240749483 ateapi: rename --egress-gateway-address to --default-egress-gateway-address (#1858)
Prefix the cluster-wide egress PEP address flag with `default-`
~'experimental-to-be-removed-' to indicate it is a temporary
cluster-level knob that will be removed~ once per-actor or per-atespace
egress gateway configuration is supported (#1591).

~I kept --egress-gateway-address as a deprecated alias for backward
compatibility but I really dont think we should. ~
I removed it, we dont need backward comp residuals pre GA

EDIT: I chatted with tim and taahir a the feedback was that its not
really experimental and can not really be removed. This flag would have
to be set for every GA user, and its not a nice experimental add on. We
dont have time to get into per-actor or per-atespace API discussions
pre-GA therfore we gotta have to live with it and discourage later if we
have a better story for that.

Related #1591


> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

/cc @bowei @EItanya
2026-09-30 15:46:52 +00:00
Lior Lieberman 2e57c900a3 move agw tests to post-merge and release-* branches (#1949)
agw e2e are non blocking. moving to only run on release branches and
post-merge for now to save runners.

Ref:
https://cloud-native.slack.com/archives/C0B6M3E2J3D/p1790620259554579?thread_ts=1790453226.572009&cid=C0B6M3E2J3D

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 20:27:58 +00:00
Lior Lieberman 6621f2b4b4 egress policy enhancements for GA (#1751)
**BREAKING** API changes prior in preparation for GA. 

~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow


- `EgressRule` is now a union of protocols: `http`, `https`,
  `tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
  Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
  listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
  request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
  `tls_passthrough`. Same wildcard grammar as before, documented once on
  `HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
  once on `HTTPRule.ports`.

- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
  placeholder value. Requests without the header pass unchanged.

- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.

- `google/protobuf/empty.proto` import dropped; nothing uses it now.

Questions

- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
  block is intercepted. is that clear enough?

follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
  the rule conflict check) and regenerate `zz_generated.validation.go`



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Fixes: #823 (there are probably more issues that i need to find)
2026-09-28 17:30:07 +00:00
Lior Lieberman 92ea03609f Comment out CNCF Project Maintainers synchronization note
Comment out the synchronization note with the CNCF Project Maintainers list.
2026-09-15 12:08:01 -07:00
Lior Lieberman c1d5b29f00 flip extproc order 2026-09-14 15:51:44 -07:00
Lior Lieberman c147737f50 docs: describe how the egress gateway enforces an EgressPolicy, leg by leg
The router README gets the gateway-side account, a leg at a time,
including the metadata that carries the CONNECT leg's decision to the
chain that dials. The egress demo README stops calling destination
authorization a follow-up.
2026-09-14 15:51:44 -07:00
Lior Lieberman 29e13a402f e2e: give egress actors policies and cover allow, deny and no-policy
The egress gateway now refuses every tunnel for an actor without an
EgressPolicy, so every e2e actor that egresses needs one. A new
e2e.EnsureEgressPolicy creates or replaces an actor's policy, and the
networking and egressmitm suites give their actors an allow-all rule at
creation, which keeps every existing egress test meaning what it meant.

The cases cover the policy itself. An actor whose policy names one
hostname is refused with 403 and "egress denied" when it asks for
another over cleartext HTTP, and for the same origin by address, since
an IP-literal Host names no host. An actor with no policy gets no tunnel
at all, which the demo app reports as a 502 from a failed connect. A
policy naming the origin's address admits a plain HTTP request to it
both by address and by name, because the request is decided on the
address the actor dialed, and it gets no tunnel for a name that resolves
anywhere else. Over HTTPS an address policy admits example.com on both
gateways: the plain gateway by address at the CONNECT, sdsmint on the
decrypted request, which it then sends to the dialed address with the
certificate checked against the name.

The one case where the gateways differ is a hostname policy over HTTPS,
and it is two tests, one per gateway, each skipping on the other.
sdsmint reads the decrypted request, so it admits the named host and
answers 403 for another; the plain gateway cannot read TLS and no
address rule allows either destination, so it closes both connections
before a byte reaches an origin. E2E_EGRESS_MITM tells the suite which
gateway it is running against. The egressmitm suite adds the same 403
to its probe, next to the trust-bundle assertions it already makes.
2026-09-14 15:51:44 -07:00
Lior Lieberman 7034e79b10 atenet-egress: classify each tunnel and dial only what the policy allowed
The handler decides address rules at the CONNECT and every readable
request on its Host and the dialed address, so both gateways have to
give it those legs and act on what it decided.

The CONNECT chain's ext_proc may answer with dynamic metadata, and the
gateway admits exactly one namespace of it, dev.ate.egress. A
set_filter_state filter placed after the ext_proc copies
dev.ate.egress:passthrough_destination into Envoy's
envoy.network.transport_socket.original_dst_address and shares it with
the upstream connection. That is the whole mechanism carrying the
decision from the process that made it to the socket that acts on it,
with no header anywhere in between.

Inside the tunnel a connection lands on one of three chains, chosen by
tls_inspector and http_inspector under the 1s timeout that already
governs sdsmint:

  * egress_cleartext, on both gateways, matches raw_buffer with an HTTP
    ALPN and runs the policy ext_proc on every request. Its answer,
    dev.ate.egress:dial, picks the route: name goes to the
    dynamic_forward_proxy cluster, which resolves the Host, and address
    to an ORIGINAL_DST cluster fed by the same filter state as the
    passthrough chains, so the bytes go to what the matching rule
    checked. There is no route without a dial, and the answer clears the
    route cache, because Envoy picks a route before the filter runs. The
    chain accepts HTTP/1.0, which http_inspector steers here and the
    codec used to refuse before any filter ran.
  * egress_tls_mitm, sdsmint only, is that same per-request check and
    routing on the requests the gateway decrypted. Both of its dials
    re-originate TLS with the SNI and certificate check taken from the
    Host, so a Host the dialed origin cannot prove fails the handshake.
    Its filter sits below the existing #ATE_MITM_EXTPROC_FILTER markers,
    which keep working for an external processor spliced in by the
    installer.
  * the passthrough chains take what is left: on the plain gateway TLS it
    will not terminate, and on both gateways anything the inspectors
    could not classify, a client that said nothing before the timeout
    included. They run no filter. Their tcp_proxy points at an
    ORIGINAL_DST cluster with no address of its own, so it dials whatever
    the CONNECT decision put in filter state and closes the connection
    when there is nothing there. A destination no address rule allowed is
    therefore unreachable for traffic the gateway cannot read, which is
    what makes deferring a decision to the request legs safe.

    There is one such chain per transport protocol rather than a single
    chain matching nothing, because Envoy buckets filter chains by
    transport protocol and never falls out of a bucket that exists. With
    egress_cleartext holding the raw_buffer bucket and filling it with
    HTTP application protocols alone, a chain matching nothing would be
    unreachable and every opaque or unclassified connection would be
    closed as no_filter_chain_match, address rule or not. A manifest test
    now pins the invariant that each transport protocol the inspectors
    set has a chain, and that each has one catching what they could not
    name.

The plain gateway gains the inner listener it never had; until now it
spliced the tunnel straight to the IP:port and could police nothing
inside it. Its TLS stays end to end encrypted, since there is no MITM CA
there, so what it can enforce for HTTPS is the address, which the
manifest says in as many words. Server-speaks-first protocols pay the 1s
sniff timeout on this gateway now, as they already did on sdsmint.

The outer hop's connection pool is keyed per actor through
envoy.network.upstream_server_name, set from the same verified
certificate as the identity. Envoy captures shared filter state when a
pool is created and a string object takes no part in the pool's key, so
without that key a second actor's tunnel could be handed a pool carrying
the first actor's identity. The clusters that target an internal
listener also cap themselves at one request per connection.

The manifest tests pin the chain names, the attributes each ext_proc
requests, the metadata namespace every leg may answer in, the filter
that turns the CONNECT leg's answer into the passthrough address and the
absence of any other writer of that key, the two routes each request leg
has and the clusters they select, and a distinct stat_prefix per
ext_proc, because a mismatch fails closed at runtime and nothing else
would notice.
2026-09-14 15:51:44 -07:00
Lior Lieberman f980a57d83 atenet: enforce the actor's EgressPolicy in the egress ext_proc handler
The egress handler so far authenticated the actor behind a CONNECT and
let everything through. It now authorizes the traffic against the actor's
EgressPolicy, on the leg where the destination is known. Which leg that
is comes from the filter chain name the dataplane asserts, the same
signal the mux already dispatches on.

  * egress, the outer CONNECT, is the only leg that sees a certificate,
    and the first to see the address the actor's kernel dialed. It keeps
    authenticating that certificate and decides the address rules,
    ip_blocks and all, against that address, once per connection. When
    one of them allows it the handler hands the address back as dynamic
    metadata, which is what lets the gateway dial it for traffic it
    cannot read and what tells the legs inside what was dialed. When
    none does but the policy names hostnames, the tunnel is opened and
    the requests inside are left to the legs that will see a name. When
    neither is true nothing inside the tunnel could ever be allowed, so
    the CONNECT is refused. A dataplane that calls out for the CONNECT
    alone sends no chain name and has no such legs behind it, so it gets
    the whole decision here and is refused rather than deferred.

  * egress_tls_mitm and egress_cleartext are the legs the gateway can
    read. Every request is decided the way the API describes: the rules
    in order, over the request's Host and the address the actor dialed,
    first match wins. It runs per request because the Host can change
    between requests on one connection. The answer also says where the
    request goes, so the bytes reach what the matching rule checked: a
    hostname match is resolved and dialed by name, an address or all
    match goes to the dialed address. Sending an address match by name
    would let an actor put any Host on an allowed address; sending a
    hostname match to the dialed address would turn the rule into a
    header check, since the actor picks both. A Host header that
    disagrees with :authority is refused rather than policed on one name
    and dialed on the other.

Identity on the request legs is the actor's SPIFFE ID, which the outer
chain sets as filter state from the certificate Envoy verified and
shares across the internal-listener hop; nothing inside the tunnel can
write it, and a callout without it is refused. The certificate's URI SAN
must name the same actor as its ActorIdentity extension, because the
CONNECT authenticates on the extension while the request legs attribute
traffic to the SAN.

Policies are read through the per-actor cache, whose TTL is a new
--egress-policy-cache-ttl flag defaulting to 10s, so a request inside a
tunnel is not a control-plane round trip.

Credential injection is not implemented yet. A matched rule that declares
one denies with 501 rather than forwarding without the secret the policy
promised.
2026-09-14 15:51:44 -07:00
Lior Lieberman 670353912a atenet: add a per-actor EgressPolicy cache for the egress gateway
The egress ext_proc handler is about to authorize every request inside
every tunnel against the actor's EgressPolicy. Fetching that policy from
ateapi per request would put the control plane on the data path, so it
is read through a per-actor cache instead.

The TTL is exactly how stale a decision can be: a create, update or
delete is visible to new requests within one TTL, and a deletion becomes
a deny. "No policy" is cached like a policy is, because deny-by-default
is the answer a flood of refused requests hits. A failed fetch is never
cached, so an ateapi outage fails only the requests it overlaps.
Concurrent fetches of one actor's policy are collapsed into a single
call, and that call outlives the caller that started it, so the burst an
actor sends the moment it wakes up costs one round trip. Expired entries
are removed once the live set has doubled, so a gateway that has seen many
actors does not keep a policy for each of them forever.
2026-09-14 15:51:44 -07:00
Lior Lieberman e517df95d4 Add the egresspolicy evaluator and share the actor SPIFFE ID format
The egress gateway is about to enforce EgressPolicy, which means two
components now interpret the same hostname patterns and CIDRs: ateapi
when it validates a policy, and atenet when it matches one. Putting the
parsers in internal/egresspolicy and having ateapi's custom validators
call them keeps "accepted by the API" and "enforceable by the gateway"
the same set by construction. The package also compiles a policy once
and evaluates it the way the API describes, in rule order over whatever
the caller knows of the destination, first match wins: a hostname rule
never matches a destination with no name, and an IP rule never matches
one whose address the caller does not know. The decision says which
kind of rule matched, because the gateway sends the bytes to what that
rule checked. HasHostnameRules answers the question that follows, which
is whether a decision point holding only an address should refuse
outright or leave the destination to one that will see a name.

Hostname patterns now reject any name whose last label is all digits,
per RFC 1123 section 2.1, instead of consulting an IP parser. That is
the same set as before for every real address and additionally refuses
dotted-quad spellings such as "01.2.3.4" that no parser accepts.

The actor SPIFFE ID moves to internal/resources with a parser next to
the builder: ateapi mints it into the certificate's URI SAN and the
egress gateway reads it back as the actor identity on the legs inside
the tunnel, so both ends need the one definition.
2026-09-14 15:51:44 -07:00
Lior Lieberman 80b652d1c6 remove https-h2 guard 2026-08-28 11:26:28 -07:00
Lior Lieberman 0f96d20a9e align grpc tests 2026-08-28 11:26:28 -07:00
Lior Lieberman b039c2f706 atenet: offer h2 on the HTTPS ingress and mirror the protocol to actors, behind a flag 2026-08-28 11:26:28 -07:00
Lior Lieberman 24cc538dc1 fixes 2026-08-10 23:28:07 -07:00
Lior Lieberman 4e462b05e1 egress: make verify-egress-demo.sh robust and executable
The gateway access-log assertion read the log once, immediately after curl
returned, and failed if it saw nothing new. That is racy twice over:

  * Envoy emits the CONNECT access-log entry asynchronously, so for an
    external destination it can land seconds after the actor's response.
  * The Actor's HTTP client keeps the tunnel alive. A repeat fetch to a host
    it already reached rides the open tunnel and produces no new entry at
    all, so a run against a warm actor failed even though egress was working.

Poll for the new entry, and fall back to any tunnel already open for this
actor's SAN before declaring failure. Also read -c envoy explicitly (the pod
also runs the ext-proc sidecar) and mark the script executable, as every other
directly-invoked script in hack/ is.
2026-08-10 23:28:07 -07:00
Lior Lieberman c9a27e5230 apply feedback 08/06 2026-08-10 23:28:07 -07:00
Lior Lieberman 359036a7fc fixes 2026-08-10 23:28:07 -07:00
Lior Lieberman 2f05a92f59 egress: add the egress demo and e2e coverage
demos/egress is a small Actor that fetches a URL it is given and echoes the
upstream status and body back, which makes the egress path observable from
outside the sandbox. hack/install-demo-egress.sh registers it as a
--deploy-demo-egress fixture and hack/verify-egress-demo.sh drives it and
checks the atenet-egress logs for the corresponding authorized CONNECT.

TestActorEgress in the networking suite covers the same path automatically:
it creates an Actor from the demo template, POSTs a fetch request through
atenet-router, and asserts 200. The suite's actor helper is parameterised by
template so the ingress test keeps using the counter fixture.
2026-08-10 23:28:07 -07:00
Lior Lieberman dd764c1e7c egress: add an Envoy egress gateway.
The gateway terminates actor's CONNECT request.
It requires downstream mTLS, so only a worker's atunnel can reach it. The gateway consists of an Envoy and an `atenet router --standalone` ext_proc sidecar.
2026-08-10 23:28:07 -07:00
Lior Lieberman 0909650285 ActorIdentity: protos, cert issuance, actorIdentity x509 extension (#670)
Also Fixes #575 

(1) mintCert + MintJWT proto changes
(2) validate that atelet is the one that calls actorIdentity Service. 
(3) verify that this atelet is only requesting a cert for an actor on
its node.
(4) verify that the actor is still running
(5) create another substatex509 extension for actor identities

JWTs needs more love, lots as TODO there

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-01 19:23:53 -07:00
Lior Lieberman 860250bf71 diagram 2026-07-31 12:51:48 -07:00
Lior Lieberman 04c299ade7 add todo comment 2026-07-31 12:51:48 -07:00
Lior Lieberman 6e2a59cc57 address feedback 2026-07-31 12:51:48 -07:00
Lior Lieberman d9ac4e24e6 networking: update docs and add an ingress e2e suite
Bring the architecture and threat-model docs in line with the atunnel
ingress path: the router no longer rewrites :authority to a worker pod IP
and forwards over plaintext port 80, it opens an mTLS tunnel to atunnel on
worker port 443, which forwards to the Actor over its private veth.

Drop cmd/atenet/atenet-diagram.png. It predates the ext_proc/ORIGINAL_DST
design and is now wrong in the part that matters most. The README section it
illustrated gains a short accurate note about the upstream hop instead of a
dangling image reference.

Add an e2e suite covering the change end to end: TestActorDirectAccess
asserts that the worker pod's port 80 is no longer a reachable Actor ingress
path (the DNAT rule is gone) and that the same Actor still answers /readyz
through atenet-router over the atunnel mTLS hop. It uses the counter demo as
its fixture, so it only needs --deploy-demo-counter.

internal/e2e/testmain.go picks up ParseSkippedFlags so that `go test` flags
(-run, -v, ...) survive pflag parsing and reach the suite.
2026-07-31 12:51:48 -07:00
Lior Lieberman 20d46051ec ingress: route actor ingress through the atunnel mTLS server
The router used to rewrite :authority to the actor's worker pod IP and let
the dynamic_forward_proxy cluster resolve it, reaching the actor over
plaintext pod-IP:80. The worker no longer DNATs pod-IP:80 to the actor, so
that path is gone.

Instead the ext_proc resolves the actor to its worker's atunnel ingress
address (IP:443) and puts it in x-ate-original-dst. An ORIGINAL_DST cluster
dials exactly that address, which leaves the request Host as the actor DNS
name -- atunnel needs it to authorize the active actor. The header is set
with OVERWRITE_IF_EXISTS_OR_ADD so a client-supplied value can never
influence the address Envoy dials.

The upstream hop is mTLS: the cluster presents the router's podidentity
credential bundle as its client cert and validates the atunnel server
against the podidentity trust bundle. Validation matches the SPIFFE URI SAN
prefix rather than the dialed pod IP, because the atunnel cert carries only
a spiffe:// URI SAN and Envoy's default SAN check against an ephemeral pod
IP would never match. atunnel in turn only accepts
spiffe://cluster.local/ns/ate-system/sa/atenet-router.

The dynamic_forward_proxy cluster, HTTP filter and DNS cache config are
removed along with the :authority rewrite.
2026-07-31 12:51:48 -07:00
Lior Lieberman cc85887613 worker: serve actor ingress through atunnel
Every worker pod now hosts an atunnel ingress server on :443 and an
atunnel egress listener, both long-lived, with per-activation
Activate/Deactivate bracketing every Run, Restore and Checkpoint.

The nftables rules ateom installs change accordingly:

  * The pod-IP:80 -> actor-veth:80 DNAT is gone. Worker port 80 is no
    longer an Actor ingress path; the only way in is the mTLS listener
    on :443, which authorizes the caller and checks the Actor is the one
    currently assigned to this worker.

  * A prerouting REDIRECT sends Actor TCP egress to the local atunnel
    egress listener, preserving SO_ORIGINAL_DST. It is only installed
    when the activation carries an egress gateway address, which nothing
    populates yet, so the masquerade path is unchanged and Actor egress
    behaves exactly as before.

The ateom Run/Restore protos gain the two fields that arm that path:
egress_gateway_address, which decides whether the redirect is installed
at all, and actor_version, the Actor resource version ate-api observed
when assigning the worker, which atunnel asserts to the egress gateway
as a lower bound on trustworthy Actor metadata. Both are consumed here
and left unset. atelet and ate-api start populating them in the egress
gateway change, which is the point at which they mean anything, so this
change adds no new requirement to the atelet wire contract.

atecontroller gives worker pods the podidentity credential + trust
bundles (the atunnel server identity), the servicedns trust bundle, and
container port 443. podidentitysigner now issues certs with
ExtKeyUsageServerAuth as well as ClientAuth, without which the worker
cannot present its podidentity cert as a TLS server cert and the
gateway handshake fails.
2026-07-31 12:51:48 -07:00
Lior Lieberman b1bd558aba atunnel: add mTLS tunnel component for actor traffic
atunnel is the in-worker component that carries Actor traffic in both
directions over authenticated channels.

  * Server: an mTLS HTTPS listener that terminates connections from the
    ingress gateway, authorizes the peer by its SPIFFE identity, checks
    that the request targets the Actor currently activated on this
    worker, and reverse-proxies to the Actor over the private veth.

  * Egress: an activation-aware TCP proxy for transparently intercepted
    Actor connections. It resolves the original destination via
    SO_ORIGINAL_DST and forwards through an EgressDialer, asserting the
    Actor identity to the far end. Dormant until an egress gateway is
    configured; the client and dialer land here so the package is
    reviewed as one unit.

Activate/Deactivate bracket an Actor activation so traffic is only
carried while an Actor is assigned to the worker, and Deactivate drains
in-flight streams before the Actor network is torn down.
2026-07-31 12:51:48 -07:00
Lior Lieberman 1b329c9872 add otel export timeout 2026-07-31 08:22:50 -07:00
Lior Lieberman 543fd3c974 fix flaky TestPlatformMetrics test 2026-07-31 08:22:50 -07:00
Lior Lieberman 92f1aa7276 Rename SessionIdentity to ActorIdentity
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.

API surface:
  service SessionIdentity        -> ActorIdentity
  MintJWTRequest.session_id      -> actor_id
  MintJWTResponse.session_jwt    -> actor_jwt
  MintCertRequest.session_id     -> actor_id
  MintCertResponse.session_certificates -> actor_certificates

Go packages:
  cmd/ateapi/internal/sessionidentity -> actoridentity
  cmd/ateapi/internal/sessionidjwt    -> actoridjwt

Flags and cluster resources:
  --session-id-jwt-pool -> --actor-id-jwt-pool
  --session-id-ca-pool  -> --actor-id-ca-pool
  Secrets, volumes and mount paths renamed to match, in both
  manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
  (--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).

Two credential identity values change with the rename:
  JWT issuer https://broker.agentic-substrate-session-id-broker.svc
          -> https://broker.agentic-substrate-actor-id-broker.svc
  SPIFFE ID spiffe://substrate-session.local/app/../session/..
          -> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.

BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.
2026-07-30 07:34:29 -07:00
Lior Lieberman 7c6d98629e Regenerate stale Python proto stubs
stubs were last regenerated on 2026-07-17 and had fallen behind ateapi.proto.
2026-07-30 07:34:29 -07:00