The passthrough chain from #2019 dialed the ORIGINAL_DST filter state,
which the CONNECT leg filled with the address the actor connected to.
The SNI picked the chain, the actor's own resolution picked the
destination, so a passthrough rule for one name let an actor reach any
IP by claiming that name in the ClientHello.
We fix it by:
- The passthrough chain is now `sni_dynamic_forward_proxy` then
`tcp_proxy` to a new raw forward-proxy cluster,
`egress_forward_proxy_passthrough`, on the shared egress_dns_cache. The
gateway resolves the SNI itself and sends the bytes to it.
- The CONNECT leg answers the dialed port under
dev.ate.egress:dialed_port; the outer chain copies it into
envoy.upstream.dynamic_port, shared with the inner listener, which both
the SNI filter and the cluster read before their configured port.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
We should not override the whole actor trust store with SSL_CERT_DIR,
removing that to defer using SSL_CERT_FILE, which is additive.
Additionally, this does not cover all languages and runtimes, adding
guidance for other env vars required and other runtimes that need union
of the trust stores in a single file.
In a followup, we are going to remove all "sdsmint" and "plain gateway"
language and just have one gateway thats capable of minting or not based
on the sni.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
do not merge, not a draft cause i do want ci running
implements:
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching
tls_passhtrough is a followup.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Prefix the cluster-wide egress PEP address flag with `default-`
~'experimental-to-be-removed-' to indicate it is a temporary
cluster-level knob that will be removed~ once per-actor or per-atespace
egress gateway configuration is supported (#1591).
~I kept --egress-gateway-address as a deprecated alias for backward
compatibility but I really dont think we should. ~
I removed it, we dont need backward comp residuals pre GA
EDIT: I chatted with tim and taahir a the feedback was that its not
really experimental and can not really be removed. This flag would have
to be set for every GA user, and its not a nice experimental add on. We
dont have time to get into per-actor or per-atespace API discussions
pre-GA therfore we gotta have to live with it and discourage later if we
have a better story for that.
Related #1591
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
/cc @bowei @EItanya
**BREAKING** API changes prior in preparation for GA.
~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow
- `EgressRule` is now a union of protocols: `http`, `https`,
`tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
`tls_passthrough`. Same wildcard grammar as before, documented once on
`HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
once on `HTTPRule.ports`.
- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
placeholder value. Requests without the header pass unchanged.
- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.
- `google/protobuf/empty.proto` import dropped; nothing uses it now.
Questions
- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
block is intercepted. is that clear enough?
follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
the rule conflict check) and regenerate `zz_generated.validation.go`
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes: #823 (there are probably more issues that i need to find)
The router README gets the gateway-side account, a leg at a time,
including the metadata that carries the CONNECT leg's decision to the
chain that dials. The egress demo README stops calling destination
authorization a follow-up.
The egress gateway now refuses every tunnel for an actor without an
EgressPolicy, so every e2e actor that egresses needs one. A new
e2e.EnsureEgressPolicy creates or replaces an actor's policy, and the
networking and egressmitm suites give their actors an allow-all rule at
creation, which keeps every existing egress test meaning what it meant.
The cases cover the policy itself. An actor whose policy names one
hostname is refused with 403 and "egress denied" when it asks for
another over cleartext HTTP, and for the same origin by address, since
an IP-literal Host names no host. An actor with no policy gets no tunnel
at all, which the demo app reports as a 502 from a failed connect. A
policy naming the origin's address admits a plain HTTP request to it
both by address and by name, because the request is decided on the
address the actor dialed, and it gets no tunnel for a name that resolves
anywhere else. Over HTTPS an address policy admits example.com on both
gateways: the plain gateway by address at the CONNECT, sdsmint on the
decrypted request, which it then sends to the dialed address with the
certificate checked against the name.
The one case where the gateways differ is a hostname policy over HTTPS,
and it is two tests, one per gateway, each skipping on the other.
sdsmint reads the decrypted request, so it admits the named host and
answers 403 for another; the plain gateway cannot read TLS and no
address rule allows either destination, so it closes both connections
before a byte reaches an origin. E2E_EGRESS_MITM tells the suite which
gateway it is running against. The egressmitm suite adds the same 403
to its probe, next to the trust-bundle assertions it already makes.
The handler decides address rules at the CONNECT and every readable
request on its Host and the dialed address, so both gateways have to
give it those legs and act on what it decided.
The CONNECT chain's ext_proc may answer with dynamic metadata, and the
gateway admits exactly one namespace of it, dev.ate.egress. A
set_filter_state filter placed after the ext_proc copies
dev.ate.egress:passthrough_destination into Envoy's
envoy.network.transport_socket.original_dst_address and shares it with
the upstream connection. That is the whole mechanism carrying the
decision from the process that made it to the socket that acts on it,
with no header anywhere in between.
Inside the tunnel a connection lands on one of three chains, chosen by
tls_inspector and http_inspector under the 1s timeout that already
governs sdsmint:
* egress_cleartext, on both gateways, matches raw_buffer with an HTTP
ALPN and runs the policy ext_proc on every request. Its answer,
dev.ate.egress:dial, picks the route: name goes to the
dynamic_forward_proxy cluster, which resolves the Host, and address
to an ORIGINAL_DST cluster fed by the same filter state as the
passthrough chains, so the bytes go to what the matching rule
checked. There is no route without a dial, and the answer clears the
route cache, because Envoy picks a route before the filter runs. The
chain accepts HTTP/1.0, which http_inspector steers here and the
codec used to refuse before any filter ran.
* egress_tls_mitm, sdsmint only, is that same per-request check and
routing on the requests the gateway decrypted. Both of its dials
re-originate TLS with the SNI and certificate check taken from the
Host, so a Host the dialed origin cannot prove fails the handshake.
Its filter sits below the existing #ATE_MITM_EXTPROC_FILTER markers,
which keep working for an external processor spliced in by the
installer.
* the passthrough chains take what is left: on the plain gateway TLS it
will not terminate, and on both gateways anything the inspectors
could not classify, a client that said nothing before the timeout
included. They run no filter. Their tcp_proxy points at an
ORIGINAL_DST cluster with no address of its own, so it dials whatever
the CONNECT decision put in filter state and closes the connection
when there is nothing there. A destination no address rule allowed is
therefore unreachable for traffic the gateway cannot read, which is
what makes deferring a decision to the request legs safe.
There is one such chain per transport protocol rather than a single
chain matching nothing, because Envoy buckets filter chains by
transport protocol and never falls out of a bucket that exists. With
egress_cleartext holding the raw_buffer bucket and filling it with
HTTP application protocols alone, a chain matching nothing would be
unreachable and every opaque or unclassified connection would be
closed as no_filter_chain_match, address rule or not. A manifest test
now pins the invariant that each transport protocol the inspectors
set has a chain, and that each has one catching what they could not
name.
The plain gateway gains the inner listener it never had; until now it
spliced the tunnel straight to the IP:port and could police nothing
inside it. Its TLS stays end to end encrypted, since there is no MITM CA
there, so what it can enforce for HTTPS is the address, which the
manifest says in as many words. Server-speaks-first protocols pay the 1s
sniff timeout on this gateway now, as they already did on sdsmint.
The outer hop's connection pool is keyed per actor through
envoy.network.upstream_server_name, set from the same verified
certificate as the identity. Envoy captures shared filter state when a
pool is created and a string object takes no part in the pool's key, so
without that key a second actor's tunnel could be handed a pool carrying
the first actor's identity. The clusters that target an internal
listener also cap themselves at one request per connection.
The manifest tests pin the chain names, the attributes each ext_proc
requests, the metadata namespace every leg may answer in, the filter
that turns the CONNECT leg's answer into the passthrough address and the
absence of any other writer of that key, the two routes each request leg
has and the clusters they select, and a distinct stat_prefix per
ext_proc, because a mismatch fails closed at runtime and nothing else
would notice.
The egress handler so far authenticated the actor behind a CONNECT and
let everything through. It now authorizes the traffic against the actor's
EgressPolicy, on the leg where the destination is known. Which leg that
is comes from the filter chain name the dataplane asserts, the same
signal the mux already dispatches on.
* egress, the outer CONNECT, is the only leg that sees a certificate,
and the first to see the address the actor's kernel dialed. It keeps
authenticating that certificate and decides the address rules,
ip_blocks and all, against that address, once per connection. When
one of them allows it the handler hands the address back as dynamic
metadata, which is what lets the gateway dial it for traffic it
cannot read and what tells the legs inside what was dialed. When
none does but the policy names hostnames, the tunnel is opened and
the requests inside are left to the legs that will see a name. When
neither is true nothing inside the tunnel could ever be allowed, so
the CONNECT is refused. A dataplane that calls out for the CONNECT
alone sends no chain name and has no such legs behind it, so it gets
the whole decision here and is refused rather than deferred.
* egress_tls_mitm and egress_cleartext are the legs the gateway can
read. Every request is decided the way the API describes: the rules
in order, over the request's Host and the address the actor dialed,
first match wins. It runs per request because the Host can change
between requests on one connection. The answer also says where the
request goes, so the bytes reach what the matching rule checked: a
hostname match is resolved and dialed by name, an address or all
match goes to the dialed address. Sending an address match by name
would let an actor put any Host on an allowed address; sending a
hostname match to the dialed address would turn the rule into a
header check, since the actor picks both. A Host header that
disagrees with :authority is refused rather than policed on one name
and dialed on the other.
Identity on the request legs is the actor's SPIFFE ID, which the outer
chain sets as filter state from the certificate Envoy verified and
shares across the internal-listener hop; nothing inside the tunnel can
write it, and a callout without it is refused. The certificate's URI SAN
must name the same actor as its ActorIdentity extension, because the
CONNECT authenticates on the extension while the request legs attribute
traffic to the SAN.
Policies are read through the per-actor cache, whose TTL is a new
--egress-policy-cache-ttl flag defaulting to 10s, so a request inside a
tunnel is not a control-plane round trip.
Credential injection is not implemented yet. A matched rule that declares
one denies with 501 rather than forwarding without the secret the policy
promised.
The egress ext_proc handler is about to authorize every request inside
every tunnel against the actor's EgressPolicy. Fetching that policy from
ateapi per request would put the control plane on the data path, so it
is read through a per-actor cache instead.
The TTL is exactly how stale a decision can be: a create, update or
delete is visible to new requests within one TTL, and a deletion becomes
a deny. "No policy" is cached like a policy is, because deny-by-default
is the answer a flood of refused requests hits. A failed fetch is never
cached, so an ateapi outage fails only the requests it overlaps.
Concurrent fetches of one actor's policy are collapsed into a single
call, and that call outlives the caller that started it, so the burst an
actor sends the moment it wakes up costs one round trip. Expired entries
are removed once the live set has doubled, so a gateway that has seen many
actors does not keep a policy for each of them forever.
The egress gateway is about to enforce EgressPolicy, which means two
components now interpret the same hostname patterns and CIDRs: ateapi
when it validates a policy, and atenet when it matches one. Putting the
parsers in internal/egresspolicy and having ateapi's custom validators
call them keeps "accepted by the API" and "enforceable by the gateway"
the same set by construction. The package also compiles a policy once
and evaluates it the way the API describes, in rule order over whatever
the caller knows of the destination, first match wins: a hostname rule
never matches a destination with no name, and an IP rule never matches
one whose address the caller does not know. The decision says which
kind of rule matched, because the gateway sends the bytes to what that
rule checked. HasHostnameRules answers the question that follows, which
is whether a decision point holding only an address should refuse
outright or leave the destination to one that will see a name.
Hostname patterns now reject any name whose last label is all digits,
per RFC 1123 section 2.1, instead of consulting an IP parser. That is
the same set as before for every real address and additionally refuses
dotted-quad spellings such as "01.2.3.4" that no parser accepts.
The actor SPIFFE ID moves to internal/resources with a parser next to
the builder: ateapi mints it into the certificate's URI SAN and the
egress gateway reads it back as the actor identity on the legs inside
the tunnel, so both ends need the one definition.
The gateway access-log assertion read the log once, immediately after curl
returned, and failed if it saw nothing new. That is racy twice over:
* Envoy emits the CONNECT access-log entry asynchronously, so for an
external destination it can land seconds after the actor's response.
* The Actor's HTTP client keeps the tunnel alive. A repeat fetch to a host
it already reached rides the open tunnel and produces no new entry at
all, so a run against a warm actor failed even though egress was working.
Poll for the new entry, and fall back to any tunnel already open for this
actor's SAN before declaring failure. Also read -c envoy explicitly (the pod
also runs the ext-proc sidecar) and mark the script executable, as every other
directly-invoked script in hack/ is.
demos/egress is a small Actor that fetches a URL it is given and echoes the
upstream status and body back, which makes the egress path observable from
outside the sandbox. hack/install-demo-egress.sh registers it as a
--deploy-demo-egress fixture and hack/verify-egress-demo.sh drives it and
checks the atenet-egress logs for the corresponding authorized CONNECT.
TestActorEgress in the networking suite covers the same path automatically:
it creates an Actor from the demo template, POSTs a fetch request through
atenet-router, and asserts 200. The suite's actor helper is parameterised by
template so the ingress test keeps using the counter fixture.
The gateway terminates actor's CONNECT request.
It requires downstream mTLS, so only a worker's atunnel can reach it. The gateway consists of an Envoy and an `atenet router --standalone` ext_proc sidecar.
Also Fixes#575
(1) mintCert + MintJWT proto changes
(2) validate that atelet is the one that calls actorIdentity Service.
(3) verify that this atelet is only requesting a cert for an actor on
its node.
(4) verify that the actor is still running
(5) create another substatex509 extension for actor identities
JWTs needs more love, lots as TODO there
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Bring the architecture and threat-model docs in line with the atunnel
ingress path: the router no longer rewrites :authority to a worker pod IP
and forwards over plaintext port 80, it opens an mTLS tunnel to atunnel on
worker port 443, which forwards to the Actor over its private veth.
Drop cmd/atenet/atenet-diagram.png. It predates the ext_proc/ORIGINAL_DST
design and is now wrong in the part that matters most. The README section it
illustrated gains a short accurate note about the upstream hop instead of a
dangling image reference.
Add an e2e suite covering the change end to end: TestActorDirectAccess
asserts that the worker pod's port 80 is no longer a reachable Actor ingress
path (the DNAT rule is gone) and that the same Actor still answers /readyz
through atenet-router over the atunnel mTLS hop. It uses the counter demo as
its fixture, so it only needs --deploy-demo-counter.
internal/e2e/testmain.go picks up ParseSkippedFlags so that `go test` flags
(-run, -v, ...) survive pflag parsing and reach the suite.
The router used to rewrite :authority to the actor's worker pod IP and let
the dynamic_forward_proxy cluster resolve it, reaching the actor over
plaintext pod-IP:80. The worker no longer DNATs pod-IP:80 to the actor, so
that path is gone.
Instead the ext_proc resolves the actor to its worker's atunnel ingress
address (IP:443) and puts it in x-ate-original-dst. An ORIGINAL_DST cluster
dials exactly that address, which leaves the request Host as the actor DNS
name -- atunnel needs it to authorize the active actor. The header is set
with OVERWRITE_IF_EXISTS_OR_ADD so a client-supplied value can never
influence the address Envoy dials.
The upstream hop is mTLS: the cluster presents the router's podidentity
credential bundle as its client cert and validates the atunnel server
against the podidentity trust bundle. Validation matches the SPIFFE URI SAN
prefix rather than the dialed pod IP, because the atunnel cert carries only
a spiffe:// URI SAN and Envoy's default SAN check against an ephemeral pod
IP would never match. atunnel in turn only accepts
spiffe://cluster.local/ns/ate-system/sa/atenet-router.
The dynamic_forward_proxy cluster, HTTP filter and DNS cache config are
removed along with the :authority rewrite.
Every worker pod now hosts an atunnel ingress server on :443 and an
atunnel egress listener, both long-lived, with per-activation
Activate/Deactivate bracketing every Run, Restore and Checkpoint.
The nftables rules ateom installs change accordingly:
* The pod-IP:80 -> actor-veth:80 DNAT is gone. Worker port 80 is no
longer an Actor ingress path; the only way in is the mTLS listener
on :443, which authorizes the caller and checks the Actor is the one
currently assigned to this worker.
* A prerouting REDIRECT sends Actor TCP egress to the local atunnel
egress listener, preserving SO_ORIGINAL_DST. It is only installed
when the activation carries an egress gateway address, which nothing
populates yet, so the masquerade path is unchanged and Actor egress
behaves exactly as before.
The ateom Run/Restore protos gain the two fields that arm that path:
egress_gateway_address, which decides whether the redirect is installed
at all, and actor_version, the Actor resource version ate-api observed
when assigning the worker, which atunnel asserts to the egress gateway
as a lower bound on trustworthy Actor metadata. Both are consumed here
and left unset. atelet and ate-api start populating them in the egress
gateway change, which is the point at which they mean anything, so this
change adds no new requirement to the atelet wire contract.
atecontroller gives worker pods the podidentity credential + trust
bundles (the atunnel server identity), the servicedns trust bundle, and
container port 443. podidentitysigner now issues certs with
ExtKeyUsageServerAuth as well as ClientAuth, without which the worker
cannot present its podidentity cert as a TLS server cert and the
gateway handshake fails.
atunnel is the in-worker component that carries Actor traffic in both
directions over authenticated channels.
* Server: an mTLS HTTPS listener that terminates connections from the
ingress gateway, authorizes the peer by its SPIFFE identity, checks
that the request targets the Actor currently activated on this
worker, and reverse-proxies to the Actor over the private veth.
* Egress: an activation-aware TCP proxy for transparently intercepted
Actor connections. It resolves the original destination via
SO_ORIGINAL_DST and forwards through an EgressDialer, asserting the
Actor identity to the far end. Dormant until an egress gateway is
configured; the client and dialer land here so the package is
reviewed as one unit.
Activate/Deactivate bracket an Actor activation so traffic is only
carried while an Actor is assigned to the worker, and Deactivate drains
in-flight streams before the Actor network is torn down.
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.
API surface:
service SessionIdentity -> ActorIdentity
MintJWTRequest.session_id -> actor_id
MintJWTResponse.session_jwt -> actor_jwt
MintCertRequest.session_id -> actor_id
MintCertResponse.session_certificates -> actor_certificates
Go packages:
cmd/ateapi/internal/sessionidentity -> actoridentity
cmd/ateapi/internal/sessionidjwt -> actoridjwt
Flags and cluster resources:
--session-id-jwt-pool -> --actor-id-jwt-pool
--session-id-ca-pool -> --actor-id-ca-pool
Secrets, volumes and mount paths renamed to match, in both
manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
(--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).
Two credential identity values change with the rename:
JWT issuer https://broker.agentic-substrate-session-id-broker.svc
-> https://broker.agentic-substrate-actor-id-broker.svc
SPIFFE ID spiffe://substrate-session.local/app/../session/..
-> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.
BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.