18 Commits
Author SHA1 Message Date
Taahir Ahmed fe82dae3f3 serverboot.Fatal: Allow additional slog arguments (#1908)
Prefactor for unified config file.
2026-09-26 01:29:52 +00:00
Tim Bai 93e691d9b1 serverboot: give the ateoms a LoggerProvider over the relay (#1857)
Step 2 of #1748, plus the consumer rule from step 1b.

`LoggingOptions` gains `ExporterConn` and `RelayCapable`, mirroring
`TracingOptions`. Both ateoms initialize logging after metrics with the
relay connection and defer a nil-guarded shutdown. `OTEL_LOGS_EXPORTER`
still defaults to `none`, so no ateom emits a record; this is the
provider the actor events land on when emission moves.

A relay-capable log resource carries `ate.otlp.relay` like traces and
metrics do. `newLoggerProvider` is split from `InitLogging` so the test
asserts that on an emitted record's resource, the way
`TestMeterProviderRelayAttribute` does for metrics.

`docs/observability.md` states the consumer rule that replaced the relay
denylist dropped in #1800: a lifecycle record is authoritative only
under ateapi's resource, because the relay admits only ateom resources.

Follow-up for the emission step, not here: the controller does not
inject `OTEL_LOGS_EXPORTER` into worker pods, so no environment
exercises log-over-relay yet.
2026-09-24 20:40:19 +00:00
Krisztian F 5f7c690108 (feat): Actor lifecycle events over otlp (#1658)
## What this does

ateapi writes a record every time an actor changes state. Until now
those records only went to the pod's stdout, and nothing reads stdout.
This sends the same records to a collector as OTLP log events.

Actor name and uid cannot be metric labels (too many values), and traces
are sampled at 1%. So these records are the only way to answer "what
state is this actor in, and since when".

## Changes

- `serverboot.InitLogging` sets up a LoggerProvider, next to the
existing tracer and meter ones.
- New `internal/actorevent` package builds the log records.
- ateapi emits at the two places that already write the stdout records.
- Two event names: `ate.actor.state_changed` and `ate.actor.crashed`.
- Both names are registered in `docs/metrics/registry/events.yaml`, so
`make verify` checks them.
- kind gets a logs pipeline and a count connector. The e2e suite reads
the counts back.
- Docs updated. `otel-collector.md` said substrate has no
LoggerProvider, which is no longer true.

## Opt-in

`OTEL_LOGS_EXPORTER` defaults to `none`. Only the kind overlay sets it
to `otlp`. The base ConfigMap is untouched, so no deployed environment
changes when this merges.

## Notes on the design

- **No slog bridge.** Only two call sites emit these records, so
emitting twice costs two lines. A bridge would also send every ateapi
log over the wire, could not set the event name, and would loop, because
SDK export errors are logged through slog.
- **Batching processor, not the simple one.** These records sit on the
actor resume path. A processor that exports inside the emit call would
add a blocking gRPC call there, so a slow collector would become control
plane latency.
- **Two event names, not one per state.** `ate.actor.state` already says
which transition happened. A crash gets its own name because it carries
two extra attributes and a higher severity.
- **Both copies are kept on purpose.** No collector in this repo reads
pod stdout, so nothing is duplicated today. `kubectl logs` keeps
working. If a filelog agent is ever added, drop one of the two. The
escape hatch is written down in `docs/metrics/substrate.yaml`.

## Dependencies

Adds `otel/log`, `otel/sdk/log` and `otlploggrpc`, all pinned at
v0.20.0. That is the release that matches the pinned `otel v1.44.0`.
v0.21.0 would pull the core modules to v1.45.0, which this change does
not need. The logs API has a
v1.47.0 release candidate upstream, so it is on its way to stable.

## Testing

- Unit tests for the exporter resolver, the record builder, and both
ateapi emit sites.
- The record builder test checks the attribute set matches what the
event name declares, in both directions.
- The ateapi tests check the OTLP record carries the same attributes as
the stdout record.
- Ran end to end on kind. Records arrive with the right event name,
severity, attributes, and with trace context on the record's own fields
rather than as attributes.
- Checked the off state too. With `OTEL_LOGS_EXPORTER` removed, the
collector receives no log records and stdout is unchanged.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-18 18:11:12 +00:00
Aditya ShantanuandAditya Shantanu 66300d5c47 workersync: gate worker registration on ateom readiness (#1243)
Fixes #1106. Takes over #1107 (adopting @Stevenjin8's approach and
commits from that PR — credit to him for the design; trailer omitted
only for CLA) and incorporates the review feedback it accumulated.

Workers were registered as schedulable the moment their pod had an IP —
long before ateom's gRPC server was listening, which is the scheduling
race in #1106 (actors dispatched to a cold ateom fail at dial time).
Now:

- ateom (gvisor + microvm) serves HTTP `/readyz` on a dedicated port
(8080) once its gRPC listener is up, flipping to 503 on SIGTERM so
draining workers stop receiving work.
- The worker Deployment probes that endpoint (replaces the earlier TCP
probe against the TLS port, which spammed handshake errors).
- `isWorkerEligible` requires `PodReady=True` in addition to an IP.
Eligibility gates **registration only** — a readiness flap never
deregisters a worker with a bound actor (pinned by
`TestSyncer_ReadinessFlapDoesNotDeregister`).
- A readiness server that cannot come up exits the process rather than
leaving a pod that looks alive but can never receive work.

Beyond #1107's head: dropped the pre-migration
`controlapi/syncer_test.go` copy (didn't compile after #1203 moved the
syncer), fixed the `workersync` fixtures for the new gate, and added
`TestSyncer_PodWithIPButNotReadyIsNotRegistered` (+ the flap test
above).

## Testing

- `go test ./...`: green, zero failures.
- `-race -count=10` on `workersync` + `serverboot`, `-race -count=5` on
`controllers`, `-race` on all of `controlapi` incl. functional tests:
green.
- Green under a hostile ambient env
(`KUBECTL_CONTEXT`/`KO_DOCKER_REPO`/`PROJECT_ID` set to garbage).

## Scope notes

- The one post-#1160 failure with a *post-restore* signature (run
32891995583: restored actor's own container `/healthz` timing out, actor
stuck RESUMING) is a different layer — ateom was demonstrably serving —
and is **not** claimed as fixed here. If it recurs, it deserves its own
issue. (The other suspected post-merge failure, run 32911674788, turned
out to be that PR's own x509 regression, not this flake.)
- Mixed-version caveat: this controller pointed at pods running a
pre-readyz ateom image leaves those workers unready/unregistered. The
supported install flows build both from the same source, so this only
matters for hand-rolled skew.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-08-26 17:32:03 -04:00
Krisztian F c3c6876985 (chore): consolidate logging to ateattr (#1035)
Actor logs used `ate.dev/actor_*` while spans and metrics use `ate.*`
registry in internal/ateattr. This PR makes ateattr the single source of
truth for everything telemetry-related.

- Renamed the six actor log labels onto the registry: 
- `ate.atespace`, `ate.actor.name`, `ate.actor.uid`,
`ate.template.namespace`, `ate.template.name`,
`ate.actor.container.name`
- Logs join traces now. Records set `trace_id`, `span_id` and
`trace_flags`, so you can go from Actor restored to the resume that
caused it. Our own lines only, not an actor's stdout: one goroutine
forwards a whole container stream and can't know which request produced
a given line. Per line correlation comes with #853.
- Actors can't fake platform labels. They already couldn't overwrite
ours, but they could invent new ones like `ate.tenant` that look
platform issued downstream. Anything under `ate.` from an actor is now
dropped.
- Fixed the asymmetry that was actually left: actor supplied label
values weren't stringified, and one non string value makes Cloud Logging
discard the labels for that whole entry.
- Note: the actor_uid bullet in the issue is stale, #841 fixed it
earlier. Lifecycle records still set five labels rather than six, on
purpose as they're about the actor, so no container produced them.

Fixes #886

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: krisztianfekete <git@krisztianfekete.org>
2026-08-19 11:02:10 -04:00
Keith Mattix II 89b83a5c89 Support CONNECT in atenet router (#715)
added CONNECT support for atenet ingress to support arbitrary actor ports
2026-08-14 16:12:27 -07:00
Chenyi Wang c178f5b74e ateom: export telemetry through an atelet unix-socket OTLP relay (#809)
ateom runs inside the worker pod that hosts the actor, and exported OTLP
straight to the collector over the pod's network. This adds a node-local
relay: ateom pushes OTLP/gRPC over a unix socket that atelet serves and
forwards to the collector, so a worker pod needs no network path of its
own to export spans and metrics.

Four things motivate it:

- Blast radius. The pod runs untrusted agent code, so allowing it egress
to the collector makes the collector reachable to anything that escapes
the sandbox. A unix socket cannot leave the node.
- Connection count. Worker pods are heavily oversubscribed; N ateoms per
node each held their own collector connection. They collapse into
atelet's single per-node one.
- Interference. ateom transparently redirects actor egress to its own
atunnel listener, and its own outbound traffic has to stay clear of the
rules it installs. A unix socket is not IP traffic.
- Shutdown loss. Teardown frees the actor's network and then the pod
goes away, which is when the spans describing teardown are still queued
in the batch processor. atelet outlives the worker pod.

The relay forwards the OTLP request verbatim rather than decoding and
re-exporting, so each ateom's own resource (service.name,
service.instance.id) survives instead of being absorbed into atelet's.

It is best-effort: an ateom that finds no socket at startup logs it and
exports directly to OTEL_EXPORTER_OTLP_ENDPOINT as before, so this is a
no-op for a cluster running an older atelet. atelet likewise declines to
serve a relay when no collector is configured, since it would accept
spans only to drop them. Both halves stay off with
--otlp-relay-socket="".

The socket lives in ateompath.BasePath, the hostPath already mounted at
the same path into atelet and into every ateom pod, so no new volume or
controller change is needed.

Also includes:
  - End-to-end tests covering the full serverboot-to-collector path.
  - Observability documentation updates for Jaeger tracing.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-14 14:33:14 -07:00
Krisztian F c155efd1ac feat(otel): onboard atecontroller to the OTLP path (#754)
atecontroller had no OTel at all. Dev-mode zap logger, and its
controller-runtime metrics were only available on a :8080 that we don't
scrape.

After this PR, logs go through the shared slog handler (plus a
`--log-level` flag to match the other binaries), and
controller-runtime's Prometheus registry is bridged onto the OTLP reader
so the reconcile/workqueue metrics actually reach the collector.
Filtering these out is a pipeline responsibility. Also added otelgrpc to
the ateapi client, which was untraced.

This unblocks #564 the workperpool metrics, cc @Angelawork, @JeffLuoo:
there's a working `MeterProvider` to use for
`ate.workerpool.desired_workers`/`ready_workers`.

Couple of things to mention for review:

-`InitMetricsPushOnly`, not `InitMetrics`, even though we do serve
:8080. That port is controller-runtime's own private registry, not the
global one `serverboot.metricsMux` serves, so a pull reader there would
collect into something we never expose.
- Bridge is pinned to v0.68.0 to match otelgrpc. Wanted to go to
v0.70.0, but that requires otel/sdk/metric 1.45.0 and pulls the whole
SDK up with it (406 vendor files instead of 81). Happy to do that bump
separately.
- zap/zapr fall out of go.mod since atecontroller was the last importer.
- I included an OTel collector image bump from the early 2024 (!) one to
latest, which was breaking exposing native histograms

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-06 10:25:03 -04:00
Krisztian F 15eecd02c9 Define a production-ready default tracing policy (#711)
Today everything outside ateapi hardcodes `ParentBased(NeverSample)`,
and setting `OTEL_TRACES_SAMPLER` does nothing because an explicit
sampler silences the SDK's env handling. This makes troubleshooting
production issues impossible via traces.

This PR gives every component a sane default and makes the standard OTel
env vars work:

- Control plane (ateapi, atelet, ateom) defaults to
`parentbased_traceidratio` 0.1, the router (data plane root) to 0.01,
per the discussion on #584.
- `OTEL_TRACES_SAMPLER` / `OTEL_TRACES_SAMPLER_ARG` override any of this
without a rebuild. Invalid values keep the component default and log a
warning instead of inheriting the SDK's fall-open-to-100% behavior.
- Envoy's `RandomSampling` is derived from the router's resolved policy,
so the two root decisions cannot drift.
- kubectl-ate without `--trace` no longer installs a tracer provider at
all: the old `NeverSample` provider injected `sampled=0`, which pinned
every parent based sampler downstream and would have defeated the server
side ratios. `--trace` still forces a full end to end trace.
- ate-controller propagates the two env vars to the ateom worker pods it
creates, same as the metric export vars.
- kind pins ateapi to `parentbased_always_on`, so the local Jaeger flow
keeps showing every API call.
- The agentgateway integration follows what we have above. its
`randomSampling` changes from `true` to `0.01` to match the data plane
default, though as static config it does not follow
`OTEL_TRACES_SAMPLER` overrides, so the two need adjusting together.


Verified on a kind cluster e2e manually. It resolved samplers logged at
startup, a `--trace` resume produced one trace across ateapi, atelet,
and ateom, an unsampled CLI call got picked up server side, Envoy
continued a sampled traceparent while sampling 0 of 30 parentless
requests at the 1% default, and an invalid env value fell back to the
component default.

Gating who may use `--trace` stays a separate follow-up (and a
discussion), and the more sophisticaed tracing policies belongs in a
collector, not in substrate.

Fixes #584
cc. @git286 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-04 11:31:04 -04:00
igooch d016eddcb1 serverboot: dynamic --log-level flag for the server binaries (#677)
Adds a dynamic `--log-level` flag (debug, info, warn, error;
case-insensitive) to
the server binaries, backed by a `slog.LevelVar` in `serverboot`.
Defaults to
`info` — no behavior change unless the flag is set.

- **serverboot**: the LevelVar behind `InitLogger`'s handler;
`SetLogLevel(string)`
to parse and set it (empty string = unset, a documented no-op);
`LogLevel()`
  returns a `slog.Leveler` for binaries that build their own handler.
- **atelet, ateapi**: new pflag, fatal on an invalid value.
- **ateom-gvisor, ateom-microvm**: flag accepted by the binary; note
their
container args are built by the WorkerPool controller with no
pass-through yet,
so wiring it is a follow-up — and that follow-up must roll out only
after this
change's ateom images are deployed (older images reject unknown flags
and would
  crashloop).
- **atenet (router + dns)**: replaces two pre-existing hand-rolled
`--log-level`
  parsers that silently fell back to `info` on a bad value; both now use
  `serverboot.SetLogLevel` and `serverboot.InitLogger`, which also adds
  `ate.dev/trace-id` correlation to atenet logs.
- **Kind manifests**: atelet runs at `--log-level=debug` on kind
(dev/CI), so e2e
suites can assert on per-item log lines; production installs keep
`info`.

Motivation: groundwork for #463 Phase 2 observability — per-item GC log
lines can
ship at Debug and be enabled per-node instead of landing at Info.

**Testing**: unit tests pin the untouched default at exactly `info`,
dynamic
raise/lower, case-insensitivity, invalid-value rejection, and the
empty-string
no-op; `go test -race` green across serverboot, atenet, atelet, ateapi.
2026-08-01 16:51:22 -07:00
Krisztian F acf5c12df4 (otel): onboard ateom to the OTLP metrics path (#562)
This PR wires `ateom` into the OTLP push path so it can emit metrics,
and unblock unblocking
https://github.com/agent-substrate/substrate/issues/550

- serverboot: new InitMetricsPushOnly (OTLP periodic reader, no
Prometheus pull surface) for binaries that run no metrics HTTP server
- ateom-gvisor + ateom-microvm call it, so the existing otelgrpc server
handler now feeds `rpc.server.*` to the collector
- atecontroller propagates its `OTEL_EXPORTER_OTLP_ENDPOINT` +
`OTEL_RESOURCE_ATTRIBUTES` into the ateom worker pods
- fixes ateom telemetry not reaching the collector 
- e2e: metrics suite asserts an `ateom` service reaches the collector

cc. @git286 

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-07-29 10:25:37 -04:00
Julian Gutierrez Oschmann 04aee6ccb8 Instrument ate-apiserver so that it can shutdown gracefully. (#523)
Make the server trap SIGTERM and initiate a graceful shutdown process.
This involves reporting the Pod as not ready, and draining all in-flight
gRPC requests so they can complete.

Part of #181
2026-07-24 16:34:47 -07:00
krisztianfekete 992aad740a chore(otel): bump semconv v1.21.0 -> v1.40.0 and pin schema_url 2026-07-08 12:43:20 -07:00
Benjamin Elder 56c4734e0f internal/actorlog: share the actor logger across runtimes
Move the ActorLogger (and SyncedWriter) out of cmd/ateom-gvisor into a shared
internal package so both ateom runtimes forward actor container logs to the pod
log with the same ate.dev/* labels. Add serverboot.InitLoggerWithWriter so a
runtime can route its own slog through the same synchronized writer as the actor
log forwarder (no interleaved lines).
2026-06-25 13:44:25 -07:00
Eitan Yarmush 504cfd21ab Replace actor eth0 move with veth networking 2026-06-10 22:37:09 -07:00
Konippi dffe585cfe feat(serverboot): set service.instance.id on the telemetry resource 2026-06-04 22:40:05 -07:00
Tim Hockin 1ac5ba32a1 Fix boilerplate in go files 2026-05-31 19:45:36 -07:00
Davanum Srinivas a436ea216a refactor: trim cmd/* boilerplate, extract internal/serverboot (#88)
Three related cleanups in cmd/{ateapi,atelet,ateom-gvisor}/main.go:

1. atelet Run/Checkpoint/Restore handlers shared three blocks of
cut-and-paste (fetchRunsc + MkdirAll, errgroup over pause + application
containers calling prepareOCIDirectory, dial-ateom + NewAteomClient, the
ateletpb -> ateompb WorkloadSpec projection). Factored into
fetchRunscAndPrep, prepareOCIBundles, dialAteom, buildAteomWorkloadSpec,
plus a uploadIfExists for the optional pages.img / pages_meta.img upload
in Checkpoint.

2. ateapi main() was a 242-line wall of inlined startup steps;
decomposed into named helpers (loadFlagsFromEnv, logFlagValues,
connectRedis with buildRedisTLSConfig + pingRedisWithRetries,
newKubeClients, buildServerCreds) so the top-level flow reads as
orchestration.

3. ateapi, atelet, and ateom-gvisor each carried their own ~30-line
initTracing + ~35-line initMetrics + slog/metrics-server boilerplate
that differed only in ServiceName, sampler choice, and whether the OTLP
exporter was wired. Lifted to a new internal/serverboot package exposing
InitLogger, InitTracing (with TracingOptions.NoExporter for ateom-gvisor
which loses egress once eth0 moves into the gvisor netns), InitMetrics,
Fatal, ShutdownProvider, and StartMetricsServer.

No behaviour change. Verified with go vet, GOOS=linux go build, and go
test ./cmd/ateapi/internal/...

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR

Signed-off-by: Davanum Srinivas <davanum@gmail.com>
2026-05-27 10:30:38 -04:00