mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
## What this does ateapi writes a record every time an actor changes state. Until now those records only went to the pod's stdout, and nothing reads stdout. This sends the same records to a collector as OTLP log events. Actor name and uid cannot be metric labels (too many values), and traces are sampled at 1%. So these records are the only way to answer "what state is this actor in, and since when". ## Changes - `serverboot.InitLogging` sets up a LoggerProvider, next to the existing tracer and meter ones. - New `internal/actorevent` package builds the log records. - ateapi emits at the two places that already write the stdout records. - Two event names: `ate.actor.state_changed` and `ate.actor.crashed`. - Both names are registered in `docs/metrics/registry/events.yaml`, so `make verify` checks them. - kind gets a logs pipeline and a count connector. The e2e suite reads the counts back. - Docs updated. `otel-collector.md` said substrate has no LoggerProvider, which is no longer true. ## Opt-in `OTEL_LOGS_EXPORTER` defaults to `none`. Only the kind overlay sets it to `otlp`. The base ConfigMap is untouched, so no deployed environment changes when this merges. ## Notes on the design - **No slog bridge.** Only two call sites emit these records, so emitting twice costs two lines. A bridge would also send every ateapi log over the wire, could not set the event name, and would loop, because SDK export errors are logged through slog. - **Batching processor, not the simple one.** These records sit on the actor resume path. A processor that exports inside the emit call would add a blocking gRPC call there, so a slow collector would become control plane latency. - **Two event names, not one per state.** `ate.actor.state` already says which transition happened. A crash gets its own name because it carries two extra attributes and a higher severity. - **Both copies are kept on purpose.** No collector in this repo reads pod stdout, so nothing is duplicated today. `kubectl logs` keeps working. If a filelog agent is ever added, drop one of the two. The escape hatch is written down in `docs/metrics/substrate.yaml`. ## Dependencies Adds `otel/log`, `otel/sdk/log` and `otlploggrpc`, all pinned at v0.20.0. That is the release that matches the pinned `otel v1.44.0`. v0.21.0 would pull the core modules to v1.45.0, which this change does not need. The logs API has a v1.47.0 release candidate upstream, so it is on its way to stable. ## Testing - Unit tests for the exporter resolver, the record builder, and both ateapi emit sites. - The record builder test checks the attribute set matches what the event name declares, in both directions. - The ateapi tests check the OTLP record carries the same attributes as the stdout record. - Ran end to end on kind. Records arrive with the right event name, severity, attributes, and with trace context on the record's own fields rather than as attributes. - Checked the off state too. With `OTEL_LOGS_EXPORTER` removed, the collector receives no log records and stdout is unchanged. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
50 lines
1.9 KiB
Go
50 lines
1.9 KiB
Go
// Copyright The OpenTelemetry Authors
|
|
// SPDX-License-Identifier: Apache-2.0
|
|
|
|
/*
|
|
Package global provides access to a global implementation of the OpenTelemetry
|
|
Logs API.
|
|
|
|
This package is experimental. It will be deprecated and removed when the [log]
|
|
package becomes stable. Its functionality will be migrated to
|
|
go.opentelemetry.io/otel.
|
|
*/
|
|
package global // import "go.opentelemetry.io/otel/log/global"
|
|
|
|
import (
|
|
"go.opentelemetry.io/otel/log"
|
|
"go.opentelemetry.io/otel/log/internal/global"
|
|
)
|
|
|
|
// Logger returns a [log.Logger] configured with the provided name and options
|
|
// from the globally configured [log.LoggerProvider].
|
|
//
|
|
// If this is called before a global LoggerProvider is configured, the returned
|
|
// Logger will be a No-Op implementation of a Logger. When a global
|
|
// LoggerProvider is registered for the first time, the returned Logger is
|
|
// updated in-place to report to this new LoggerProvider. There is no need to
|
|
// call this function again for an updated instance.
|
|
//
|
|
// This is a convenience function. It is equivalent to:
|
|
//
|
|
// GetLoggerProvider().Logger(name, options...)
|
|
func Logger(name string, options ...log.LoggerOption) log.Logger {
|
|
return GetLoggerProvider().Logger(name, options...)
|
|
}
|
|
|
|
// GetLoggerProvider returns the globally configured [log.LoggerProvider].
|
|
//
|
|
// If a global LoggerProvider has not been configured with [SetLoggerProvider],
|
|
// the returned Logger will be a No-Op implementation of a LoggerProvider. When
|
|
// a global LoggerProvider is registered for the first time, the returned
|
|
// LoggerProvider and all of its created Loggers are updated in-place. There is
|
|
// no need to call this function again for an updated instance.
|
|
func GetLoggerProvider() log.LoggerProvider {
|
|
return global.GetLoggerProvider()
|
|
}
|
|
|
|
// SetLoggerProvider configures provider as the global [log.LoggerProvider].
|
|
func SetLoggerProvider(provider log.LoggerProvider) {
|
|
global.SetLoggerProvider(provider)
|
|
}
|