Since #2726 the canonical main process's stdout and stderr are captured
in pipes that feed only the in-memory replay buffer used by sandbox
connect. Agent output therefore never reaches the container's own stdout
and stderr, so it is missing from kubectl logs, docker logs, and podman
logs and from anything that collects container logs. Before #2726 the
entrypoint inherited the container's descriptors and its output appeared
there.
Copy the main process's output to the launcher's stdout and stderr in
addition to the replay buffer, restoring the earlier behavior:
- Output is copied byte for byte to the matching stream from a
forwarder thread per stream, after it is published to the replay
buffer. When the container runtime falls behind on a stream, that
stream's reader waits instead of dropping output, so backpressure
reaches the agent as it did with inherited descriptors, while the
other stream and attachments keep receiving output.
- Before the main process's exit is published, the output readers
finish and queued output is drained to the container log, so an
agent's final lines are not lost at shutdown. A 30 second deadline
covers both; when it expires, readers waiting on the container log
are released and drain the pipes into the replay buffer only, so a
stalled container log cannot block exit reporting.
- PTY-mode processes are not copied. The terminal stream carries escape
sequences and echoed input, and terminal commands never reached the
container log before #2726.
- Exec, SSH, and SFTP sessions are not copied.
Launcher log lines keep their existing format and remain in the
container's stderr. They are written as whole lines, and a newline is
inserted first when the agent left stderr mid-line, so launcher and
agent lines do not merge.
The Docker and VM drivers appended the tail of the workload's output to
failure messages: Docker the workload container's log, and the VM driver
the guest console, which carries the launcher's stdout and stderr. Those
messages land in the sandbox's Ready condition and in platform events
that the gateway republishes to the sandbox event stream. With agent
output in that log, those messages would carry arbitrary agent output,
including anything sensitive the agent prints, into gateway status and
events. The supervisor starts its health endpoint only after the agent
starts, so every Docker failure path could include agent output, and the
VM driver reports one whenever the VM or host supervisor exits. Forward
only the supervisor's log tail, matching the Podman driver, which reads
the workload log solely to match fixed launcher markers and never
forwards raw workload output. The workload's output remains available
through docker logs and the VM's rootfs-console.log.
Document where main process output appears in the logging docs and the
cluster debugging skill.
Closes#3928
Signed-off-by: Kris Hicks <khicks@nvidia.com>
679b19067 added global.image.registry (ghcr.io/nvidia) as the fallback for
empty per-image registries and split e2e image references into registry and
repository. Locally built images such as openshell/gateway:<tag> have no
registry host, so the chart rewrote them to ghcr.io/nvidia/openshell/* and
the k3d cluster could not pull them. Clear global.image.registry in the
Kubernetes e2e wrapper, which sets every image's registry explicitly.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
log_response has always logged every gateway response at INFO, including
health probes and the GetSandboxConfig and provider-readiness polls each
supervisor makes. #3915 demoted the request spans for those polled paths
to DEBUG, which stripped the request{method path} prefix from the log
line at INFO but left the line itself, so the gateway log fills with
bare 'response status=200' lines several times per second.
Follow the span's level: polled requests log their response at DEBUG,
or WARN on a 5xx so probe and poll failures stay visible.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Store operation spans and request spans for supervisor-polled RPCs
(GetSandboxConfig, ReportProviderReadiness) use DEBUG level, so the
default INFO filter no longer exports them. The provider credential
refresh worker opens its span only when a state has work.
Refs #2698
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Previously, Helm installations could not enable the gateway OCSF JSONL
destination through chart values because generated `gateway.toml` omitted the
`openshell.gateway.ocsf_log` table.
Now, setting `server.ocsfLog.enabled` renders the path, optional schema
version, rotation, retention, and queue limits into gateway configuration.
Output is disabled by default. The default path, `/tmp/gateway-ocsf.jsonl`,
is writable in the gateway container with either the StatefulSet or
Deployment workload, so enabling output does not require persistent storage.
Invalid schema versions, rotation values, non-positive limits, or an empty
path while enabled fail chart rendering.
Additionally, `server.extraVolumes` and `server.extraVolumeMounts` add
operator-supplied volumes to the gateway pod, so operators who want records
to survive restarts can place the OCSF path on persistent storage without
replacing chart-generated configuration.
The gateway pod's default termination grace period rises from 5 to 30
seconds. Gateway shutdown can spend up to 10 seconds on supervisor session
cleanup before allowing 5 seconds to drain queued OCSF records, so the
5-second default risked a SIGKILL before the final records were written.
The grace period is only an upper bound: the gateway exits as soon as its
shutdown completes.
Refs #2762
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Remove architecture/. It was a constant source of merge conflicts, became
an effectively append-only log of the project, and was of dubious value.
Design records live in rfc/, crate details in crate READMEs, and user
documentation in docs/.
Move the git-ignored plans directory from architecture/plans to plans/,
keeping the old .gitignore entry. Remove the arch-doc-writer agents and
update AGENTS.md, CONTRIBUTING.md, skills, the feature request template,
and links in proto/, rfc/, and examples/ that pointed into architecture/.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
The test reserved a loopback port, released it, and rebound it inside the
workload. A concurrent nextest process could claim the port in between,
failing the bind and surfacing only as a RecvError on the ready channel.
Bind port 0 in the workload and send the assigned address instead, and
report the workload error when the listener never becomes ready.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(server): write gateway OCSF events to JSONL
Previously, gateway security activity was available only in diagnostic output,
and events not associated with a sandbox, such as TLS certificate reloads, had
no independent structured record.
Now, configuring `openshell.gateway.ocsf_log` writes every gateway-produced
OCSF record to a bounded JSONL destination independently of `RUST_LOG`. The
destination supports daily or disabled rotation, retention limits, queue bounds,
and optional schema downgrade to OCSF 1.1 or 1.3.
Additionally, existing gateway emitters (TLS reloads, service routing, and
policy approval and auto-approval audits) emit structured events, so they reach
the JSONL destination, console shorthand, and the affected sandbox's log
stream. Records identify the gateway by its configured name in `device.uid` and
`device.name`, shared across replicas, with `device.hostname` identifying the
replica and `device.os` the gateway's operating system. Metrics and warnings
expose known best-effort losses.
Refs #2762
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(mxc): attribute ETW events to gateway
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
The snap installer tests added in #3656 checked the gateway config mode
with GNU stat -c, which BSD stat on macOS rejects, so mise run ci failed
locally on macOS. Check the mode with find -perm instead, which matches
the exact mode on both GNU and BSD systems.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(driver-kubernetes-secrets)!: store provider credentials in one namespace
The Kubernetes Secrets credential driver now stores every credential in its
configured namespace in all workspace modes and rejects handles that reference
any other namespace before contacting the Kubernetes API. The gateway reaches
credential Secrets through the Role in that namespace; this allows removing the
Secret rules from the ClusterRole.
- Remove the workspace_mode, gateway_id, and allow_reference_namespace driver
settings and stop rendering them from Helm. Configurations that set them fail
at startup. Existing credential state is not migrated.
- Add server.credentialDrivers.kubernetesSecrets.createNamespace to provision a
dedicated credential namespace. The namespace is kept on uninstall, adopted
by a reinstall of the same release, and left untouched when something else owns
it.
- Update the gateway config reference, Kubernetes setup docs, 0.1.0
upgrade guide, compute-runtime architecture, and cluster debugging skill.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(helm): reduce gateway Secret permissions
Remove the gateway's Secret list permission in every workspace mode and
grant source Secret reads through a Role in the sandbox namespace.
Bootstrap Secret cleanup deletes Secrets by exact name instead of listing
them.
- Grant get on the copied client TLS and image-pull Secrets through a Role in
the sandbox namespace. The ClusterRole keeps get and patch on those names
for the ownership check and server-side apply into workspace namespaces.
- Delete sandbox and supervisor bootstrap Secrets by exact name, derived from
the runtime generation recorded on the Sandbox and, on restart, the target
generation, tolerating 404. The generation annotation is cleared only after
cleanup succeeds, and each bootstrap Secret has a Pod owner reference, so
garbage collection removes any generation the driver does not name.
- Drop Secret list from the ClusterRole and the shared-mode sandbox Role.
- Extend the managed e2e RBAC checks to Secret list.
- Update the Kubernetes setup and sandbox runtime docs, compute-runtime
architecture, and cluster debugging skill.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(driver-kubernetes): stage workspace Secrets per runtime generation
Every Secret the Kubernetes driver writes into a workspace namespace is
now scoped to one sandbox runtime generation, immutable, and created
with create only, so the gateway never reads, patches, or adopts an
existing Secret there. This removes the gateway's cluster-wide get and
patch on the copied client TLS and image-pull Secret names.
- In managed mode, create an immutable copy of each configured
image-pull Secret per generation, named os-pull-<id>-<generation>-<n>
and owned by the generation's workload and supervisor Pods. Pods and
the restarted Sandbox template reference those names, and generation
cleanup deletes them by name. A Secret already holding a generation
name fails the create.
- Outside shared mode, stage the gateway client TLS material into the
supervisor bootstrap Secret instead of copying the client TLS Secret
into the workspace namespace.
- Remove the fixed-name TLS and image-pull copies, the target ownership
read, and the ClusterRole get and patch rule on the copied names.
Source reads stay in the sandbox-namespace Role.
- Update the managed e2e to expect generation image-pull Secrets and
client TLS material in the supervisor bootstrap Secret, and to check
that the gateway cannot read the copied names in workspace namespaces.
- Update the gateway config and compute driver references, Kubernetes
setup and sandbox runtime docs, compute-runtime architecture, driver
README, and cluster debugging skill.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(helm)!: grant operator-mode Secret permissions through the workspace chart
The operator-mode gateway ClusterRole grants no Secret permissions.
The openshell-workspace chart Role, installed in each operator-managed
namespace, grants the gateway create and delete on Secrets for sandbox
runtime generations. Operator-managed namespaces require the workspace
chart.
- Fail the chart tests on any ClusterRole rule that includes Secrets in
operator and shared modes.
- Install the workspace chart when the operator e2e provisions a
namespace, and assert that the gateway has no Secret permissions in a
namespace without it.
- Update the Kubernetes setup docs, 0.1.0 upgrade guide, compute-runtime
architecture, and cluster debugging skill.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(server): retry runtime identity persistence on sandbox create
With multiple gateway replicas, another replica can update a new sandbox
record between the create path's read and its compare-and-swap write of the
runtime identity. The create path made one attempt, so the conflict failed
the request and deleted the backend sandbox. Persist the identity through
the same retrying helper that start and startup recovery use, which rereads
the record and retries while the sandbox stays in the same generation and a
Provisioning or Ready phase.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* ci(e2e): run Kubernetes HA tests one at a time
The two HA tests scale and roll the shared gateway Deployment. Their
in-process lock does not apply under nextest, which runs each test in its
own process, so one test could delete a gateway pod while the other was
executing through it. Put both tests in a nextest test group limited to one
thread in the e2e-kubernetes profile. The override matches test names
because a binary() filter fails in workspaces that lack the HA test binary.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
PRs labeled test:e2e-kubernetes now run the Kubernetes HA suite and the
Kubernetes credential-driver suite (Kubernetes Secrets and Vault). Both stay
optional and off in merge groups.
- Read the test:e2e-kubernetes label for the HA and credential-driver
lanes in Branch E2E Checks.
- Point gator at test:e2e for Helm and Kubernetes coverage and at
test:e2e-kubernetes for gateway high availability and credential
driver storage.
Refs #3481
Signed-off-by: Kris Hicks <khicks@nvidia.com>
The credential driver e2e has not passed end to end, and the disabled
kubernetes-credential-drivers CI lane hid three problems.
The test broke when the JSON format of provider list changed.
Continuation-token pagination (#3249) changed
provider list --output json from a bare array of providers to an object
with next_page_token and a providers array. The test still parsed the
output as an array, so it failed before checking either storage
backend. Read the providers array from the new object instead.
Its sandbox name was about 58 characters, but sandbox names are
DNS-routable and limited to 19, so sandbox creation was rejected. Build
a short unique name instead.
The sandbox guard deletes its sandbox from a detached thread on drop, so
the test deleted the provider while the sandbox still existed. The
gateway rejects deleting a provider that is attached to a sandbox, the
test ignored that error, and the credential Secret remained. Delete the
sandbox explicitly before returning from the sandbox check.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
The RFC-0012 architecture migration preserved sandbox OCSF JSON enablement but
introduced a regression: ocsf_schema_version was no longer honored.
This fixes the regression so that schema downgrades occur again.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Previously, Network Activity could be constructed without a source or
destination endpoint, allowing connection, accept, relay, and configuration
events to violate the OCSF 1.8 endpoint constraint.
Now, NetworkActivityBuilder requires a source or destination endpoint at
compile time. Connection failures identify the workload peer or genuine
transparent destination, listener failures identify the listening endpoint,
and mediation-lane failures use Application Lifecycle rather than fabricated
network endpoints. Malformed forward requests use HTTP Activity with a
method-only request, generated 400 response, and workload peer.
Additionally, Unix relay-channel events use Base Event, policy-validation
warnings use Config State Change, and the unused bypass monitor is removed
because the current isolation architecture no longer uses it.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
As part of the RFC-0012 changes, the Podman runtime moved trusted networking
into a paired supervisor running as the resolved non-root identity, while the
capability-free sandbox retained ownership of seccomp-mediated I/O. The
supervisor still generated interception TLS material under the root-owned
/etc/openshell-tls directory, and the sandbox treated the documented ENOENT
notification race as a fatal listener failure.
This meant TLS interception could fail with a permission error, and an exiting
target process could stop the network broker while the kernel was preparing its
notification.
Now, we store generated supervisor TLS material in its writable /run tmpfs and
retry ENOENT notification races while preserving fatal handling for other
listener errors.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
The deterministic fixture still modeled Alpine after the supervisor image
switched to Debian, causing the wrapper to exit before producing parity
evidence. Update its image expectations and launch evidence to match the
current harness contract.
Force the intentional replacement of the read-only staged artifact so BSD mv
does not prompt when the deterministic parity test runs from a terminal.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
setuptools-scm can select jj's Git-bridge dev tag when uv resolves the local
package. Pin a development-only version for every Python task that performs
that resolution.
The 0.0.0 value is valid local editable-package metadata for development tasks;
it is neither a release version nor a Git tag. Wheel builds remain unpinned and
derive their published version from the actual release tag.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Previously, metadata events omitted both HTTP request and response objects,
early proxy rejections used HTTP Activity without request context, and
unsupported-scheme events did not expose enough safe HTTP context to satisfy
the OCSF 1.8 schema.
Now, metadata events include a method-only request and their actual HTTP
response codes without recording the metadata URL. Unsupported-scheme events
also include a method-only request plus the generated 400 response. Authority
mismatches and credential-resolution denials use HTTP Activity with their
generated 403 or 500 responses, and HTTP activity IDs are derived from the
request method.
Additionally, HttpActivityBuilder now enforces the OCSF request-or-response
constraint at compile time.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites.
- Remove test-only shell interception and capture generated gateway config
directly.
- Allow parity tests to use supplied supervisor binaries without resolving a
Linux target.
- Normalize temporary-directory paths and use portable RPM config installation.
- Set a valid setuptools-scm version for Python protobuf generation in Jujutsu
checkouts.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Previously, every event from a sandbox reused the sandbox ID as its event
ID. Consumers deduplicating security records could mistake separate events
for the same record, and missing device types or empty image objects could
prevent schema validation.
Give each event its own ID, retain the sandbox association separately, and
classify the environment as Other/Sandbox while keeping the OS separate.
Omit unknown container details instead of emitting empty objects.
Security tooling can now distinguish events from the same sandbox and read
their identity consistently after serialization.
Refs #1055
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Centralize compute-driver RPC descriptors, stream instrumentation, provider
routing, and standalone installation in openshell-otel. Use typed RPC
constants so gateway and in-process driver paths cannot panic on unknown
operation strings or repeat runtime method parsing.
Emit semantic-convention rpc.service and rpc.method attributes, preserve
trace context and resource identity across deployment modes, and route both
RPC boundary and backend crate spans to each selected driver provider. Leave
consumer-dropped watch spans unset while recording observed terminal status,
and avoid reboxing untraced external-driver streams.
Derive each driver tracing identity from Cargo package and crate metadata and
attach its descriptor to the compute-driver registration, keeping provider
selection and target routing tied to the registered implementation. Share
tracing setup and round-trip test support across Docker, Podman, Kubernetes,
and VM, and update the gateway tracing documentation.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Characterization tests for behavior that later work changes. No behavior
change.
- Pin the `openshell logs` output format, so a change to how that text is
produced shows up as a diff rather than passing silently.
- Cover the supervisor log push layer.
- Assert OcsfEvent survives a JSON round trip. Serialization is written
by hand and deserialization dispatches on class_uid into independent
per-variant paths, so a field can serialize correctly and still be
dropped or rejected on the way back in.
Refs #1055
Signed-off-by: Kris Hicks <khicks@nvidia.com>
OCSF_VERSION has been 1.8.0; two descriptions still said 1.7.0.
The default_gateway_id doc comment described a hostname fallback the
implementation never had. A per-replica default would be wrong here:
gateway_id is embedded in the JWT iss/aud, so it has to be stable across
replicas.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
This changes the DNS resolver test fixture to retry until TCP and UDP can share
an ephemeral port.
It fixes an issue seen in CI where a port that was successfully assigned for
one protocol failed for the other.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Ctrl-C closes gateway listeners while supervisor sessions are still current,
causing benign HTTP/2 broken pipes to be logged as warnings. Track gateway
shutdown so recognized transport closes are treated as expected.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(helm): refresh kubeconfig for existing k3d clusters
Docker can recreate the k3d load balancer on a new API port. Start existing
clusters and prefer fresh k3d entries so create does not retain a stale
endpoint.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* fix(dev): conditionally enable local OTLP export
Probe port 4317 before adding OTLP configuration for the VM, Docker,
and Podman gateway tasks. Document the startup behavior and troubleshooting
for local collector availability.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(gateway): add installation name configuration
Add a first-class operator-assigned gateway name with TOML, CLI, environment,
and Helm configuration surfaces. Local gateways default to openshell, while
Helm defaults to the chart fullname; operators sharing a collector across
namespaces or clusters can set a globally distinct name.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(gateway): identify gateways in exported traces
Attach the configured gateway installation name and compute driver to the
gateway OpenTelemetry resource so operators can filter traces from multiple
installations that share a collector. Forward the gateway name and OTLP
endpoint to managed external drivers so their distinct service resources carry
the same installation identity.
Keep service.name stable per process type, omit blank resource values, and
leave per-span operation names and request attributes unchanged.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes.
Export standalone driver spans through OTLP/gRPC as the distinct
openshell-driver-kubernetes service, preserve gateway trace context, record
lifecycle operations and gRPC failures, and flush spans on shutdown.
Kubernetes currently runs in-process when selected as a built-in gateway
driver. Use the temporary server-boundary shim shared with Podman and Docker
so traces retain the shape they will have when Kubernetes moves to a
separate process. Move the common ComputeDriver RPC tracing layer into
openshell-otel to keep all drivers aligned.
Propagate the active W3C context through the controller-reserved Sandbox
annotation and enable Agent Sandbox OTLP export in the local k3s workflow.
This connects asynchronous controller reconciliation spans to the originating
OpenShell create trace.
Expose gateway OTLP configuration through Helm and add an Aspire collector
to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace
UI forwarding for Kubernetes and local container gateway development.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Agents will often give a ton of feedback, but it can be hard to quantify the
impact of the issue the feedback is about.
This change updates the PR review feedback template to be more human-readable,
rooting any concerns in user-visible behavior when appropriate, and comparing
the new behavior to old behavior so that PR authors can make a determination of
whether they want to accept or reject the feedback.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Make the local k3s gateway workflow match the Docker and Podman flows by
registering and selecting successful plaintext Skaffold deployments with the
OpenShell CLI. Derive the registration name from the worktree-specific k3d
cluster name so parallel worktrees retain independent gateway metadata.
Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on
localhost:8090, and update the development and debugging guidance to use the
active registered gateway instead of one-off endpoint flags.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Mirror the VM and Podman driver tracing setup for Docker. Export Docker
driver spans through OTLP/gRPC as the distinct openshell-driver-docker
service, preserve gateway trace context, record lifecycle and asynchronous
provisioning operations, and report gRPC failures.
Docker currently runs in-process when selected as a built-in gateway driver.
Add the same temporary server-boundary shim used by Podman so traces retain
the shape they will have when Docker moves to a separate process. Generalize
the gateway provider selection for both in-process drivers and share the OTLP
collector fixture across Docker, Podman, and VM tracing tests.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Previously, Podman could be selected through automatic driver detection or with
`mise run gateway -- --driver podman`, but it did not have a dedicated task
like the Docker and VM drivers.
This adds a gateway:podman task and moves the Podman-specific setup into its
own script. The generic gateway task now delegates Podman launches to that
script.
Additionally:
Unlike Docker, which rebuilds and bind-mounts the supervisor binary, Podman
uses a dev-tagged supervisor image that can become stale. The default Podman
supervisor image is therefore rebuilt on each launch.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Mirror the VM driver tracing setup for Podman. Export standalone driver
spans through OTLP/gRPC as the distinct openshell-driver-podman service,
propagate W3C context across ComputeDriver RPCs, record bounded RPC names
and failures, and flush buffered spans during graceful shutdown.
Podman still runs in-process when selected as a built-in gateway driver.
Add a temporary tracing shim that partitions gateway and Podman spans by
target into separate tracer providers while preserving their shared trace
and parentage. The shim also emits the same ComputeDriver server boundary
that the tonic layer emits out of process, keeping the observable trace
shape stable when Podman is eventually extracted.
Trace container create preparation, image and storage setup, lifecycle
operations, and cleanup. Document the service boundary and cover it with
isolated and repeated tracing tests.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
This is convenient when you want to run a local gateway pointed at a remote
compute driver so that the supervisor can reach across the network to the
gateway which is listening on 0.0.0.0.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Use the same compact persona, workflow, impact, reproduction, and
environment prompts for bug reports and feature requests. Keep logs
optional and specific to bug reports.
Remove filing-time agent diagnostics so maintainers can evaluate user needs
apart from investigation output, which becomes stale over time. Require
contributors to investigate current behavior after humans accept the work, and
treat state:accepted or roadmap placement as that signal.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
On macOS, bind the standalone Podman gateway to IPv6 loopback while
registering localhost as the TLS endpoint. This keeps IPv4 loopback
available for the callback-only listener, matching the e2e fix in
commit 4cb77a9.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Update the docs for `sandbox create` and `exec` to dissuade use of `--env` for
secrets, and enhance the docs for `--provider` to explain what it's for.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Keep the completed instrumented handler future in an explicit pinned box and
drop it after the await. This releases the handler-side request span clone
before the disconnect test checks producer ownership, avoiding compiler- and
platform-dependent retention of an unfinished span.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Continue distributed traces across the gateway-to-driver process boundary and
export VM driver spans to the same OTLP/gRPC collector. The driver reports as
the distinct openshell-driver-vm service.
Updated the gateway architecture and configuration reference with a generic
external-driver forwarding contract.
Instrumented:
- Every RemoteComputeDriver RPC injects the active W3C trace context into
tonic metadata. Managed VM readiness and runtime initialization give startup
capability probes stable parent operations rather than isolated root spans.
- A tonic service layer creates fixed, low-cardinality server spans for every
ComputeDriver RPC. New handlers inherit tracing automatically; failures record
OpenTelemetry error status and the gRPC status code.
- Background provisioning remains attached to CreateSandbox after the RPC
returns without extending the RPC span lifetime.
- Provisioning records image preparation, bootstrap image resolution, overlay
preparation, lifecycle configuration, pre-launch hooks, guest preparation,
and launcher spawn as child spans.
- VM startup reconciliation roots one trace for the persisted-sandbox scan,
with per-sandbox restore and provision operations beneath it. The root remains
open until all spawned restore tasks finish.
- Delete cleanup records its own child operation.
Design notes:
- The gateway forwards its configured OTLP endpoint to managed external
drivers. SDK `OTEL_*` variables continue to own sampling, batching, limits,
headers, and transport tuning.
- The VM driver has its own tracer provider and service resource so trace
backends preserve the service boundary.
- RPC operation names come from an explicit method mapping, keeping cardinality
bounded without parsing the protobuf descriptor set at runtime.
- Propagation uses a remote SpanContext for spawned provisioning. This keeps
one trace while allowing the CreateSandbox server span to finish when the
RPC response is sent.
- Startup restoration is independent of gateway requests. It begins at the VM
driver reconciliation span rather than attaching to an unrelated RPC.
- Existing tracing events remain on the logging path. The OpenTelemetry layer
exports spans only and excludes the SDK exporter callsites to avoid recursive
traces.
- Export configuration failures do not prevent the driver from serving, and
buffered spans are drained during graceful shutdown.
- Trace fields identify drivers, sandboxes, images, lifecycle phases, and gRPC
outcomes without recording credentials, sandbox tokens, or request query
parameters.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Run the constructor gateway-discovery error test only on Linux, matching the
production branch that inspects the Podman bridge gateway. macOS uses its
Podman machine callback path and correctly skips that inspection.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Give each span-assertion test its own thread-scoped in-memory exporter so
parallel tests cannot contaminate or reset captured spans. Keep a bare global
tracing registry only to preserve callsite interest, and serialize scoped
subscriber changes because tracing caches that interest process-wide.
Remove the test-only OTLP collector, polling delivery barrier,
transport-specific test, and direct opentelemetry-proto dependency. Seed the
expected persistence conflict before tracing begins so its assertion window
contains only the operation under test.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Extract common OpenTelemetry provider construction into openshell-otel so
OpenShell services share one OTLP/gRPC export implementation.
The shared crate owns:
- typed exporter setup errors and non-fatal provider enablement;
- endpoint trimming and URI validation before lazy exporter connection;
- fixed and environment-or-default service-name policies;
- service version and caller-supplied resource attributes;
- batch tracer-provider construction; and
- span-only tracing layers that exclude OpenTelemetry exporter callsites.
Migrate the gateway to the shared provider while retaining its configurable
service name, error marking, and tracing test collector. Add the shared crate
to the architecture inventory and document the tracing boundary.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Add a published Issue Triage and Lifecycle page and align AGENTS.md,
CONTRIBUTING.md, README.md, the PR template, and the issue-handling
skills on the state:*/agent:* label model.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Add an opt-in OTLP/gRPC trace exporter to the gateway. Export is enabled
by the presence of an `[openshell.gateway.otlp]` table with an endpoint;
there is no separate toggle.
Instrumented:
- Inbound request server spans, named for the RPC (`$service/$method`) or
`{method} {path}` for plain HTTP. They continue valid W3C `traceparent`
context when present and start a new trace otherwise. gRPC spans also
carry `rpc.system`, `rpc.service`, `rpc.method`, and trailer-derived
`rpc.grpc.status_code`.
- Compute driver calls (create, delete, list, get, validate, watch) as
client spans anchored on the `ComputeDriver` contract.
- Store reads and writes as children of the current request or loop span.
- Work with no inbound request: compute driver initialization, the sandbox
reconcile sweep, provider credential refresh tick, and driver watch events.
Each roots one operation trace so its child work does not arrive as anonymous
single-span traces.
This is deliberately not exhaustive. Auth, policy evaluation, and
middleware remain uninstrumented, as do store lifecycle calls (`ping`,
`close`) that a readiness poll would turn into a span per tick. The aim is
a useful trace tree at a reviewable size; coverage can grow against real
traces.
Design notes:
- The TOML table owns whether and where to export. The SDK `OTEL_*`
variables own how; sampling, batching, and limits are not mirrored into
gateway config.
- The OpenTelemetry layer exports spans only. Existing `tracing` events
remain on the stdout and sandbox-log paths and are not copied into trace
payloads.
- Telemetry never blocks the gateway. A malformed endpoint logs an error
and disables export rather than failing startup, and buffered spans are
drained during graceful shutdown.
- Failed spans carry error status without a separate `error.type` attribute.
Request spans use HTTP status and gRPC response trailers; driver spans use
the returned gRPC status; autonomous loop spans record failed results
explicitly. Store spans exempt `UniqueViolation` and `Conflict`, because
those errors report expected contention such as a held lease or an
optimistic-concurrency retry.
- The compute driver is reachable only through `TracedDriver::call`, so a
call cannot skip its span. This is the client half of a client/server pair
and the single place to inject context if drivers move out of process.
- Tests share one process-wide subscriber and in-memory exporter because
`tracing` caches callsite interest globally.
Inbound W3C trace context is propagated into gateway request spans. Context
is not yet injected into outbound driver calls, so a future out-of-process
driver would still need propagation at the `TracedDriver` seam.
The Helm chart is intentionally unchanged, so OTLP export cannot yet be
enabled on a chart-deployed gateway.
Refs #2507
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Previously, Docker was auto-detected when the CLI was installed or a candidate
Unix socket existed. Neither check verified that the Docker API was responsive.
A similar check was done when auto-detecting Podman in the past, but was
replaced in 1f07bf04 with a probe of candidate Podman sockets instead.
This change applies the functional API probing approach introduced for Podman
in 1f07bf04 to Docker. It also makes Docker driver initialization use the same
socket-selection mechanism as Docker auto-detection instead of Bollard’s local
defaults. This means the previously auto-detectable Docker socket paths
$HOME/.docker/run/docker.sock and $XDG_RUNTIME_DIR/docker.sock will actually be
usable.
When no working compute driver can be auto-detected, the gateway exits early
with a message saying as much:
> configuration error: no compute driver configured and auto-detection found no
> suitable driver; set --drivers or OPENSHELL_DRIVERS to kubernetes, podman,
> docker, or vm
This makes for a better user experience when installing OpenShell without an
available supported compute driver.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Dependabot updated the Helm setup action in release-canary but missed the same
reference in the release composite action.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Run cargo fmt for both the root workspace and the standalone e2e Rust workspace
from the rust format and format-check tasks. Apply rustfmt to the e2e sources
so the expanded formatting check passes.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Apply the Linux cfg to the Path import so native macOS lint runs do not report
it as unused when the only call site is compiled out.
This fixes `mise run rust:lint` on macOS.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Add policy schema, proto, provider profile, OPA, and L7 proxy support for
`protocol: json-rpc` and `protocol: mcp`. Generic JSON-RPC endpoints match
exact method names only, with `method: "*"` as the all-method sentinel;
wildcard/glob methods and params matchers are rejected.
Parse JSON-RPC request bodies and batches in the forward proxy, deny
response-shaped client frames, limit receive-stream GET allowance to MCP
endpoints, and redact params in decision logs. Preserve L7 rule params on the
proto load path so MCP `tools/call` tool filters behave like YAML-loaded
policies.
Add MCP conformance coverage, JSON-RPC L7 e2e coverage, and docs for the new
protocols and current matcher limitations.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Co-authored-by: ddurst <267424412+ddurst-nvidia@users.noreply.github.com>
This fixes an issue where you may run e.g. `mise run e2e:python`, then after
Python is upgraded in mise.toml, subsequent runs of `e2e:python` fail because
the Python version is out of sync.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Add a create-rfc skill that directs agents through the OpenShell RFC process
and template. Expand the RFC template with clearer section guidance and
suggested lengths.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
The Ubuntu Snap canary downloads its artifact from a different workflow run
(the triggering Release Dev run) via run-id. Cross-run downloads require
authentication, so pass github.token to actions/download-artifact.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Auto-detection previously treated Podman as available only when the podman CLI
was visible on PATH. However, package manager services can run with a
restricted PATH, which lets Docker be selected even when a Podman API socket is
reachable. Additionally, podman may symlink /var/run/docker.sock to podman's
machine unix socket, which would be incorrectly detected as Docker. Worse
still: the podman machine may not even be running.
This replaces the Podman binary check with a functional HTTP probe against the
standard Podman socket paths. The probe requires /_ping to answer with a
Libpod-Api-Version header before treating the socket as Podman, which lets the
gateway select the embedded Podman driver only when the API is usable.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
Install the Snap built by the triggering Release Dev workflow by setting
merge-multiple: true on the artifact download. actions/download-artifact
otherwise extracts each artifact into its own subdirectory, leaving the
package at release/snap-linux-amd64/*.snap, so the install glob
./release/*.snap matched nothing. Merging flattens the artifact's contents
directly into release/ where the dangerous local snap install expects it.
Harden the Snap canary setup by enabling snapd.socket, waiting for snap
seeding (snap wait system seed.loaded), and running every step with strict
shell options (set -euo pipefail) so failures surface immediately.
Register the snapped gateway with the CLI as the documented local plaintext
snap-docker gateway, and print version and snap services, before running
openshell status so the canary verifies a configured and reachable gateway
instead of only the install.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
The RPM canary needs to exercise the install.sh user-service path, but a GitHub
Actions job container does not boot with systemd as PID 1. The Fedora RPM
canary needs to exercise the install.sh user-service path, but a GitHub Actions
job container does not boot with systemd as PID 1. This means the Fedora RPM
canary was incomplete as compared to the others.
With this change, we run Fedora as a nested privileged systemd container
instead, wait for systemd to become reachable, then start the root user manager
so systemctl --user works for the RPM gateway unit, achieving parity with the
other canary tests.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
release-helm and tag-ghcr-release now depend on the release job.
This is to prevent a GHCR image or helm chart from being published when some
other aspect of the release fails.
Signed-off-by: Kris Hicks <khicks@nvidia.com>