* chore(agents): simplify contributor instructions and workflows Closes #3980 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(contributing): scope verification to affected components Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(contributing): standardize issue branch naming Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
16 KiB
Agent Instructions
This file is the primary instruction surface for agents contributing to OpenShell. It is injected into your context on every interaction — keep that in mind when proposing changes to it.
See CONTRIBUTING.md for build instructions, task reference, project structure, and the full agent skills table.
Project Identity
OpenShell is built agent-first. We design systems and use agents to implement them — this is not vibe coding. The product provides safe, sandboxed runtimes for autonomous AI agents, and the project itself is built using the same agent-driven workflows it enables.
Skills
OpenShell has two skill collections:
skills/contains public, installable skills for using and operating OpenShell. These skills must work outside a source checkout and use installed CLI help plus published documentation as their sources of truth..agents/skills/contains internal contributor and maintainer workflows for developing OpenShell. Your repository-aware harness can discover and load them natively.
Do not rely on this file for a full inventory. The detailed public and contributor skill tables are in CONTRIBUTING.md (for humans).
Architecture Overview
| Path | Components | Purpose |
|---|---|---|
crates/openshell-cli/ |
CLI binary | User-facing command-line interface |
crates/openshell-conformance/ |
CLI conformance library | Reusable driver-agnostic scenarios and command runner |
crates/openshell-conformance-cli/ |
Conformance CLI | Legacy local list and run entrypoint pending follow-up cleanup |
crates/openshell-server/ |
Gateway server | Control-plane API, sandbox lifecycle, auth boundary |
crates/openshell-sandbox/ |
Sandbox runtime | Capability-free workload launcher, process identity, and seccomp-mediated I/O |
crates/openshell-supervisor/ |
Supervisor runtime | Gateway session, policy evaluation, credentials, and upstream networking |
crates/openshell-binary-identity/ |
Binary identity | Shared trusted procfs executable identity resolution for isolation backends |
crates/openshell-isolation-interface/ |
Isolation backend interface | RFC 0012 IsolationBackend trait and types; the supervisor-facing runtime contract |
crates/openshell-sandbox-backend/ |
OpenShell sandbox backend | OpenShellRuntimeBackend and the authenticated OpenShell Sandbox Protocol shared with openshell-sandbox |
crates/openshell-policy/ |
Policy engine | Filesystem, network, and process constraints |
crates/openshell-policy-schema/ |
Authored policy schema | Dependency-light YAML/JSON representation, bounded parsing, and pure authored-language semantics |
crates/openshell-bootstrap/ |
Gateway metadata | Gateway registration metadata, auth token storage, mTLS bundle storage |
crates/openshell-gateway-interceptors/ |
Gateway interceptors | Intercepts and transforms configured gRPC requests at the gateway routing boundary |
crates/openshell-ocsf/ |
OCSF logging | OCSF v1.8.0 event types, builders, shorthand/JSONL formatters, tracing layers |
crates/openshell-otel/ |
OpenTelemetry support | Shared OTLP trace provider, resource, and tracing-layer construction |
crates/openshell-otel-test-support/ |
OpenTelemetry test support | Shared loopback OTLP collector fixture for tracing tests |
crates/openshell-core/ |
Shared core | Common types, configuration, error handling |
crates/openshell-extension-core/ |
Extension core | Shared extension identity, JWT claims, bearer-token rotation, and TLS transport primitives |
crates/openshell-gateway/ |
Gateway binary composition | Links selected first-party compute drivers into the backend-agnostic server registry |
crates/openshell-sdk/ |
Shared client SDK | Async Rust gateway client (gRPC transport, TLS, OIDC refresh, edge tunnel); consumed by CLI, TUI, and @openshell/sdk |
crates/openshell-providers/ |
Provider management | Credential provider backends |
crates/openshell-tui/ |
Terminal UI | Ratatui-based dashboard for monitoring |
crates/openshell-driver-kubernetes-secrets/ |
Kubernetes Secrets credential driver | In-process CredentialDriver backend for OpenShell-managed K8s Secret storage |
crates/openshell-driver-vault/ |
Vault credential driver | In-process CredentialDriver backend for Vault-compatible KV storage |
crates/openshell-driver-db-credstore/ |
Database credential driver | In-process CredentialDriver backend for gateway database credential storage |
crates/openshell-driver-kubernetes/ |
Kubernetes compute driver | In-process ComputeDriver backend for K8s sandbox pods |
crates/openshell-driver-docker/ |
Docker compute driver | In-process ComputeDriver backend for local Docker sandbox containers |
crates/openshell-driver-podman/ |
Podman compute driver | In-process ComputeDriver backend for local Podman sandbox containers |
crates/openshell-driver-vm/ |
VM compute driver | Standalone libkrun-backed ComputeDriver subprocess (embeds its own rootfs + runtime) |
crates/openshell-driver-mxc/ |
Microsoft MXC compute driver | In-process Windows AppContainer and isolation-session compute backend |
crates/openshell-prover/ |
Policy prover | Policy verification and proof generation |
crates/openshell-prover-cli/ |
Policy prover CLI | Standalone local policy boundary checks |
crates/openshell-server-macros/ |
Server macros | Compile-time helpers for gateway RPC authorization |
crates/openshell-supervisor-middleware/ |
Middleware runtime | Generic middleware registry, remote service integration, and chain execution |
crates/openshell-supervisor-middleware-builtins/ |
Built-in middleware | First-party in-process middleware implementations |
crates/openshell-supervisor-network/ |
Network supervisor | Proxying, L7 enforcement, policy evaluation, and provider credential injection |
crates/openshell-supervisor-process/ |
Process supervisor | Process lifecycle, namespace, and bypass monitoring |
crates/openshell-vfio/ |
VFIO support | PCI and GPU passthrough preparation and lifecycle |
python/openshell/ |
Python SDK | Python bindings and CLI packaging |
sdk/typescript/ |
TypeScript SDK | Native Connect client, curated sandbox API, and generated protobuf types |
proto/ |
Protobuf definitions | gRPC service contracts |
deploy/ |
Docker, Helm, K8s | Dockerfiles, Helm chart, manifests |
docs/ |
Published docs | MDX pages, navigation, and content assets |
fern/ |
Docs site config | Fern site config, components, and theme assets |
skills/ |
Public agent skills | Installable workflows for using and operating OpenShell |
.agents/skills/ |
Contributor agent skills | Repository-aware workflows for developing OpenShell |
.agents/agents/ |
Agent personas | Sub-agent definitions (e.g., reviewer) |
Public API Conventions
Follow proto/README.md for public protobuf API design. It is the canonical source for entity-reference naming, workspace selectors, field design, and schema evolution.
Sandbox Logging (OCSF)
When adding or modifying log emissions in openshell-sandbox, determine whether the event should use OCSF structured logging or plain tracing.
When to use OCSF
Use an OCSF builder + ocsf_emit!() for events that represent observable sandbox behavior visible to operators, security teams, or agents monitoring the sandbox:
- Network decisions (allow, deny, bypass detection)
- HTTP/L7 enforcement decisions
- SSH authentication (accepted, denied, nonce replay)
- Process lifecycle (start, exit, timeout, signal failure)
- Security findings (unsafe policy, unavailable controls, replay attacks)
- Configuration changes (policy load/reload, TLS setup, provider attachments, settings)
- Application lifecycle (supervisor start, SSH server ready)
When to use plain tracing
Use info!(), debug!(), warn!() for internal operational plumbing that doesn't represent a security decision or observable state change:
- gRPC connection attempts and retries
- "About to do X" events where the result is logged separately
- Internal SSH channel state (unknown channel, PTY resize)
- Zombie process reaping, denial flush telemetry
- DEBUG/TRACE level diagnostics
Choosing the OCSF event class
| Event type | Builder | When to use |
|---|---|---|
| TCP connections, proxy tunnels, bypass | NetworkActivityBuilder |
L4 network decisions, proxy operational events |
| HTTP requests, L7 enforcement | HttpActivityBuilder |
Per-request method/path decisions |
| SSH sessions | SshActivityBuilder |
Authentication, channel operations |
| Process start/stop | ProcessActivityBuilder |
Entrypoint lifecycle, signal failures |
| Security alerts | DetectionFindingBuilder |
Nonce replay, bypass detection, unsafe policy. Dual-emit with the domain event. |
| Policy/config changes | ConfigStateChangeBuilder |
Policy load, Landlock apply, TLS setup, provider attachments, settings |
| Supervisor lifecycle | AppLifecycleBuilder |
Sandbox start, SSH server ready/failed |
Severity guidelines
| Severity | When |
|---|---|
Informational |
Allowed connections, successful operations, config loaded |
Low |
DNS failures, non-fatal operational warnings, LOG rule failures |
Medium |
Denied connections, policy violations, deprecated config |
High |
Security findings (nonce replay, Landlock unavailable) |
Critical |
Process timeout kills |
Example: adding a new network event
use openshell_ocsf::{
ocsf_emit, NetworkActivityBuilder, ActivityId, ActionId,
DispositionId, Endpoint, Process, SeverityId, StatusId,
};
let event = NetworkActivityBuilder::new(crate::ocsf_ctx())
.activity(ActivityId::Open)
.action(ActionId::Denied)
.disposition(DispositionId::Blocked)
.severity(SeverityId::Medium)
.status(StatusId::Failure)
.dst_endpoint(Endpoint::from_domain(&host, port))
.actor_process(Process::new(&binary, pid))
.firewall_rule(&policy_name, &engine_type)
.message(format!("CONNECT denied {host}:{port}"))
.build();
ocsf_emit!(event);
Key points
crate::ocsf_ctx()returns the process-wideEventContext. It is always available (falls back to defaults in tests).ocsf_emit!()is non-blocking and cannot panic. It stores the event in a thread-local and emits viatracing::info!().- The shorthand layer and JSONL layer extract the event from the thread-local. The shorthand format is derived automatically from the builder fields.
- For security findings, dual-emit: one domain event (e.g.,
SshActivityBuilder) AND oneDetectionFindingBuilderfor the same incident. - Never log secrets, credentials, or query parameters in OCSF messages. The OCSF JSONL file may be shipped to external systems.
- The
messagefield should be a concise, grep-friendly summary. Details go in builder fields (dst_endpoint, firewall_rule, etc.).
Sandbox Infra Changes
- If you change sandbox infrastructure, ensure the relevant sandbox e2e path succeeds.
Network Sockets
- On latency-sensitive TCP streams, disable Nagle's algorithm so small
request/response frames don't stall on delayed ACKs. Use
openshell_core::net::set_tcp_nodelay_best_efforton an accepted or already-connected stream, oropenshell_core::net::connect_tcp_nodelay_best_effortwhen dialing. - This applies to loopback/localhost TCP too — the delayed-ACK stall is a timer behavior, not wire latency.
- You should skip it for unix domain sockets (no Nagle). It's not critical for test-only connections, though using it on any non-UDS TCP stream — tests included — is fine and preferred.
Commits
- Always use Conventional Commits format for commit messages
- Format:
<type>(<scope>): <description>(scope is optional) - Common types:
feat,fix,docs,chore,refactor,test,ci,perf - Sign off on each commit for DCO compliance. Use the
--signoffoption togit committo add theSigned-off-byfooter to ensure the user's configured email address is used. - Never mention Claude or any AI agent in commits (no author attribution, no Co-Authored-By, no references in commit messages)
Go SDK (sdk/go/)
- The Go SDK lives in
sdk/go/with module pathgithub.com/NVIDIA/OpenShell/sdk/go. - Run
mise run go:cifor the full SDK CI pipeline (lint, build, test, proto-check, docs-check). - Proto bindings are generated with
mise run go:proto:genfrom the.protofiles inproto/. - Domain types in
sdk/go/openshell/v1/types/must not import proto packages. - Converters in
sdk/go/openshell/v1/internal/converter/deep-copy slices and maps at boundaries. - Tests use bufconn for in-process gRPC and testify for assertions.
TypeScript SDK (sdk/typescript/)
- Run
mise run sdk:ts:cifor codegen, proto lint, Biome lint, type checking, unit tests, coverage, and build validation. - Proto bindings are generated with
mise run sdk:ts:protofrom the files selected insdk/typescript/buf.gen.yaml. - Generated files under
sdk/typescript/src/gen/are build outputs and must not be committed. - Keep the curated API free of generated wire types; expose full generated messages and RPCs through
@nvidia/openshell-sdk/raw. - The release workflow publishes the package to GitHub Packages. Branch checks exercise the publish path with
npm publish --dry-run.
Python
- Always use
uvfor Python commands (e.g.,uv pip install,uv run,uv venv)
Docker
- Always prefer
misecommands over direct docker builds (e.g.,mise run docker:buildinstead ofdocker build)
Cluster Infrastructure Changes
- If you change gateway deployment infrastructure (e.g., Helm values/templates, gateway image packaging, or deploy logic in
openshell-cli), update thedebug-openshell-clusterskill inskills/debug-openshell-cluster/SKILL.mdto reflect those changes.
Skill Maintenance
When behavior, commands, or development workflows change, review the related agent skills in the same branch. Use the sync-agent-infra skill for the maintenance map and consistency checks.
Documentation
- Put crate-specific implementation details in the relevant crate
README.mdand design proposals inrfc/. - When changes affect user-facing behavior, make the smallest necessary changes to the relevant published docs pages under
docs/and navigation indocs/index.yml. Explain exactly what users need to know; avoid repeating information across pages or exhaustively describing internals that do not affect users. - When changing gateway TOML fields, driver-specific config options, config defaults, or Helm rendering of
gateway.toml, updatedocs/how-it-works/gateways/configuration.mdxin the same branch. fern/contains the Fern site config, components, preview workflow inputs, publish settings, and publishing documentation infern/README.md.- Follow the docs style guide in docs/CONTRIBUTING.mdx: active voice, minimal formatting, no filler introductions,
shellfences for copyable commands, and no duplicate body H1. - Fern PR previews run through
.github/workflows/branch-docs.yml. Release Dev publishesdev, and Release Tag publishes an immutable stable version pluslatest. Both production paths call.github/workflows/sync-docs.ymlonce. - Use the
update-docs-from-commitsskill to scan recent commits and draft doc updates.
Security
- Never commit secrets, API keys, or credentials. If a file looks like it contains secrets (
.env,credentials.json, etc.), do not stage it. - Do not run destructive operations (force push, hard reset, database drops) without explicit human confirmation.
- Scope changes to the issue at hand. Do not make unrelated changes in the same branch.