mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
* refactor(sandbox): extract run_networking from run_sandbox
Lifts TLS state generation, network namespace setup, proxy startup,
bypass monitor spawn, and SSH-side proxy URL / netns FD computation
out of run_sandbox into a sibling async fn `run_networking` that
returns a Networking struct. The identity cache moves with it (only
consumed by the proxy). Entrypoint PID allocation moves just above
the call site so it can be passed in.
No behavior changes — same OCSF emits, same async order, same RAII
lifetimes for the proxy and bypass-monitor handles, now held by the
returned Networking value in run_sandbox's frame.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(sandbox): extract run_process and lift netns to run_sandbox
Lifts the post-networking tail of `run_sandbox` (zombie reaper, SSH
server, supervisor session, process spawn, OPA probe, policy poll loop,
denial aggregator, wait/exit) into a sibling async fn `run_process`.
Also moves network namespace creation out of `run_networking` into a new
`create_netns_for_proxy` helper invoked from `run_sandbox`, so
`run_networking` is purely the proxy component (OPA evaluation, TLS
interception, credential injection, inference routing, gRPC control
API). The netns is then borrowed into both `run_networking` and
`run_process`.
No behavior change.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* chore(workspace): scaffold openshell-supervisor-networking and openshell-supervisor-process crates
Add empty placeholder crates that subsequent commits will populate as the
sandbox decomposition proceeds. Both crates compile clean as part of the
workspace and are picked up automatically by the existing
`members = ["crates/*"]` glob.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift DenialEvent to openshell-core
The DenialEvent struct is emitted by both the proxy/L7 layer (networking-side)
and the bypass monitor (process-side), and crosses the run_networking ->
run_process API boundary. Move it to openshell-core so the eventual
supervisor-networking and supervisor-process crates can both reference it
without depending on each other. DenialAggregator and the channel/flush
helpers stay in openshell-sandbox for now.
A thin `pub use openshell_core::DenialEvent;` re-export from
denial_aggregator.rs keeps every existing `crate::denial_aggregator::DenialEvent`
call site resolving without further edits.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift normalize_path to openshell-core
Move the lexical path-normalization helper from openshell-policy to
openshell-core::paths so it can be reached from crates that sit below
openshell-policy in the dependency graph. openshell-policy keeps its
existing public API via a `pub use` re-export, so all current call sites
(e.g. openshell-sandbox/src/policy.rs) continue to resolve unchanged.
This is a prerequisite for lifting openshell-sandbox/src/policy.rs into
openshell-core: that file's `From<ProtoFilesystemPolicy>` impl calls
normalize_path, and lifting it as-is would cycle through openshell-policy.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift SandboxPolicy and friends to openshell-core
Move openshell-sandbox/src/policy.rs (SandboxPolicy, NetworkPolicy,
ProxyPolicy, FilesystemPolicy, LandlockPolicy, ProcessPolicy, NetworkMode,
LandlockCompatibility, plus their Proto* TryFrom/From impls) to
openshell-core/src/policy.rs.
Both prospective supervisor leaves (networking and process) dispatch on
SandboxPolicy. Hosting it in openshell-core lets either leaf reach for it
without depending on the other (or on the future orchestrator).
The From<ProtoFilesystemPolicy> impl now calls the in-crate
openshell_core::paths::normalize_path lifted in the previous commit, which
is what made this move cycle-free.
Update all crate::policy::* call sites in openshell-sandbox to
openshell_core::policy::*.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move child_env from openshell-sandbox
child_env (proxy_env_vars, tls_env_vars) is process-side only — consumed
by process.rs and ssh.rs. With the orchestrator staying in
openshell-sandbox (Shape A), openshell-sandbox depends on the new leaf
crates, so process-only modules can land in
openshell-supervisor-process directly.
Add openshell-supervisor-process as a path dependency of
openshell-sandbox. Update process.rs and ssh.rs to import from
openshell_supervisor_process::child_env.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move skills from openshell-sandbox
Move the static skills installer (and its embedded resource directory)
out of openshell-sandbox into openshell-supervisor-process. The module
is process-side only — invoked once during sandbox start to drop
agent skill files into the workspace — and has no cross-leaf consumers.
Adds miette as a dependency and tempfile as a dev-dependency on
openshell-supervisor-process.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move mechanistic_mapper from openshell-sandbox
Move the mechanistic mapper (HTTP method/path → operation classifier
that derives policy proposals from connection summaries) out of
openshell-sandbox into openshell-supervisor-networking. Single internal
caller (run_policy_poll_loop in lib.rs) and only depends on
openshell-core + tracing — no cross-leaf entanglement.
First population of the openshell-supervisor-networking crate; adds
openshell-core and tracing as dependencies.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift procfs to openshell-core
Move procfs (PID lookups, ancestor walking, /proc/net/tcp socket-owner
resolution, file SHA256 hashing) from openshell-sandbox into
openshell-core. The module is consumed cross-leaf — by bypass_monitor
on the process side and by identity / proxy on the networking side —
so it has to sit below both leaves.
Adds tracing, sha2, and hex as dependencies on openshell-core.
Updates the three call sites in openshell-sandbox to import from
openshell_core::procfs.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move identity from openshell-sandbox
Move BinaryIdentityCache (path → SHA256 cache used to identify the
process behind an outbound connection) from openshell-sandbox into
openshell-supervisor-networking. The cache is consumed only by the
networking-side proxy and the orchestrator; with procfs already in
openshell-core there are no remaining cross-leaf dependencies.
Adds miette as a dependency and tempfile as a dev-dependency on
openshell-supervisor-networking. Adds a Default impl for
BinaryIdentityCache to satisfy clippy::new_without_default now that
the type is publicly exposed across crates.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move agent-proposals flag from openshell-sandbox
Move AGENT_PROPOSALS_ENABLED, agent_proposals_enabled(), and the
test-only ProposalsFlagGuard out of openshell-sandbox into
openshell-supervisor-process::proposals. The flag is read only by the
process-side policy_local route handler and the orchestrator; lifting
it to openshell-core would have made core carry sandbox-owned runtime
state without buying anything.
The test-only ProposalsFlagGuard is still consumed from networking-side
l7/rest tests today (until the wider Q2 OCSF/gRPC injection work lands).
Expose it via a new optional `test-helpers` feature on
openshell-supervisor-process so test crates opt in explicitly without
pulling tokio sync primitives into production builds.
openshell-sandbox keeps its existing crate-private path
(`crate::AGENT_PROPOSALS_ENABLED`, `crate::test_helpers`) via re-exports
so call sites and tests are unchanged.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift secrets to openshell-core
Move crates/openshell-sandbox/src/secrets.rs to crates/openshell-core/src/secrets.rs so both supervisor leaves can reach SecretResolver and the placeholder helpers without depending on openshell-sandbox.
Add base64 to openshell-core deps (only stdlib + base64 are used). Promote previously pub(crate) constructors and methods on SecretResolver to pub since cross-crate callers (provider_credentials, proxy/L7 tests) now name them across the crate boundary. Update import paths in proxy.rs, l7/{rest,relay,websocket}.rs, and provider_credentials.rs from crate::secrets to openshell_core::secrets.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift provider_credentials to openshell-core
Move crates/openshell-sandbox/src/provider_credentials.rs to crates/openshell-core/src/provider_credentials.rs. Both supervisor leaves now name ProviderCredentialState in their function signatures (run_networking takes &ProviderCredentialState, run_process takes ProviderCredentialState by value), and under Shape A leaves can't depend on openshell-sandbox, so the type must live in openshell-core.
The orchestrator (run_sandbox in openshell-sandbox) remains the only writer: it constructs ProviderCredentialState::from_environment and the policy poll loop calls install_environment on credential rotation. Both leaves stay pure readers via snapshot()/resolver().
Update import paths in proxy.rs, ssh.rs, and lib.rs from crate::provider_credentials to openshell_core::provider_credentials.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* style: rustfmt import ordering
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(ocsf): move SandboxContext singleton from openshell-sandbox
Move the process-wide OCSF SandboxContext OnceLock + LazyLock fallback + getter from openshell-sandbox/src/lib.rs into a new openshell-ocsf::ctx module. The type already lives in openshell-ocsf, so its singleton lives next to it.
Add openshell_ocsf::ctx::set_ctx() and openshell_ocsf::ctx::ctx(). The orchestrator (run_sandbox) now calls set_ctx during startup. Sandbox keeps a pub(crate) use openshell_ocsf::ctx::ctx as ocsf_ctx; re-export so the 138 existing crate::ocsf_ctx() call sites resolve unchanged.
When the sandbox modules themselves migrate into the leaf crates, they'll import openshell_ocsf::ctx directly and the re-export goes away.
Under Shape A neither leaf can depend on openshell-sandbox; both already depend on openshell-ocsf to construct events, so this adds no new dep edge.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift grpc_client to openshell-core
Both prospective leaves (supervisor-networking and supervisor-process)
need CachedOpenShellClient, AuthedChannel, and the connect/fetch
helpers. Under Shape A the leaves cannot depend on openshell-sandbox,
so the type has to live below them. openshell-core already pulls in
tonic and miette; this enables tonic's channel/tls features and adds
tokio as a direct dep.
Updates all crate::grpc_client::* call sites in openshell-sandbox to
openshell_core::grpc_client::*. No re-export shim — the call-site
count was small enough to update directly.
See architecture/plans/sandbox-split-design-choices.md for the full
rationale and trade-offs.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move denial_aggregator from openshell-sandbox
DenialAggregator and FlushableDenialSummary belong with the proxy and L7
layer that emit denials. Moves the file into openshell-supervisor-networking;
adds tokio as a regular dep there since DenialAggregator uses
tokio::sync::mpsc.
Drops the pub use openshell_core::DenialEvent re-export inside the moved
file (no longer needed cross-crate). Updates bypass_monitor.rs, proxy.rs,
and lib.rs to import openshell_core::DenialEvent directly.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move log_push from openshell-sandbox
LogPushLayer is a process-side tracing layer that streams sandbox logs
to the gateway via gRPC. Moves into openshell-supervisor-process; adds
openshell-core, openshell-ocsf, tokio-stream, tracing, and
tracing-subscriber as direct deps there.
Updates the only external call site (openshell-sandbox/src/main.rs) to
import from openshell_supervisor_process::log_push.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move bypass_monitor from openshell-sandbox
bypass_monitor reads /dev/kmsg for nftables drop log lines and emits
denial events. Pure process-side concern, called only from
run_networking which spawns it on the netns. Moves into
openshell-supervisor-process; all deps (openshell-core, openshell-ocsf,
tokio, tracing) were already declared there.
Replaces crate::ocsf_ctx() shim calls inside the moved file with
openshell_ocsf::ctx::ctx() — first leaf-side caller to import the OCSF
context singleton directly instead of going through openshell-sandbox's
re-export.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move debug_rpc from openshell-sandbox
debug_rpc is the CLI subcommand handler that exercises authenticated
gRPC calls (issue-token, refresh-token, get-config, etc.). Pure
process-side concern, called only from openshell-sandbox/main.rs.
Adds base64, hex, serde_json, sha2, and tonic (with channel/tls
features) as direct deps on openshell-supervisor-process. Updates the
single call site in main.rs.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move supervisor_session from openshell-sandbox
supervisor_session opens a bidirectional gRPC stream that lets the
gateway initiate shells inside the sandbox. Pure process-side concern,
called only from run_process. Adds uuid as a direct dep on
openshell-supervisor-process.
Replaces crate::ocsf_ctx() shim calls inside the moved file with
openshell_ocsf::ctx::ctx() — same pattern as bypass_monitor.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): lift managed_children tracker from openshell-sandbox
The MANAGED_CHILDREN set tracks PIDs of supervisor-spawned children
(entrypoint + SSH sessions) so the orchestrator's SIGCHLD reaper can
distinguish them from incidental zombies. Pure process-side concern,
moves to openshell_supervisor_process::managed_children with three
public fns: register, unregister, is_managed.
Updates lib.rs reaper, process.rs, and ssh.rs to call through the new
module path. Drops the now-unused HashSet import from lib.rs.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move sandbox hardening from openshell-sandbox
Lift the process-only hardening pieces (landlock, seccomp, PreparedSandbox,
prepare/enforce, log_sandbox_readiness, top-level apply, and
apply_supervisor_startup_hardening) from crates/openshell-sandbox/src/sandbox/
to crates/openshell-supervisor-process/src/sandbox/.
Leave netns.rs and nft_ruleset.rs in openshell-sandbox for now, since both
eventual leaf crates (supervisor-networking and supervisor-process) read from
NetworkNamespace and its final home is decided when run_networking and
run_process are extracted.
Replace crate::ocsf_ctx() shims in landlock.rs and the new linux/mod.rs with
direct openshell_ocsf::ctx::ctx() calls. Update call sites in lib.rs,
process.rs, and ssh.rs to import sandbox from openshell_supervisor_process
while keeping the netns import unchanged.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift proposals flag from openshell-supervisor-process
Move proposals.rs (AGENT_PROPOSALS_ENABLED OnceLock + agent_proposals_enabled
reader + test_helpers::ProposalsFlagGuard) from openshell-supervisor-process
to openshell-core so both eventual leaf crates can read it without depending
on each other.
The flag is process-wide singleton state initialised once during sandbox
startup and read by both the policy.local route (networking-side) and the
skills installer (process-side) — same shape as openshell_ocsf::ctx.
Move the test-helpers Cargo feature alongside it: openshell-core gains the
feature, openshell-supervisor-process loses it, and openshell-sandbox's
dev-dependency now enables openshell-core/test-helpers. Update the sandbox
re-export shim to point at openshell_core::proposals.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(core): lift netns + nft_ruleset from openshell-sandbox
Move NetworkNamespace and the nft_ruleset bypass-rule generator from
crates/openshell-sandbox/src/sandbox/linux/ to crates/openshell-core/src/netns/.
Both eventual leaf crates (supervisor-networking and supervisor-process) read
from NetworkNamespace, so it must live somewhere both can depend on without
violating the Shape A no-leaf-to-leaf rule.
Replace crate::ocsf_ctx() shims in netns with direct openshell_ocsf::ctx::ctx()
calls, matching the pattern used in already-migrated process modules. Update
super::nft_ruleset references inside netns to nft_ruleset since the module
is now a sibling sub-module of netns/mod.rs.
Add openshell-ocsf and uuid as linux-only dependencies of openshell-core, and
gate pub mod netns on target_os = "linux" since the implementation uses
netlink, ip(8), and namespace fds. Delete the now-empty sandbox/{mod.rs,
linux/mod.rs} stubs and update NetworkNamespace import paths in lib.rs and
process.rs to point at openshell_core::netns.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move process.rs and ssh.rs from openshell-sandbox
Lift the entrypoint process spawn module and the embedded SSH server
module into openshell-supervisor-process. openshell-sandbox now
re-exports ProcessHandle/ProcessStatus and calls
openshell_supervisor_process::ssh::run_ssh_server directly.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move proxy, l7, opa, policy_local from openshell-sandbox
Lift the egress proxy, L7 enforcement modules, OPA engine, and policy.local
advisor API into openshell-supervisor-networking. Move accompanying data
files (sandbox-policy.rego), test fixtures (testdata/), and integration
tests (system_inference, websocket_upgrade). Sandbox lib.rs now references
these via openshell_supervisor_networking::* and ProxyHandle::start_with_bind_addr
is exposed as pub for the orchestrator call site.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(sandbox): hoist policy poll loop and denial aggregator into orchestrator
Move the symlink-resolver, policy poll loop, and denial-aggregator flush
spawns out of run_process and into run_sandbox so run_process no longer
needs OpaEngine, retained_proto, the local policy context, the sandbox
name, the gateway endpoint for telemetry, the OCSF flag, or the denial
receiver. These long-running orchestrator-owned tasks now live alongside
the other sandbox-startup wiring, matching the design log decision in
architecture/plans/sandbox-split-design-choices.md (Q5).
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move run_process from openshell-sandbox
Lift the workload supervision entry point (zombie reaper, SSH server
spawn, supervisor session, entrypoint child spawn, exit-with-timeout)
into its own module in openshell-supervisor-process. The orchestrator
in openshell-sandbox now calls openshell_supervisor_process::run::run_process
directly. With this move run_process names only types from openshell-core,
openshell-ocsf, openshell-supervisor-process itself, std, and tokio —
no openshell-supervisor-networking dependency.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move bypass_monitor from supervisor-process
Bypass detection is network-policy enforcement: it parses nftables LOG
entries from /dev/kmsg and emits OCSF NetworkActivity / DetectionFinding
events plus DenialEvents into the same channel the proxy feeds. Its
lifetime is tied to the network namespace, not to the workload child.
Moving it to openshell-supervisor-networking puts it next to the proxy
and the denial aggregator that consume its output, and unblocks moving
run_networking out of openshell-sandbox without a leaf-to-leaf dep.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move inference route helpers from openshell-sandbox
Move build_inference_context, partition_routes, bundle_to_resolved_routes,
spawn_route_refresh, the InferenceRouteSource enum, and the route refresh
interval helpers into a new openshell-supervisor-networking::inference_routes
module along with their unit tests. The orchestrator now calls into the
networking leaf for inference context construction; the leaf owns its own
route bundle resolution end-to-end.
The new module is named inference_routes to avoid colliding with the
existing l7::inference module, which handles request-time HTTP parsing
and pattern matching rather than route bundle setup.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-networking): move run_networking from openshell-sandbox
Move the Networking handle struct, run_networking, and the Linux-only
create_netns_for_proxy helper into a new openshell-supervisor-networking::run
module. The orchestrator in openshell-sandbox now invokes
openshell_supervisor_networking::run::{create_netns_for_proxy, run_networking}
and reads the Networking fields directly; the leaf owns the entire
networking-stack startup path (CA generation, proxy task, bypass monitor,
inference context, denial channel) end-to-end.
The Networking RAII handle fields (proxy, bypass_monitor) are now public
without leading underscores so the public API satisfies clippy's
pub_underscore_fields lint while still serving as drop guards held by the
orchestrator's frame.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* fix(workspace): align Cargo deps and call sites for split crates
The recent module lifts left two Linux-only gaps that the macOS host
workspace check skipped:
- openshell-core's netns module needs libc, tempfile, and nix on Linux,
but only openshell-ocsf and uuid were carried over.
- openshell-supervisor-process's seccomp/landlock modules need landlock
and seccompiler, which still lived on openshell-sandbox.
- openshell-sandbox's runtime_pid_limit branch referenced an unqualified
process:: that pointed at the old in-crate module.
Move landlock/seccompiler to supervisor-process, add the missing core
deps, qualify the call sites, and drop sandbox deps that no longer have
runtime users (landlock, seccompiler, target-gated tempfile/uuid, the
unix libc/rustix block).
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): rename openshell-supervisor-networking to openshell-supervisor-network
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): own denial-aggregator flush end-to-end
Move the denial-aggregator spawn and flush_proposals_to_gateway out of
run_sandbox and into run_networking. The networking leaf already owns
every other input (proxy + bypass_monitor as producers, denial channel,
mechanistic_mapper, denial_aggregator) and already opens its own gRPC
connections (inference_routes, policy_local) — the orchestrator was the
only piece left straddling the boundary.
Networking now drives the full path: producers -> channel -> aggregator
-> flush -> gateway. Drops denial_rx from Networking; adds sandbox_name
to run_networking so SubmitPolicyAnalysis can resolve by sandbox name
(falls back to ID when unset). Same shape as log_push in the process leaf.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): own symlink-resolution task
Move the OPA binary-symlink resolver out of run_sandbox and into
run_networking. The task probes /proc/<entrypoint_pid>/root/ until the
workload's mount namespace is accessible, then rebuilds the OPA engine
with resolved binary paths so policy rules match canonical names instead
of symlinks.
Both inputs (Arc<OpaEngine>, retained_proto) are networking-leaf concerns
and were already plumbed into run_networking; the entrypoint_pid Arc is
read lazily after the process leaf populates it. Adds retained_proto as
a parameter and spawns the resolver early in run_networking so the probe
loop starts before the proxy comes up.
Same shape as the denial-flush move: networking owns its own background
task end-to-end; the orchestrator stops hosting work that doesn't
conceptually belong to it.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move seccomp install into run_process
The supervisor seccomp prelude is part of "set up the workload-side
process tree", not part of orchestration. Move the call site from
run_sandbox into the top of run_process and drop the now-unused
re-export from openshell-sandbox::lib.
Timing is preserved: by the time the orchestrator calls run_process,
run_networking has already returned, so netns + nftables setup is
complete.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move check_runtime_pid_limit into run_process
The PID-limit precondition is process-side: it gates whether the workload
child can be spawned at all. Move the call from run_sandbox into the top
of run_process, alongside the seccomp prelude. Same shape as the seccomp
move — function already lives process-side, only the call site relocates.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move validate_sandbox_user to process crate
The sandbox-user check is a precondition for privilege-dropping the
workload child; it has no relevance to networking. Move the function
next to drop_privileges in openshell-supervisor-process::process and
call it from the top of run_process.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move prepare_filesystem to process crate
Creating and chowning read_write directories is workload-side
preparation, not orchestration. Move prepare_filesystem and its
prepare_read_write_path helper (plus tests) into
openshell-supervisor-process::process and call from run_process,
alongside validate_sandbox_user.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-process): move startup skill install into run_process
The eager initial-settings fetch + agent skill install is process-side:
the install materializes files the workload's filesystem sees. The
orchestrator still owns the AGENT_PROPOSALS_ENABLED OnceLock init
because the policy poll loop also reads it; only the early fetch and
install hop into run_process.
Behavior unchanged. Best-effort: any RPC or install failure is logged
but does not fail startup.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): own PolicyLocalContext construction
Move the PolicyLocalContext construction from run_sandbox into
run_networking. The orchestrator was building it solely to thread it into
the networking leaf and to share it with the policy poll loop; now
run_networking builds it from inputs it already takes (retained_proto,
openshell_endpoint, sandbox_name|sandbox_id) and exposes it on the
returned Networking struct.
The orchestrator's poll loop now grabs the Arc clone from
networking.policy_local_ctx, so the orchestrator no longer imports
openshell_supervisor_network::policy_local.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* feat(supervisor): add --mode flag to gate network/process leaves
Add a --mode flag (default "network,process") that selects which
supervisor leaves run in the current process. Two new shapes are
unlocked without splitting the binary:
--mode=network # network-only sidecar
--mode=process # process-only supervisor
--mode=network,process # combined (default; current behavior)
In network-only mode the orchestrator skips run_process and waits on
SIGINT/SIGTERM before tearing down the proxy. The entrypoint PID stays
at 0 for the lifetime of the process, which silently degrades the
proxy's binary-identity TOFU and the bypass monitor's PID enrichment;
this is correct in a split-pod topology where the workload's /proc
lives in another pod.
In process-only mode run_networking is skipped entirely. SSH sessions
get no proxy URL, no netns FD, and no CA paths, matching what a
split-pod consumer would expect when network enforcement is delegated
to a sidecar.
The policy poll loop continues to run unconditionally; its OPA-reload
and policy.local hooks already gate on the resources only present when
network is enabled, and the env-refresh / proposals-toggle hooks
remain active in process mode.
Closes a step toward the RFC-0001 supervisor topology proposed in
issue #1305 by drew.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* style(supervisor-process): rustfmt long debug! line
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): pull DenialEvent down from core
DenialEvent is only emitted and consumed inside openshell-supervisor-network
(proxy, bypass monitor, denial aggregator). It never crossed the leaf
boundary, so the earlier lift to openshell-core was speculative. Move it
back into the network crate where its only callers live.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): pull procfs down from core
procfs was lifted to openshell-core under the assumption it would be
shared cross-leaf, but on the current branch all three callers
(bypass_monitor, identity, proxy) live in openshell-supervisor-network.
No file in openshell-supervisor-process imports it. Move the module to
the network crate and drop sha2/hex from openshell-core, which were
pulled in only for procfs.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* style(supervisor-network): run cargo fmt
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* fix(supervisor-network): add libc dev-dependency for procfs tests
The procfs/bypass_monitor/proxy test modules use libc::{fork, exec,
fcntl, kill, waitpid} but the dep wasn't declared in this crate's
Cargo.toml. It was previously satisfied transitively when these
modules lived in openshell-core; the move left the test target
unable to resolve libc.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(sandbox): move denial aggregator to orchestrator
The denial aggregator and mechanistic mapper consume denial events
produced by the proxy and (subsequently) the bypass monitor. With both
supervisor leaves becoming pure producers of `DenialEvent`, the
consumer-side aggregation belongs in the orchestrator, not in either
leaf.
Move `denial_aggregator.rs` and `mechanistic_mapper.rs` from
`openshell-supervisor-network` to `openshell-sandbox` (the
orchestrator). The orchestrator now owns the unbounded denial channel:
it constructs `(tx, rx)`, hands `tx` to `run_networking` for the proxy
to clone, drains `rx` via the aggregator task, and runs the gateway
flush helper itself.
`run_networking`'s signature gains a `denial_tx` parameter and loses
its internal channel construction, aggregator spawn, and
`flush_proposals_to_gateway` helper. `DenialEvent` stays in
`openshell-supervisor-network` for now; a follow-up commit will lift
it to `openshell-core` alongside the bypass monitor relocation.
* refactor(supervisor-process): pull bypass monitor down from network
`bypass_monitor` is process-isolation machinery: it tails the kernel
log via `dmesg --follow`, parses nftables LOG lines emitted from the
workload's network namespace, resolves PIDs via `/proc`, and emits
OCSF events plus optional `DenialEvent`s. None of this touches the
proxy, OPA, TLS, or any other supervisor-network state — it only
shared the denial channel because both feed the same aggregator.
Move `bypass_monitor.rs` from `openshell-supervisor-network` to
`openshell-supervisor-process` (as `bypass_monitor/mod.rs`). Spawn it
in `run_process` where the netns name and entrypoint PID are already
in scope. The orchestrator hands an extra `bypass_denial_tx` clone of
the denial channel sender to `run_process` for this purpose.
Lift `DenialEvent` from `openshell-supervisor-network` to
`openshell-core`. Both supervisor leaves now produce it, so it needs
a shared location that neither leaf depends on. This reverses an
earlier commit that pulled the type into the network leaf when it was
the only producer.
Copy the minimal subset of `/proc` parsers used by `bypass_monitor`
into a private `bypass_monitor::procfs` submodule. The alternative —
extracting a shared procfs crate — is a much larger refactor that
this commit does not need; supervisor-network's `procfs.rs` continues
to serve the proxy and identity cache.
* refactor(supervisor-process): derive ssh netns fd inside run_process
The ssh_netns_fd was computed in run_networking purely to forward it
through the Networking struct and back into run_process. supervisor-network
never read it. Move the derivation to run_process where the
NetworkNamespace handle is already in scope.
* refactor(supervisor-process): derive ssh proxy url inside run_process
The ssh_proxy_url was computed in run_networking purely to forward it
through the Networking struct and back into run_process. supervisor-network
never read it. Move the derivation to run_process where the
NetworkNamespace handle and SandboxPolicy are already in scope.
After this commit the Networking struct no longer carries any SSH-shaped
fields, and supervisor-network reads only host_ip from the netns (for the
proxy bind address).
* refactor(supervisor-network): take proxy bind ip directly instead of netns
run_networking only ever read host_ip from the netns it was passed (the
SSH plumbing reads moved to run_process in earlier commits). Replace the
NetworkNamespace parameter with a plain Option<IpAddr> the orchestrator
extracts. supervisor-network's run module no longer references the netns
type for any consumer, only for create_netns_for_proxy (which still lives
in this crate; relocates next).
* refactor(supervisor-process): move netns ownership out of core
Relocates the NetworkNamespace handle, nft ruleset builder, and
create_netns_for_proxy constructor into openshell-supervisor-process.
The orchestrator (openshell-sandbox) phantom-owns the RAII handle for
the duration of run_sandbox; supervisor-network no longer references
the type at all.
Drops uuid, libc, nix, openshell-ocsf, and tempfile from core's Linux
target deps (all were exclusive to netns). tempfile becomes a Linux
runtime dep on supervisor-process for nft ruleset application.
* chore(sandbox): prune leaf-only deps from orchestrator manifest
cargo-machete flagged 26 direct dependencies that were carried over
from the pre-split monolith and are no longer used by the orchestrator
itself: regorus, russh, rcgen, tokio-rustls, ipnet, apollo-parser,
openshell-router, anyhow, base64, bytes, flate2, glob, hex, hmac, nix,
rand_core, rustls-pemfile, serde, serde_yml, sha1, sha2, thiserror,
tokio-stream, uuid, webpki-roots.
These now live (transitively) in openshell-supervisor-network and
openshell-supervisor-process where they are actually consumed.
* chore(deps): prune unused deps from supervisor crates
- Drop unused `url` from openshell-supervisor-network.
- Mark `prost` and `prost-types` as cargo-machete-ignored in
openshell-core: they have no source-level `use`, but the tonic-
generated proto code references them via `::prost::Message` etc.
- openshell-supervisor-process is already clean.
* fix(supervisor-network): wait for entrypoint PID before symlink probe
The OPA symlink-resolution task reads entrypoint_pid once at the top
of the spawned closure. Because the spawn happens before run_process
publishes the workload PID, the load returns 0, the probe path bakes
in as /proc/0/root/, and the loop exhausts its retries against a path
that does not exist on Linux. The reload never fires, so policies that
whitelist symlinked binaries (e.g. /usr/bin/python3 → python3.11) get
silent denials when the workload exec's the realpath.
Split the wait into two phases: 5s polling entrypoint_pid for a
non-zero value, then the existing 5s window probing /proc/<pid>/root/.
Distinct warn messages on each timeout so future debugging can tell
"PID never published" apart from "container fs never appeared".
* fix(sandbox): restore GPU procfs baseline (#1522)
Signed-off-by: Evan Lezar <elezar@nvidia.com>
(cherry picked from commit 5102cb9413)
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* fix(supervisor-process): use renamed tonic tls-native-roots feature
Upstream renamed the tonic `tls` feature to `tls-native-roots`. The
supervisor-process Cargo.toml still referenced the old name, which broke
the workspace build after merging upstream.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* refactor(supervisor-network): relocate token_grant and spiffe_endpoint
Upstream's SPIFFE-backed token grant feature landed in
crates/openshell-sandbox/src/. After the supervisor split, the L7
enforcement code in supervisor-network calls into token_grant, which
would require supervisor-network to depend back on sandbox.
Move token_grant.rs and spiffe_endpoint.rs into supervisor-network
where the only callers live, add the reqwest and spiffe deps to
supervisor-network's Cargo.toml, and drop them from sandbox.
Also fix two stale `openshell_core::proto::` self-references in
openshell-core (a pre-existing breakage that surfaced once the rest of
the merge compiled).
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
* fix(supervisor-process): broaden Path import cfg to all unix targets
The `Path` import was gated on `cfg(any(test, target_os = "linux"))`,
but `prepare_read_write_path` is gated on `cfg(unix)` — broader. On
non-Linux unix the function still referenced `&std::path::Path`
explicitly, so upstream's qualified path was load-bearing.
After the supervisor split, lint runs on Linux where `Path` IS in
scope, so `unused_qualifications` fires. Broaden the import cfg to
match the function's cfg and use the bare `Path` name everywhere.
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
---------
Signed-off-by: Radoslav Hubenov <rrhubenov@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>