mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-02 07:34:45 +08:00
main
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
70a0f6c547 | refactor(cli): remove gateway lifecycle management (#1221) | ||
|
|
d45c1a704e |
fix(scripts): eliminate xargs subshell dependency in docker-cleanup.sh (#1207)
Replace xargs usage with native docker/podman multi-argument inspect calls. The previous implementation failed because xargs spawns subshells that don't inherit the ce() function from container-engine.sh. Instead of piping container IDs through xargs, collect them into an array and pass them directly to `ce inspect`, which accepts multiple IDs. This eliminates the subshell issue entirely and simplifies the code. Fixes the docker:cleanup mise task that was failing with: xargs: ce: No such file or directory Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
f8fb382146 |
fix(scripts): handle docker cleanup when no containers are running (#977)
The docker-cleanup.sh script failed when no containers were running because grep -v returned exit code 1 on empty input, causing the script to abort due to set -euo pipefail. Add || true to the volume detection pipeline so the script succeeds when there are no running containers (in_use_volumes will be empty, which is the correct behavior). Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
d44d8a1e27 |
feat: Openshell driver podman (#904)
* feat(podman): add Podman compute driver for rootless sandbox management Adds openshell-driver-podman, a new compute driver that manages OpenShell sandboxes as rootless Podman containers via the Podman REST API over a Unix socket. Enables local workstation sandboxes without Kubernetes. Driver features: - Bridge networking with ephemeral host-port mapping for rootless SSH reachability - Named volumes for workspace storage, Podman native health checks, GPU via CDI - Supervisor binary sideloaded via image volume mount (BYOC-compatible) - SSH handshake secret injected via Podman secrets API (not plaintext env) - Typed ContainerSpec structs, input validation, and path-traversal guards - Cgroups v2 required; fails fast on v1 hosts - Bounded event stream buffer; watch stream reconnection handled by server watch_loop - Graceful shutdown and standalone driver binary with gRPC bridge Rootless-specific fixes: - Skip drop_privileges when user namespace lacks SETUID/SETGID/DAC_READ_SEARCH caps - Add /run/netns tmpfs mount for ip netns in rootless containers - Use secret_env map (not secrets array) for env-var injection in libpod API - Resolve SSH endpoint to 127.0.0.1:<host_port> instead of unreachable bridge IP Server/sandbox hardening: - Split loopback and link-local SSRF gates; Podman/VM drivers allow loopback - Close SSRF bypass in SSH tunnel Host path by resolving DNS before connecting - Prevent OPENSHELL_* env var override by user-supplied spec environment maps - Disable SQLite pool idle_timeout/max_lifetime for in-memory databases - Emit deleted_event on 404-during-inspect instead of regressing sandbox phase - Key delete cleanup by stable sandbox_id to survive container label drift CLI fixes: - Restore --name as a named flag on sandbox create (not positional) - Fix exec command arg parsing to not consume sandboxed-command flags - Propagate SSH verbosity via OPENSHELL_SSH_LOG_LEVEL Build tooling: - Add tasks/scripts/container-engine.sh: auto-detects Podman or Docker, exposes unified ce_* helpers; all build/cluster/VM scripts updated to use it - Add docker:build:supervisor mise task for standalone supervisor image - Add openshell-driver-podman to Dockerfile.images pre-fetch/build stages - Add e2e/rust/e2e-podman.sh and e2e:podman mise task for full lifecycle testing Signed-off-by: Adam Miller <admiller@redhat.com> * fix(driver-podman): derive grpc endpoint from server bind port When a user starts the gateway on a non-default port (e.g. --port 8081), sandbox containers were receiving OPENSHELL_ENDPOINT pointing at the default port 8080. The driver's auto-detection fallback read OPENSHELL_BIND_ADDRESS from the environment, which was stale or unset, and fell back to DEFAULT_SERVER_PORT. Add gateway_port to PodmanComputeConfig and thread config.bind_address.port() from the server into the driver so the fallback uses the actual listening port. Remove the OPENSHELL_BIND_ADDRESS env var read and the extract_port_from_bind_address helper which are no longer needed. Add --gateway-port / OPENSHELL_GATEWAY_PORT to the standalone driver binary for parity when the driver is run outside the embedded server path. Signed-off-by: Adam Miller <admiller@redhat.com> * fix(driver-podman): address PR feedback on env test safety and cluster DNS docs Replace hand-rolled unsafe TempEnvVar RAII guard with temp_env::with_vars and a static ENV_LOCK mutex, fixing a data race in parallel test execution. The prior safety comment incorrectly claimed Cargo runs tests single-threaded. Update debug-openshell-cluster skill to accurately document the DNS proxy strategy (setup_dns_proxy + public DNS fallback) and clarify the separation between cluster DNS and sandbox agent DNS enforcement. Signed-off-by: Adam Miller <admiller@redhat.com> * fix(e2e): resolve CI failures in auth timeout, test harness, and formatting - Short-circuit browser_auth_flow when OPENSHELL_NO_BROWSER=1 instead of waiting the full 120s AUTH_TIMEOUT for a callback that never arrives - Add timeout to SandboxGuard::create() and create_with_upload() to prevent indefinite hangs (matches create_keep() which already had one) - Add missing '--' separator in no_proxy test before command args - Add #![cfg(feature = "e2e")] gate to sandbox_lifecycle.rs - Run cargo fmt on openshell-driver-podman - Refine cluster DNS docs for Podman in debug-openshell-cluster skill Signed-off-by: Adam Miller <admiller@redhat.com> * refactor(server): remove allows_loopback_endpoints from ComputeRuntime SSRF protection is now handled at the network and proxy layers (openshell-core net.rs, openshell-sandbox proxy.rs) rather than requiring per-driver flags on ComputeRuntime. Update architecture docs to reflect supervisor relay SSH transport and add rootless networking deep-dive. Signed-off-by: Adam Miller <admiller@redhat.com> --------- Signed-off-by: Adam Miller <admiller@redhat.com> |
||
|
|
454327d890 |
feat(policy): add policy recommendation plumbing (#204) (#222)
* feat(policy): add policy recommendation plumbing — denial aggregation, transport, approval pipeline, and mechanistic recommendations Implement the infrastructure layer for automated policy recommendations (#204): - Proto: 9 new RPCs and messages for draft policy lifecycle (submit, get, approve, reject, approve-all, edit, undo, clear, history) - Persistence: SQLite/Postgres migrations and store methods for draft_policy_chunks and denial_summaries tables - Server: Full gRPC handler implementations with mechanistic mapper that auto-generates NetworkPolicyRule proposals from denial summaries - Sandbox: DenialAggregator with MPSC channel, deduplication, periodic flush to gateway via SubmitPolicyAnalysis - CLI: 'openshell draft' subcommand with get/approve/reject/approve-all/undo/clear/history operations - TUI: Draft recommendations panel accessible from sandbox policy view - Docs: Architecture documentation in architecture/policy-advisor.md * feat(policy): add L7-aware mechanistic mapper and policy advisor CTF example Add L7 rule generation to mechanistic mapper (build_l7_rules, generalise_path, looks_like_id) with 3 new unit tests. Add examples/policy-advisor/ with a 7-gate CTF script, restrictive sandbox policy, and walkthrough README. * fix(policy): use sandbox name for denial flush and add TUI draft badges Fix denial aggregator passing sandbox UUID instead of name to SubmitPolicyAnalysis, which caused 'sandbox not found' errors on flush. Add notification badges to the TUI sandbox list and detail header showing pending draft recommendation counts. * fix(policy): deduplicate draft chunks and tolerate overlapping OPA rules Skip draft chunk creation when a pending/approved chunk already covers the same host:port endpoint, preventing duplicate rules across denial aggregator flush cycles. Rewrite three OPA complete rules (network_policy_for_request, matched_network_policy, matched_endpoint_config) to tolerate multiple matching policies without triggering a "complete rule conflict" error. network_policy_for_request becomes a boolean, matched_network_policy uses a set comprehension with min(), and matched_endpoint_config uses an array comprehension with index-0 selection. * feat(tui): interactive draft actions, highlight bar, and detail popup Rework the draft recommendations panel to match the logs UX: - Highlight bar (green accent + background) instead of arrow marker - Viewport-aware j/k scrolling with g/G for top/bottom - Enter opens a full-screen detail popup showing endpoints, binaries, rationale, security notes, and action hints Add approve/reject/approve-all draft actions: - [a] approve selected chunk, [x] reject, [A] approve all pending - Actions work from both the list view and the detail popup - gRPC calls run async; result updates status bar and refreshes data - Nav bar shows all available keybindings Fix draft count refresh: sandbox_draft_counts now refreshes on every tick (not just Dashboard), so the detail header badge updates in real time. Improve badge labels: show 'N pending' instead of a bare number in both the dashboard sandbox list and sandbox detail header. * refactor(policy): DB-level draft chunk dedup with hit counter and timestamps Replace the in-memory HashSet dedup in SubmitPolicyAnalysis with a database-level upsert. New denormalized columns on draft_policy_chunks: - host, port: extracted from proposed_rule at insert time - hit_count: incremented on conflict (same sandbox + host + port) - first_seen_ms, last_seen_ms: track when the endpoint was first and most recently proposed A partial unique index (WHERE status IN ('pending','approved')) ensures only one active chunk per endpoint per sandbox; rejected/superseded chunks don't block new proposals. Surface hit_count and first/last_seen in: - CLI: 'openshell draft get' shows 'Hits: N (first ..., last ...)' - TUI: detail popup shows hits row; list view shows 'Nx' suffix * fix(policy): optimistic retry on policy version conflicts + structured logging merge_chunk_into_policy and remove_chunk_from_policy now retry up to 5 times on UNIQUE constraint violations (version conflicts from concurrent approvals). Each attempt re-reads the latest policy, re-merges the rule, and increments the version. This eliminates the race condition where rapid successive approvals would fail with a DB error. Add structured tracing to all draft action handlers: - ApproveDraftChunk: logs rule_name, host, port, hit_count before merge and version + policy_hash after success - RejectDraftChunk: logs rule_name, host, port, reason - ApproveAllDraftChunks: logs pending_count at start, per-chunk merge progress, and final summary with chunks_approved/skipped - UndoDraftChunk: logs before/after with rule_name and version - Retry attempts log as warnings with attempt number and conflicting version * wip: forward proxy fix, mapper allowed_ips, TUI polish, CTF rewrite * fix(tui): use correct --gateway flag for ssh-proxy ProxyCommand * chore: add Docker cleanup script for stale images, volumes, and build cache * feat(tui): approve-all confirmation modal and CTF cleanup Add [A] confirmation popup that snapshots pending chunks, shows a scrollable list, and approves each chunk individually on confirm. This prevents approving chunks that arrived after the modal opened. Remove transient issue #205 reference from CTF victory banner. * fix(tui): correct import ordering for rustfmt * wip: stateful toggle model, rename to network rules Draft chunks now follow a toggle state machine: pending -> approved | rejected (initial decision) approved <-> rejected (toggle) One row per (sandbox_id, host, port) via expanded unique index. Rejecting an approved rule removes it from the active policy. Re-approving a rejected rule merges it back. Rename CLI from 'draft' to 'rule', TUI from 'Draft Recommendations' to 'Network Rules'. State-aware keybindings: approved shows [x] Revoke, rejected shows [a] Approve. Fix sandbox detail hiding delete confirmation behind pending message. * refactor(policy): move mapper sandbox-side, slim schema, per-binary granularity Move mechanistic mapper from gateway to sandbox so all analysis runs sandbox-side (N sandboxes = N independent pipelines). Gateway is now a thin validate + persist + approval layer. Architectural changes: - Move mechanistic_mapper.rs from navigator-server to navigator-sandbox - Sandbox flush flow: aggregator drains -> mapper runs -> proposals sent - Gateway SubmitPolicyAnalysis: validate + persist only, no mapper - Drop denial_summaries table (write-only, zero readers) - Consolidate migrations 003+004+005 into single 003 Schema slimming: - Drop 5 unused columns from draft_policy_chunks (stage, denial_refs, supersedes_chunk_id, analysis_mode, decided_by) - Add per-binary granularity: binary column, widen unique index to (sandbox_id, host, port, binary) - Mapper groups by (host, port, binary), one proposal per triple - Merge appends binary to existing rule; revoke removes just that binary CTF & UX: - 7-gate CTF: add Gate 3 (curl -> ifconfig.me:80) for per-binary demo - TUI shows binary short name in list, full path in detail popup - CLI output shows binary field - Idempotent rule names, hit_count accumulates real denial counts - Rationale text no longer bakes in stale denial count |