mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-04 00:23:53 +08:00
dev
57
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6e369f2396 |
chore(agents): simplify contributor instructions and workflows (#3987)
* chore(agents): simplify contributor instructions and workflows Closes #3980 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(contributing): scope verification to affected components Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * docs(contributing): standardize issue branch naming Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
0e8d9e55f9 |
test(e2e): remove schema parity campaign (#3864)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
cf1bbb965d |
docs: remove the architecture directory (#3799)
Remove architecture/. It was a constant source of merge conflicts, became an effectively append-only log of the project, and was of dubious value. Design records live in rfc/, crate details in crate READMEs, and user documentation in docs/. Move the git-ignored plans directory from architecture/plans to plans/, keeping the old .gitignore entry. Remove the arch-doc-writer agents and update AGENTS.md, CONTRIBUTING.md, skills, the feature request template, and links in proto/, rfc/, and examples/ that pointed into architecture/. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
9244868056 |
docs: refresh architecture and agent guides (#3705)
* docs: refresh architecture and agent guides Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: describe updated security architecture neutrally Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: highlight new isolation primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sandboxes): clarify how to disconnect Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align architecture and guides with current navigation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(extensibility): streamline extension authentication guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
d91b1999a0 |
feat(api)!: use sandbox names as canonical RPC references (#3272)
* feat(api)!: use sandbox names as canonical references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(cli): update forward color fixture for workspace scope Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): use sandbox names for settings lookup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox request fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): use canonical sandbox receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): harden sandbox mutation handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): update rebased sandbox references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(api)!: standardize canonical entity references Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(api): codify protobuf API conventions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve workspace selector semantics Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): restore workspace selector parity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): preserve descriptive name fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(api): update e2e request fixtures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(core): omit workspace selector during bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): remove proto convention checker Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(api): refresh schema fingerprints after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(cli): use canonical provider receipt field Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
58b5f8f976 |
feat(prover): add standalone policy boundary checker (#3289)
* feat(prover): add standalone policy maximum checker Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): simplify check scope schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align containment and cancellation with runtime Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): stabilize containment checks in CI Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): avoid solver in fast-path guard test Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align string containment with runtime Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): reject ambiguous z3 string escapes Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): align containment with runtime boundaries Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover-cli): harden cancellation and invalid input Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(packaging): install policy prover with OpenShell Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): clarify installation and check results Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(build): describe prover distribution directly Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): rename maximum policy to boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): localize fail-closed validation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): cover fail-closed CLI surfaces Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(prover): allow CI load for REST solver proof Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): use canonical policy schema for containment Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): preserve uncertainty for runtime binary globs Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): bound policy validation work Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): document validation resource limits Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(prover): make containment API extensible Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(prover): define containment API contract Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(ci): integrate prover with consolidated builds Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(ci): declare release packaging dependency Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
2ccef97769 |
feat(policy): establish one canonical authored policy representation (#3334)
* chore(policy): restart schema implementation Signed-off-by: Johnny Greco <jogreco@nvidia.com> * feat(policy-schema): add canonical authored policy model Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): use canonical authored schema Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(prover): project canonical policy documents Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(policy): describe shared schema boundary Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): close reviewed parser gaps Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): fail closed on unsupported fields Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy): preserve partial process identities Signed-off-by: Johnny Greco <jogreco@nvidia.com> * fix(policy-schema): harden authored policy inspection Signed-off-by: Johnny Greco <jogreco@nvidia.com> * refactor(policy): rename policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * revert(policy): restore policy schema crate Signed-off-by: Johnny Greco <jogreco@nvidia.com> * test(e2e): serialize OIDC PKCE scenarios Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
c1f2e7189f |
feat(isolation): implement the RFC 0012 sandbox architecture (#2942)
* feat(isolation): add RFC 0012 backend contract Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * refactor(isolation): name the interface crate explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): expose trusted host gateway Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(agents): inventory the MXC driver Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add mediated DNS transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): tighten interface error and digest contracts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): remove unrelated driver inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): define capability-free launch contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): seal confirmed boundary state Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate confirmation for external backend implementations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): clarify mediated DNS identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): unify typed network mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): bind launches to sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(mxc): initialize extended sandbox status Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add boundary protocol and Linux primitives Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden signals and separate process status from transport Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate remote confirmation through public contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): validate wire state and propagate snapshot failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(isolation): import owned agent specification explicitly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(isolation): describe mediated DNS channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): bound mediation attach without nested retries Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): generalize loopback protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add transport-neutral session authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): separate sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime boundary controls Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): add terminal boundary operation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): split supervisor and sandbox runtimes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): harden boundary isolation and lifecycle ownership Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): reject private root redirects and adopt typed errors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve accept thread ownership on musl Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): isolate credential probes from filtered threads Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): return retained exec exit status to independent waiters Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound network mediation and preserve socket authorization Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bound control admission and retire stale mediation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): select migrated drivers per stack layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): implement loopback connector Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(isolation): authenticate the Sandbox Protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): consume dedicated backend crate Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(sandbox): align topology session fixture Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): align projected bootstrap bundle Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): validate refreshed credentials before rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): fail closed across supervisor disconnects Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): repair rebased sandbox CI Signed-off-by: Drew Newberry <anewberry@nvidia.com> * build(runtime): publish separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(config): configure the sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate sandbox binary linkage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(isolation): use backend and runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): use a scratch runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): refresh schema and dependency policy Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): bind reconnects to supervisor process Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs: align runtime split operational guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(security): document Kubernetes runtime RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): enforce runtime lifecycle invariants Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(compute): identify sandbox start generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): restore sandbox launch sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): support authenticated runtime replacement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): bind sandbox session successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): retry pending sandbox successors Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(vm): run the supervisor outside the guest workload Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): use unified build toolchain Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): use sandbox runtime terminology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): own guest network bootstrap Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): expose guest init version Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): select native supervisor artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): guard guest init Linux symbols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): scope Linux test imports Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): avoid guest interface casts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): reconcile admitted sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): share resolved sandbox identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): surface host supervisor failures Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): include guest logs on supervisor exit Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate and clean runtime generations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): keep shared paths in the base layer Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): isolate workloads behind the host supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve host gateway alias resolution Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(docker): use separate sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): restore startup validation after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): narrow supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): close companion isolation gaps Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): align mediated network expectations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(docker): exercise mediated network paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): attach supervisor to managed network Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): defer supervisor recovery until gateway is ready Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): preserve workloads during session rotation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(docker): remove unrelated configuration RFC changes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): add proxy-pod isolation topology Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): use stable sandbox service authority Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(kubernetes): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): adapt proxy pods to current runtime APIs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): describe the single runtime placement Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(kubernetes): simplify sandbox orchestration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): validate deployment prerequisites Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): update Trivy Helm profile inventory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(kubernetes): update Trivy scan inventory count Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): reuse preloaded runtime images in e2e Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): type and clean runtime resources Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): make sandbox restarts recoverable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): preserve supervisor egress Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): adopt isolated sandbox and supervisor containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): stage bootstrap archives at named volume destinations Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): rotate launch-scoped authentication Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use sandbox backend protocol Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): use host networking for supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(podman): split sandbox and supervisor images Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): repair rebase integration Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(podman): name the sandbox runtime directly Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provision supervisor CA runtime storage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): address isolation review findings Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): inspect Debian supervisor provenance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): use libpod-compatible tmpfs options Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind verified sandbox runtime binary Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): provide external driver data directory Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): start sandbox before joining user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): separate supervisor user namespace Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make sandbox starts generation-aware Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): rotate restored sandbox sessions Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): bind sandbox session lineage Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(isolation): add TCP and DNS benchmark harnesses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): align benchmark timing and supported protocols Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): report TCP benchmark metrics accurately Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(perf): cancel failed worker startup Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(docker): build matching local supervisor image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(podman): make local sandbox smoke test runnable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): wire local sandbox runtime image Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(kubernetes): narrow sandbox service RBAC Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(ci): validate split runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): harden runtime session handling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(supervisor): add standalone network proxy role Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): remove implementation companion notes Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(vm): standardize runtime release name Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(vm): pin renamed runtime artifacts Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): persist sandbox runtime identity Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(runtime): restore branch validation Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(isolation): reconcile main after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(network): close unframed HTTP 1.0 responses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * chore(isolation): preserve upstream OCSF updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(security): close credential and TLS replay paths Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(auth): make sandbox refresh retries idempotent Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
3c0f58872e |
refactor(ocsf): rename SandboxContext to EventContext (#3263)
Rename the shared OCSF context and its callers without changing event behavior. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
a0814443f1 |
feat(docs): publish versioned release snapshots (#3149)
* feat(docs): add version availability labels Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): use supported Python for sync Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(docs): publish versioned docs from releases Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): format dev version label Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(docs): upgrade Fern CLI to 5.112.0 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): make release publishing monotonic Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): preserve snapshot release identity Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(fern): document versioned publishing Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(docs): cover explicit snapshot rollback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(docs): use Fern refs for versions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): bundle components for ref versions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * revert(docs): keep complete version copies Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): publish latest and dev channels Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
8bc7955263 |
feat(skills): separate public and contributor workflows (#2899)
* feat(skills): separate public and contributor workflows Closes #2736 Publish the four user-facing OpenShell skills from the top-level skills directory, mark contributor workflows internal, and update portability guidance, validation, and documentation. Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): clarify public skill audit scope Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): use markdown documentation links Signed-off-by: Johnny Greco <jogreco@nvidia.com> * docs(skills): align public and contributor guidance Signed-off-by: Johnny Greco <jogreco@nvidia.com> --------- Signed-off-by: Johnny Greco <jogreco@nvidia.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
c8f13205e3 |
ci(release): publish prerelease artifacts (#3093)
Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
bb70461878 |
test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): add portable CLI conformance baseline Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(conformance): add standalone CLI runner Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): run conformance in gateway lanes Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
4eaa1051f8 |
docs: correct some comments and references (#3059)
OCSF_VERSION has been 1.8.0; two descriptions still said 1.7.0. The default_gateway_id doc comment described a hostname fallback the implementation never had. A per-replica default would be wrong here: gateway_id is embedded in the JWT iss/aud, so it has to be stable across replicas. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
e457974a52 |
feat(docker): export driver traces over OTLP (#2851)
Mirror the VM and Podman driver tracing setup for Docker. Export Docker driver spans through OTLP/gRPC as the distinct openshell-driver-docker service, preserve gateway trace context, record lifecycle and asynchronous provisioning operations, and report gRPC failures. Docker currently runs in-process when selected as a built-in gateway driver. Add the same temporary server-boundary shim used by Podman so traces retain the shape they will have when Docker moves to a separate process. Generalize the gateway provider selection for both in-process drivers and share the OTLP collector fixture across Docker, Podman, and VM tracing tests. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
7fc6138981 |
docs(agent): warn on missing workflow labels (#2815)
Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
bdabb54cb3 |
fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(extension-core): verify gateway JWTs Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): keep inbound verification external This should become an extension SDK package. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(security): harden the extension authentication contract Follow-up hardening on the alpha extension authentication mechanism. Claim contract: - Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They share a signing key with sandbox-to-gateway admission tokens and were otherwise separated by audience alone, so a verifier that neglects to check `aud` could accept a gateway credential. The header is a second, independent discriminator. - Publish OIDC-shaped discovery at `/.well-known/openid-configuration` so a service configured with only the gateway URL can learn the exact expected issuer and the JWKS location. It is shaped, not compliant: `issuer` is the gateway identity, not the serving URL. Audience agreement: - `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`. The audience is otherwise configured independently on each side of the boundary, where a mismatch surfaces only as an opaque authentication failure on every call. OpenShell now compares the two and fails at startup. An empty field keeps the check off for existing services. Compatibility: - Add `allow_insecure_transport` per registration. Enabling gateway JWT signing previously made any plaintext endpoint a hard startup failure, including the endpoint form used in our own documentation. The opt-out attaches no credential, is refused by the gateway if a supervisor asks for one, and warns at every startup. - Make the transport requirement kind-aware. A middleware endpoint must be reachable from every sandbox supervisor, so only interceptors may use a gateway-local Unix socket. Credential lifecycle: - Replace the process-global slot map with a supervisor-owned `ExtensionCredentialStore` shared explicitly across the gateway connections the supervisor opens, removing test-order coupling. - Rotate only when a credential is missing or has passed four fifths of its lifetime. Configuration polling ran every ten seconds against fifteen-minute credentials, so each poll re-ran gateway effective-policy resolution and re-minted the gateway token. - Bound credential minting per sandbox, since each request resolves the caller's effective policy. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs: record alpha extension authentication in RFC appendices Restore the RFC 0009 and 0010 bodies to their accepted text and move every extension-authentication update into appendices instead. An RFC records a decision at a point in time; superseding detail belongs alongside it rather than rewritten into it. RFC 0009's appendix carries the shared contract: claims, authorization, key distribution, the `allow_insecure_transport` replacement for the body's `allow_insecure`, and residual risks. RFC 0010's records only what differs for interceptors and links to it. The existing protocol-extensions appendix, which parked the phase 2 transport question, now points forward to what was built. Also document the audience handshake, the discovery endpoint, the `typ` requirement, and `jti` replay guidance in the extensibility and gateway configuration pages, and correct the middleware transport guidance: middleware endpoints must be reachable from sandbox supervisors, so Unix sockets are not an option there. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): abstract extension server trust Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(core): update middleware manifest example Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): preserve unsigned gateway compatibility Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): reject cross-domain token replay Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
7547edc7ff |
docs(issues): Center reports on user stories (#2615)
Use the same compact persona, workflow, impact, reproduction, and environment prompts for bug reports and feature requests. Keep logs optional and specific to bug reports. Remove filing-time agent diagnostics so maintainers can evaluate user needs apart from investigation output, which becomes stale over time. Require contributors to investigate current behavior after humans accept the work, and treat state:accepted or roadmap placement as that signal. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
35fb27ef14 |
feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk) (#2122)
* feat(sdk): add TypeScript SDK (@nvidia/openshell-sdk) First native, per-language SDK for the OpenShell gateway: a thin, idiomatic TypeScript client over proto-generated gRPC stubs (connect-es), no FFI. Covers the v0.1 surface — sandbox lifecycle (create/get/list/delete + waitReady/ waitDeleted), health, and streamed exec. - sdk/typescript/: package, client/transport/errors, protoc + protoc-gen-es codegen (gen/ gitignored, absorbed into dist/ at build), committed lockfile. - tasks/typescript.toml: sdk:ts install/proto/typecheck/build/ci/publish; sdk:ts:typecheck wired into `check`; sdk-typescript job in branch-checks (typecheck, build, and a --dry-run publish that validates the release path). - Enforce SPDX headers on .ts/.tsx/.mts/.cts (skip node_modules and gen/); back-fill docs/_components/jsx.d.ts and fern/components/CustomFooter.tsx. - release.py gains an npm version format; release-tag.yml publishes to GitHub Packages on tag, stamping the version (0.0.0 placeholder in git); prerelease builds publish under the `next` dist-tag, not `latest`. Ships as @nvidia/openshell-sdk on GitHub Packages pre-GA; public npm (@openshell/sdk) follows at GA with an unchanged public API. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): adopt TypeScript 6, tidy @types/node range - typescript ^5.7.2 -> ^6.0.3 (6.0 is now `latest`; the old caret capped at 5.x) - @types/node ^24.0.0 -> ^24 (same range, tidier) No source changes; codegen, typecheck, and build pass on 6.0.3. Verified the emitted d.ts still type-check for downstream consumers on TypeScript 5.0.4 through 5.9.3, so this does not raise the SDK's consumer TS floor. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * refactor(sdk): group operations under a composable SandboxClient Reshape the client from flat methods (createSandbox, listSandboxes, exec) to a scoped SandboxClient reached as `client.sandbox.create/get/list/delete/exec` (+ waitReady/waitDeleted), mirroring the CLI's noun-verb model and the Python SDK's SandboxClient. SandboxClient is also usable standalone via SandboxClient.connect(); OpenShellClient composes it over a single shared transport, so future service/provider clients reuse one connection. health() stays top-level as a gateway call. No behavior change; types are unchanged. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): generate TypeScript SDK stubs with buf Replace the protoc gen.sh with `buf generate` + buf.gen.yaml. `buf` (@bufbuild/buf) is a package devDependency and self-compiles the protos, so the TS SDK no longer depends on the mise-pinned protoc; it drives the same connect-es plugin. Generation stays limited to the client-surface closure (openshell/sandbox/datamodel) via the input paths. Output is byte-identical to the previous protoc + protoc-gen-es pipeline. Lays the groundwork for a shared buf.yaml (lint/breaking/LSP) as a follow-up. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * build(proto): add repo-level buf module with lint Declare proto/ as a single buf v2 module in a root buf.yaml so buf generate, lint, breaking, and the editor LSP resolve imports the same way. Lint uses STANDARD with six documented exceptions for deviations the current protos intentionally make: the flat proto/ layout with nested packages (DIRECTORY_SAME_PACKAGE, PACKAGE_DIRECTORY_MATCH) and the established API shape with unsuffixed services and reused request/response messages (RPC_REQUEST_RESPONSE_UNIQUE, RPC_REQUEST_STANDARD_NAME, RPC_RESPONSE_STANDARD_NAME, SERVICE_SUFFIX). Every other STANDARD rule now enforces on future protos. Breaking uses FILE. Code generation stays package-scoped in sdk/typescript/buf.gen.yaml since it binds to that package's connect-es plugin and output dir; its inputs are unchanged and regeneration is byte-identical. Wire the check in via a proto:lint mise task that runs buf from the SDK devDependencies. It is a dependency of both sdk:ts:ci (so the TypeScript SDK CI job enforces it) and the top-level lint aggregate (so local pre-commit covers it). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): publish as unscoped openshell-sdk on public npm Rename the package from @nvidia/openshell-sdk to the unscoped openshell-sdk and target public npm (registry.npmjs.org) instead of GitHub Packages. GitHub Packages requires a scope matching the owning org, and the @openshell scope is blocked by an unrelated existing package, so an unscoped name on public npm is the lowest-friction distribution path and needs no org approval. Rework the release-tag publish job to auth against registry.npmjs.org with NPM_TOKEN (the job now only needs packages: read to pull the CI image). Update the README install instructions and usage imports. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk): publish @nvidia/openshell-sdk to GitHub Packages Revert the unscoped-name switch. GitHub Packages only accepts scoped names matching the owning org, so shipping there first (which needs no external npm org or NPM_TOKEN, just the repo's GITHUB_TOKEN) requires the @nvidia scope. Keeping the @nvidia/openshell-sdk name also lets a later public-npm release use the same install specifier, so adding public npm becomes a second publish step rather than a rename. Restore the GitHub Packages publish auth in the release-tag job and the scoped install instructions in the README (keeping the buf codegen note). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): add streaming exec, forward, ssh, provider, and config methods Grow SandboxClient to the surface the first two consumers need. execStream yields stdout/stderr chunks as they arrive and exec now drains it, keeping its buffered ExecResult and signature unchanged. execInteractive is the TTY + stdin transport primitive (start-first framing, output/write/resize/close/done, no terminal glue). forward binds a local TCP listener that tunnels each accepted connection into the sandbox for the process lifetime, minting and revoking a per-socket SSH session token around a forwardTcp bidi. Adds createSshSession / revokeSshSession, attach/detach/listProviders, and getConfig / setPolicy / setSetting (sandbox-scoped, network-policy-only, with an optional wait poll). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * build(sdk-ts): add Biome and Vitest tooling The TypeScript SDK had no formatter or linter and no test runner. Add Biome (format + lint, generated src/gen excluded) enforcing 2-space indent, single quotes, semicolons, and a 120-column width, and reformat the existing hand-written sources accordingly. Add Vitest for unit tests. Wire sdk:ts:format, sdk:ts:lint, and sdk:ts:test mise tasks into the fmt/lint aggregates, the root test suite, and sdk:ts:ci so they run in CI. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * test(sdk-ts): cover the sandbox surface with in-memory transport tests Exercise SandboxClient against an in-memory OpenShell service built with createRouterTransport: request assembly and id resolution, u64/int64 rendered as strings, enum lowercasing, fromConnect code mapping, the exec/execStream drain plus a backward-compat check on exec, execInteractive start-first ordering and done resolution, and a forward() byte relay against a loopback echo with close() teardown. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * docs(sdk-ts): document the new surface and connect/upload/download boundaries Document execStream, execInteractive, forward, ssh sessions, providers, and config/policy in the SDK README, and record the intentional boundaries: interactive connect / PTY ownership, upload/download (no file-transfer RPC), and detached forwards stay out of scope. Note the Biome/Vitest dev commands. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): support mTLS client authentication Add clientCert and clientKey to ConnectOptions so the SDK can authenticate to the default local gateway, which uses mTLS user authentication. Without a client certificate and key the SDK could verify the server but never authenticate the caller, so it could not connect to the standard Docker, VM, Homebrew, or Linux-package gateway. Validate the pair as both-or-neither and pass cert and key through to the Node TLS options for https gateways. The h2c path is unchanged. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * chore(sdk-ts): drop the demo script and its tsx dependency Remove src/demo.ts, the demo npm script, the tsx devDependency, and the tsconfig build exclude for the demo. The demo was never part of the published package, and dropping it also removes the only place that logged part of an SSH session token. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk-ts)!: harden exec streaming, waits, SSH, and forwarding Address review feedback on the sandbox surface. - Make the streamed command exit code observable from idiomatic for-await: the terminal exit is now an in-band ExecStreamEvent ({ type: 'exit', exitCode }) rather than the async generator return value, which for-await discards. A stream that ends without an exit event now throws instead of reporting success. - Bound waitReady, waitDeleted, and the setPolicy wait by their timeout: each poll RPC carries a per-iteration deadline and the waits accept an AbortSignal, so a stalled call can no longer leave a wait pending forever. Add waitTimeoutSecs to SetPolicyOptions. - Validate the CreateSshSession response against the proto charset and range contract before returning it or using its token, since the values feed an OpenSSH ProxyCommand. - Respect socket backpressure when relaying forwarded responses: pause reading the gRPC stream when the local socket buffer is full and resume on drain so memory stays bounded. - Expose create-time sandbox policy: add policy and an advanced rawSpec passthrough to SandboxSpec so the safety boundary is expressible at creation and new spec fields do not require an SDK change. BREAKING CHANGE: execStream and the interactive exec output now yield a terminal { type: 'exit', exitCode } event; consumers iterating the stream must handle that arm. The exit code is no longer the async generator return value. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): export error contract, enum unions, and caller cancellation Address Tier-1 review feedback on the TypeScript SDK public surface (PR #2122). - Errors: export SdkError and SdkErrorCode so callers can use instanceof and exhaustively switch on .code. fromConnect preserves the originating ConnectError as .cause and its status as .connectCode, maps Aborted to a new 'aborted' code for optimistic-concurrency conflicts, and maps Canceled and DeadlineExceeded to 'canceled'. errorCode() behavior is unchanged. - Enums: replace the string-typed phase, status, scope, and policySource fields with lowercase literal unions (SandboxPhaseName, HealthStatus, SettingScopeName, PolicySourceName) backed by exhaustive Record maps. The unions are a hand-maintained mirror of the generated proto enums; a new drift test pins each literal to its generated member name. - Cancellation: accept an optional AbortSignal on exec, execInteractive, and forward, threaded into both sandbox resolution and the streaming RPC. forward tears down its local listener on abort. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * feat(sdk-ts): add raw escape hatch for uncurated gateway RPCs The curated sub-clients reduce proto messages to ergonomic subsets (for example get() drops created_at_ms, the full spec, conditions, runtime endpoints, and current_policy_version), and not every gateway RPC has a typed helper yet. Rather than ship methods that exist but throw, expose a generated client for the full surface. OpenShellClient.raw and SandboxClient.raw are generated clients covering every gateway RPC, returning the verbatim wire messages so proto distinctions the curated types smooth over are preserved. .transport exposes the shared connection for building extra clients over one socket. Generated request/response types are published at the new @nvidia/openshell-sdk/raw subpath. Curated methods stay the default; raw is the always-available floor. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk-ts): address review feedback on exec, forward, and auth transport Settle exec's `done` promise before yielding the exit event so a consumer that breaks on exit no longer leaves it pending forever, and give it a lone rejection handler plus a finally-settle so a stream error or early abandon can never surface as an unhandled rejection or a hang. Attach an 'error' listener to each accepted forward socket synchronously, before forwardConnection awaits CreateSshSession; a peer reset in that window previously emitted an unhandled 'error' and crashed the process. Reject ambiguous or unsafe transport configs at buildTransport: oidcToken and edgeToken together (silently OIDC-only), and any auth token sent over plaintext http:// to a non-loopback host unless allowInsecureAuth is set. Wrap versionPin so a non-u64 expectedResourceVersion raises SdkError('invalid_config') instead of a raw BigInt SyntaxError, and raise the Node engine floor to >=20.3 for AbortSignal.any(). Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk-ts): address review feedback Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(sdk-ts): defer published sdk guide Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
c5f8366cd3 |
feat(sdk/go): add Go SDK foundation, types, and sandbox client (A) (#2271)
* feat(sdk/go): add Go SDK foundation, types, and sandbox client (A) Add the Go SDK module with the full API contract and a working sandbox client as the first vertical slice. All other resource clients are present as stubs returning Unimplemented errors, to be replaced with real implementations in subsequent PRs. Contents: - Module setup (go.mod, Makefile, mise.toml) - All domain types (types/ package) - Full ClientInterface with all sub-client accessors - Shared infrastructure (errors, auth, gRPC connection, logging) - Sandbox client with converter and tests (fully functional) - Stub clients for remaining resources (exec, file, health, provider, profile, config, refresh, policy, service, ssh, tcp) Part of the Go SDK decomposition plan (#2270). Implements #2044. * fix(sdk/go): address review feedback on PR #2271 - Make scheme parsing drive transport selection: http:// uses plaintext gRPC, https:// or no scheme uses TLS. Add regression tests. - Add Resources and DriverConfig fields to SandboxTemplate and update both converter directions (SandboxFromProto/SandboxSpecToProto). - Regenerate proto bindings from current canonical proto sources to eliminate drift (SigV4/MCP fields, params matchers, reserved fields). - Run gofmt/goimports on all handwritten Go files. Signed-off-by: Roland Huß <rhuss@redhat.com> * fix(sdk/go): address principal engineer review findings - Remove dead boolCount function that would fail golangci-lint (#1) - Emit EventAdded for the first watch event instead of EventModified, matching k8s watch semantics (#7) - Add mutex locking to all mock server methods that access the shared sandboxes map, fixing latent race conditions (#12) - Skip HealthCheck integration test that calls an unimplemented stub (#13) - Scope doc.go examples: mark sections for sub-clients not yet available in this PR with "available in a future release" (#4) - Document Config.Timeout/RetryPolicy/Logger and WatchOptions fields as reserved for future use (#2, #6) Signed-off-by: Roland Huß <rhuss@redhat.com> * refactor(sdk/go): migrate mise config to centralized task include Move Go SDK mise configuration from standalone sdk/go/mise.toml into the project's centralized pattern: - Add Go tools (go, golangci-lint, protoc-gen-go, protoc-gen-go-grpc) to root mise.toml [tools] section - Create tasks/go.toml with all SDK tasks using go: namespace prefix and dir=sdk/go for working directory - Update sdk/go/Makefile to reference namespaced task names - Update proto:sync default path for monorepo layout Addresses review feedback from drew on PR #2271 regarding mise convention alignment. Signed-off-by: Roland Huß <rhuss@redhat.com> * refactor(sdk/go): remove UPSTREAM_VERSION standalone repo artifact Remove sdk/go/proto/UPSTREAM_VERSION file and its exclusion from proto:check. This was a leftover from the standalone repo prototype. In a monorepo, proto drift is detectable via git diff between sdk/go/proto/ and proto/ directly. Signed-off-by: Roland Huß <rhuss@redhat.com> * refactor(sdk/go): switch proto generation from protoc to buf Replace raw protoc invocations with buf for Go SDK proto code generation, aligning with the TS SDK approach (PR #2122). - Add repo-level buf.yaml declaring proto/ as the buf module with lint and breaking change detection config - Add sdk/go/buf.gen.yaml configuring buf to generate Go code directly from root proto/ (no more vendored .proto copies) - Delete vendored .proto source files from sdk/go/proto/ - Rewrite go:proto:gen and go:proto:check mise tasks to use buf - Remove go:proto:sync and go:proto:clean tasks (no longer needed) - Add proto target to sdk/go/Makefile - Add buf 1.72.0 to root mise.toml tool dependencies - Include options.proto in generation (was stripped from vendored copies) - Regenerate all .pb.go files via the new buf pipeline Signed-off-by: Roland Huß <rhuss@redhat.com> * test(sdk/go): add proto-converter field coverage detection Use protobuf reflection to enumerate all fields on key proto messages (SandboxSpec, SandboxTemplate, SandboxStatus, SandboxCondition, SandboxPolicy) and compare against explicit handled/skipped sets in the converter tests. Unhandled fields produce warnings (t.Log), not failures, so proto contributors are not forced to fix SDK converters in the same PR. Stale entries in the handled set (removed proto fields) do fail, since they indicate the converter references something that no longer exists. A follow-up CI workflow will create GitHub issues when converter drift lands on main. Signed-off-by: Roland Huß <rhuss@redhat.com> * fix(sdk/go): bump Go to 1.26 and fix errcheck lint violations The upstream go.mod now has `toolchain go1.26.4`, which requires Go 1.26 to build golangci-lint. Bump the mise.toml Go version from 1.25 to 1.26 and wrap deferred Close() calls in test helpers to satisfy errcheck. Assisted-By: 🤖 Claude Code * feat(sdk/go): add ObjectMeta fields (annotations, workspace, deletion_timestamp) Add three new proto ObjectMeta fields to Sandbox and Provider domain types: Annotations (map), Workspace (string), and DeletionTimestamp (*time.Time). Update converters in both directions, deep-copy maps at the proto/SDK boundary, and add TimeFromMillisPtr/MillisFromTimePtr helper functions. Assisted-By: 🤖 Claude Code * chore(sdk/go): regenerate proto bindings after rebase Pick up workspace fields from upstream PR #2445 (Wire authorization into workspace model). All request messages now include workspace parameter in the generated Go bindings. Assisted-By: 🤖 Claude Code * feat(sdk/go): add workspace scoping to all RPC interfaces Add workspace parameter to every sandbox-scoped RPC method across all interfaces (Sandbox, Exec, File, Service, SSH, TCP, Config, Policy, Provider, Profile, Refresh). The workspace string is passed as the second parameter after ctx, following the convention workspace then resource-name. Key changes: - SandboxInterface: all 10 methods gain workspace parameter - sandbox_client.go: passes Workspace field in every proto request - ListOptions: add AllWorkspaces field for cross-workspace queries - All stub interfaces updated to match new signatures - All sandbox client tests updated with "default" workspace Assisted-By: 🤖 Claude Code * chore(sdk/go): remove coverage.out from tracking Assisted-By: 🤖 Claude Code * fix(sdk/go): address review feedback from mrunalp - Add RefreshStrategyAWSStsAssumeRole to match proto enum value 6, fulfilling the "all domain types upfront" contract - Wrap context.DeadlineExceeded and context.Canceled in StatusError so IsDeadlineExceeded() and IsCancelled() helpers work correctly - Return error from mapToStruct/SandboxSpecToProto instead of silently discarding structpb.NewStruct failures on invalid template maps Signed-off-by: Roland Huss <rhuss@redhat.com> * fix(sdk/go): address remaining review items - Wire go:ci into root ci task so SDK is tested in repository CI - Fix gofmt formatting on converter files - Add goimports to mise.toml tools - Add coverage.out to .gitignore - Add Go SDK section to AGENTS.md and CONTRIBUTING.md - Add regression tests for context-error wrapping (IsDeadlineExceeded, IsCancelled) and invalid template map rejection - Remove panic from SandboxToProto, return error instead Signed-off-by: Roland Huss <rhuss@redhat.com> * fix(sdk/go): pin goimports version and update lockfile Pin goimports to 0.48.0 instead of "latest" and regenerate mise.lock to include the new entry. Signed-off-by: Roland Huss <rhuss@redhat.com> * fix(sdk/go): TLS.Insecure means skip-verify, not plaintext Align TLS.Insecure semantics with the Rust SDK: Insecure: true now uses TLS with InsecureSkipVerify (skip cert verification) instead of switching to plaintext. Only the http:// scheme triggers plaintext. This fixes token auth against dev/k3d gateways: StaticToken and RefreshableToken require transport security, which real TLS (even with InsecureSkipVerify) satisfies, but plaintext does not. For http:// + token auth (dev gateways without TLS), wrap the auth provider to override RequireTransportSecurity, matching the Rust SDK's behavior where http:// accepts any auth mode. Transport decision table (matches Rust SDK crates/openshell-sdk): http:// + any TLS config -> plaintext (TLS config ignored) https:// + Insecure: true -> TLS, skip cert verify https:// + Insecure: false -> TLS, full verification no scheme -> same as https:// Signed-off-by: Roland Huss <rhuss@redhat.com> * feat(sdk/go): add missing policy proto fields Add 6 previously silently dropped fields to the network policy types and converters, preventing security-relevant data loss on round-trip: NetworkEndpoint fields 19-23: - CredentialSigning: SigV4 re-signing mode - SigningService: AWS service name for SigV4 - SigningRegion: AWS region override for SigV4 - JsonRpcMaxBodyBytes: JSON-RPC body inspection limit - Mcp: MCP-specific policy options (new McpOptions type) L7Allow and L7DenyRule field 9: - Params: MCP params matcher map for tools/call filtering New type McpOptions with StrictToolNames and AllowAllKnownMcpMethods optional booleans matching the proto definitions. Signed-off-by: Roland Huss <rhuss@redhat.com> * fix(sdk/go): enforce coverage test and extend to policy messages Change coverage_test.go from t.Logf (silent) to t.Errorf so that unhandled proto fields fail the test immediately. Add coverage tests for NetworkEndpoint (23 fields), L7Allow (8 fields), L7DenyRule (8 fields), and McpOptions (2 fields). Any new proto field that is not in the handled set or explicitly skipped now breaks the build, closing the silent-drift gap. Signed-off-by: Roland Huss <rhuss@redhat.com> * ci(sdk/go): add Go SDK job to branch-checks workflow Add a Go SDK job to branch-checks.yml that runs mise run go:ci (lint, build, test, proto-check, docs-check) on every PR. This ensures the SDK is tested in CI, not just locally. Signed-off-by: Roland Huss <rhuss@redhat.com> * fix(sdk/go): address should-fix review items #6 Fix broken godoc examples: add workspace parameter to all method calls in doc.go that were broken after workspace scoping. #7 Add Err field to Event[T]: Watch error events now carry the underlying error instead of discarding it. #8 Separate Unauthenticated from PermissionDenied: add ErrorUnauthenticated code and IsUnauthenticated() helper. gRPC Unauthenticated (401) now maps to its own code instead of collapsing into PermissionDenied (403). #9 Add Unwrap to StatusError: replace dead Details field with Cause error field. StatusError.Unwrap() returns Cause, enabling errors.Is/As unwrapping. FromGRPCError and contextError both populate Cause. Signed-off-by: Roland Huss <rhuss@redhat.com> * ci(sdk/go): add go:format:check to CI pipeline Add gofmt format verification to go:ci. Catches unformatted Go files before they reach the PR. Fix formatting on coverage_test.go. Signed-off-by: Roland Huss <rhuss@redhat.com> * chore(sdk/go): remove Makefile in favor of mise tasks All build, lint, test, and proto-gen tasks are already defined in tasks/go.toml and invoked via mise. The Makefile was a leftover that duplicated this and raised questions in review. Signed-off-by: Roland Huß <rhuss@redhat.com> * feat(sdk/go): sync proto bindings and add credential handle support Regenerate Go proto bindings after rebase to pick up new CredentialHandle message and Provider.credential_handles and profile_workspace fields from upstream. Add domain types, converter support, and proto field coverage tests for Provider and CredentialHandle. Signed-off-by: Roland Huß <rhuss@redhat.com> * fix(sdk/go): reject plaintext auth leak and fix watch error handling Reject http:// addresses when the auth provider requires transport security instead of silently stripping the requirement. Remove the insecureAuthWrapper that overrode RequireTransportSecurity. Fix watch stream error handling: use blocking send for terminal errors so they are never silently dropped when the channel is full, and wrap mid-stream errors with converter.FromGRPCError so SDK error helpers like IsUnavailable work on watch Event.Err. Signed-off-by: Roland Huß <rhuss@redhat.com> * fix(sdk/go): address review findings from multi-agent code review - WaitReady now detects SandboxDeleting phase and returns immediately instead of polling indefinitely - Watch goroutine defers streamCancel() to prevent context leaks - Fix StopOnTerminal=false test to keep stream open (was wrong-reason pass due to stream ending, not StopOnTerminal logic) - Add EventDeleted test covering the Deleting phase branch - Add provider converter unit tests for CredentialHandle round-trip, nil handling, and empty maps Signed-off-by: Roland Huß <rhuss@redhat.com> --------- Signed-off-by: Roland Huß <rhuss@redhat.com> Signed-off-by: Roland Huss <rhuss@redhat.com> |
||
|
|
5548405fcb |
feat(credentials): add provider credential storage drivers (#2437)
* feat(credentials): add provider credential storage drivers Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential update handling Signed-off-by: Taylor Mutch <taylormutch@gmail.com> * fix(credentials): harden credential driver security, correctness, and performance Address review findings from the credential storage drivers PR: - Route additional_credentials through the driver on refresh to prevent silent data loss for multi-credential providers (e.g. AWS STS) - Clean up stored credential handles on CAS failure during refresh to prevent orphaned secrets in external backends - Enforce namespace validation in the Kubernetes Secrets driver to prevent cross-namespace credential access when allow_reference_namespace is not enabled - Cache Vault Kubernetes auth tokens with 80% TTL to avoid re-authenticating on every credential operation - Parallelize resolve_credentials in all three drivers using try_join_all for faster sandbox startup - Add existingSecret support for the KEK Secret to fix helm template/GitOps workflows where lookup returns empty and regenerates the key - Document RBAC blast radius for the Kubernetes Secrets credential driver and recommend a dedicated namespace Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add optimistic concurrency, fix thundering herd, parallelize operations Use resourceVersion optimistic concurrency with retry loop for K8s Secret ownership checks to prevent TOCTOU races. Switch Vault token cache from RwLock to Mutex with double-check pattern to prevent thundering herd on cache miss. Parallelize credential store and delete operations across independent keys using try_join_all. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): handle partial failures, add delete retry, consolidate cleanup Replace try_join_all with join_all in credential store/delete operations to handle partial failures — successfully-stored handles are cleaned up when another key fails. Add retry loop with conflict detection to db-credstore delete_credential, matching the K8s driver pattern. Consolidate 4 manual cleanup_pre_stored_provider_credentials call sites into a single error handler using an async block. Remove inconsistent .trim() from db-credstore validate_handle_owner. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): fix retry loop guard and remove unprotected validation Remove attempt-count guard from 409/Aborted match arms in retry loops so the post-loop Status::aborted error is reachable after exhausting retries. Previously, last-attempt conflicts fell through to the catch-all error arm, producing misleading Status::unavailable errors. Remove duplicate validation calls that ran after prepare_provider_credential_update but outside the cleanup-protected async block, which would leak pre-stored handles on failure. Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> * fix(credentials): add workspace/provider UUID to credential backend paths Include workspace and provider ID in credential backend object paths to ensure cross-workspace uniqueness and prevent credential collision (GATOR-1806c9be-01). - Updated credential driver proto to include workspace and provider_id fields - Modified Vault driver to include workspace/provider_id in managed_secret_path - Modified Kubernetes Secrets driver to include workspace/provider_id in credential_owner_id and managed_secret_name - Updated all credential runtime calls to pass workspace/provider_id - Updated tests to use the new signatures This prevents two workspaces sharing the same external credential store from colliding on provider names, which was a critical security issue (CWE-639). * fix(credentials): preserve provider-level expiration for handle-backed credentials Compute effective expiration from both provider and driver values using the earliest non-zero timestamp and skip expired values before insertion (GATOR-1806c9be-02). - Modified resolve_provider_handles to check provider credential_expires_at_ms - Skip expired credentials during resolution instead of returning them - Use effective expiration (min of provider and driver) in resolution results - Fix inference.rs to preserve earliest expiration when merging This ensures handle-backed credentials respect the same expiration semantics as inline credentials. * fix(credentials): stage refresh changes under new handles before validation Stage credential replacements under new immutable handles instead of reusing existing handles to prevent overwriting committed values before validation/CAS (GATOR-1806c9be-03). - Stage credentials with empty existing_handles map to force new handle creation - Validate and CAS before the new values are committed to backend storage - Delete old handles only after successful CAS - On CAS failure, delete only the newly staged handles - This prevents CWE-362/CWE-367 race conditions where failed refreshes could still modify or delete the active credential The fix ensures that a rejected refresh cannot modify the backend object still referenced by the committed provider record. * fix(credentials): add timeouts to credential driver RPCs Apply configured timeouts to both startup capability negotiation and runtime RPCs to prevent indefinite hangs (GATOR-1806c9be-05). - Add DEFAULT_CREDENTIAL_DRIVER_RPC_TIMEOUT_SECS constant (30s) - Apply timeout to GetCapabilities during startup connection - Apply timeout to all runtime RPCs (store, delete, resolve) - Use tokio::time::timeout to bound the entire GetCapabilities operation during startup, not just the socket connection - Return contextual deadline errors on timeout This prevents a faulty or overloaded driver from hanging gateway operations indefinitely. * fix(credentials): fix test to use consistent workspace/provider identity The Kubernetes auth Vault resolve test was constructing a managed path with test-workspace/test-provider-id but sending default/prov-123 in the request, causing validation to reject the request (GATOR-18e32351-01). - Update test to use test-workspace and test-provider-id in the request to match the logical_path construction - This ensures the test exercises the intended code path and validates Kubernetes auth resolution properly The test now passes and correctly validates identity enforcement. * fix(credentials): use unique staging ID for refresh to avoid overwrites Stage refresh replacements under genuinely distinct immutable handles using a unique staging ID to prevent overwriting committed values (GATOR-1806c9be-03). - Generate a unique staging ID using UUID for each refresh operation - Use this staging ID when storing credentials instead of the real provider ID - Pass the same staging ID during cleanup on failure to delete only staged objects - This ensures deterministic paths (Vault) and object names (K8s) don't collide with the committed provider's credentials The fix prevents failed refreshes from silently replacing active credentials or breaking providers by deleting still-referenced backend objects. * fix(credentials): wrap credential driver RPCs in local timeouts Add local tokio::time::timeout wrappers around credential driver RPCs to bound non-compliant or stalled UDS peers (GATOR-1806c9be-05). - Wrap StoreCredential, DeleteCredential, and ResolveCredentials in local timeouts - Return contextual deadline_exceeded errors when timeouts occur - Keep existing gRPC timeout metadata for compliant implementations - GetCapabilities during startup was already wrapped in previous commit This ensures a faulty local driver cannot hang gateway operations indefinitely, even if it accepts the connection but never responds to the RPC. * fix(credentials): preserve ownership for staged refreshes Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(credentials): bound startup capability probe Signed-off-by: Seth Jennings <sjenning@redhat.com> * test(provider): authenticate credential handler requests Signed-off-by: Seth Jennings <sjenning@redhat.com> * fix(ci): grant actions read to credential driver e2e Signed-off-by: Seth Jennings <sjenning@redhat.com> --------- Signed-off-by: Taylor Mutch <taylormutch@gmail.com> Signed-off-by: Varsha Prasad <varshaprasad96@gmail.com> Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> Signed-off-by: Seth Jennings <sjenning@redhat.com> Co-authored-by: Taylor Mutch <taylormutch@gmail.com> Co-authored-by: Varsha Prasad Narsing <varshaprasad96@gmail.com> |
||
|
|
8c7dd148a9 |
perf(net): set TCP_NODELAY on latency-sensitive TCP hops (#2220)
* perf(net): set TCP_NODELAY on all tunnel and proxy TCP hops The sandbox tunnel added ~44 ms of latency to every small request/response because no socket in the path disabled Nagle's algorithm, so sub-MSS writes waited on delayed ACKs at each hop. Set TCP_NODELAY on every latency-sensitive TCP socket: - gateway: accepted connections on the public listener (gRPC relay frames and WS tunnel writes) - CLI: edge tunnel local accept + underlying WebSocket TCP stream, insecure TLS connector (tonic's default connector already does this), and service-forward accepted sockets - supervisor: direct-tcpip connect into the sandbox netns, TCP relay target dials, egress proxy accepted connections, and all upstream CONNECT/HTTP dials (via a new connect_upstream helper) Setting TCP_NODELAY on connect is best-effort: a failure only costs latency, so we log and continue rather than fail the connection. The sandbox SSH transport rides a unix domain socket and gRPC client channels use tonic defaults (nodelay on), so no change is needed there. Fixes #2219 Signed-off-by: Jim Meyer <jim@meyer4hire.com> * refactor(net): house TCP_NODELAY helper in a shared net module Address review feedback on the TCP_NODELAY change: - Move the shared set-nodelay helper out of supervisor_session into a new crate-private `net` module in openshell-supervisor-process, so ssh and supervisor_session no longer reach across modules through a pub(crate) item. - Make the best-effort comments at each call site terse and consistent. No behavior change; the benchmark ladder reproduces the same numbers. Signed-off-by: Jim Meyer <jim@meyer4hire.com> * refactor(net): consolidate TCP_NODELAY helpers into openshell_core::net Move the best-effort TCP_NODELAY helpers into the shared openshell_core::net module so every crate dials and configures sockets the same way: - Add set_tcp_nodelay_best_effort (accepted/existing streams) and connect_tcp_nodelay_best_effort (dial + set) with unit tests. - Migrate all call sites in openshell-cli, openshell-server, and the supervisor crates to the shared helpers. - Remove the crate-private net module from openshell-supervisor-process. - Document socket guidance in AGENTS.md (Network Sockets). Signed-off-by: Jim Meyer <jim@meyer4hire.com> * perf(net): set TCP_NODELAY on exec bridge and metadata server The gateway-side single-use SSH-over-relay loopback bridge and the sandbox IMDS metadata server were missed latency-sensitive TCP hops. Set TCP_NODELAY on the accepted client connection and both russh client dials of the exec bridge — interactive keystrokes and line-buffered PTY output are the most tinygram-heavy traffic in the system — and on the metadata server's accepted connections. Also log unrecognized MaybeTlsStream variants in the edge tunnel so a future TLS-backend change surfaces a silent TCP_NODELAY miss instead of skipping it quietly. Signed-off-by: Jim Meyer <jim@meyer4hire.com> * perf(net): set TCP_NODELAY on openshell-sdk socket paths The openshell-sdk crate landed on main with its own copies of the CLI's hand-rolled sockets, which the CLI and TUI are meant to consume. Give them the same treatment as the CLI equivalents: - edge_tunnel: the accepted local tunnel connection and the WebSocket's underlying TCP socket (plain and rustls variants). - transport: the dial in InsecureTlsConnector, tonic's custom-connector path. Only these hand-rolled sockets need it. Tonic's own connector defaults tcp_nodelay to true and applies it itself, so plain Endpoint::connect callers were already covered. Signed-off-by: Jim Meyer <jim@meyer4hire.com> --------- Signed-off-by: Jim Meyer <jim@meyer4hire.com> |
||
|
|
1a25439cab |
refactor(otel): share OTLP trace provider setup (#2567)
Extract common OpenTelemetry provider construction into openshell-otel so OpenShell services share one OTLP/gRPC export implementation. The shared crate owns: - typed exporter setup errors and non-fatal provider enablement; - endpoint trimming and URI validation before lazy exporter connection; - fixed and environment-or-default service-name policies; - service version and caller-supplied resource attributes; - batch tracer-provider construction; and - span-only tracing layers that exclude OpenTelemetry exporter callsites. Migrate the gateway to the shared provider while retaining its configurable service name, error marking, and tracing test collector. Add the shared crate to the architecture inventory and document the tracing boundary. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
02e890ccf1 |
docs: document issue lifecycle labels and roadmap sequencing (#2524)
Add a published Issue Triage and Lifecycle page and align AGENTS.md, CONTRIBUTING.md, README.md, the PR template, and the issue-handling skills on the state:*/agent:* label model. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
dae926160b |
docs(agents): keep project skills synchronized (#2349)
* docs(agents): refresh CLI and debugging skills Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(agents): keep project skills synchronized Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(agents): cover middleware workflows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
d556748771 |
feat(supervisor-middleware): add network egress middleware (#2027)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
e6f319c76d |
feat(sdk): add openshell-sdk crate (#1862)
* feat(sdk): add openshell-sdk crate Additive extraction of the shared async gRPC client core (transport, TLS, OIDC single-flight refresh, edge tunnel, high-level sandbox surface, raw escape hatch) as a new workspace crate. No existing consumers yet; CLI/TUI migration follows in a separate PR. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * refactor(sdk,core): reuse shared JWT exp decoder for refresh deadlines Extract the signature-unverified JWT exp decode out of openshell-core/grpc_client.rs into openshell_core::jwt::parse_exp_secs, and have the openshell-sdk refresh path reuse it to derive a proactive refresh deadline from a bearer JWT when the caller does not advertise expires_at. Addresses review feedback to reuse pre-existing logic rather than reimplement JWT expiry handling per client. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): harden refresh single-flight and redact tokens in Debug Addresses review feedback on the openshell-sdk refresh path. - Make the single-flight cleanup cancellation-safe. The in-flight slot is now cleared by the shared refresh computation itself (epoch-guarded) rather than the leader's post-await code. Previously, if the leader future was dropped (e.g. an FFI caller cancelling its promise) after a follower drove the refresh to completion, the completed future was stranded in the slot and later refresh_now() calls re-joined it, pinning the client to a stale or already-rejected token. Adds a regression test that cancels the leader and asserts the next refresh starts a fresh attempt. - Redact bearer secrets from Debug. RefreshedToken and the oidc RefreshTokenInput/RefreshTokenOutput now use manual Debug impls that omit the access/refresh token fields via finish_non_exhaustive, matching the house style (e.g. SecretResolver, SandboxJwtIssuer). Prevents a stray {:?} or a containing struct's derived Debug from writing tokens to logs. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): fail refresh when the new token can't be encoded as metadata store_bearer now returns an error instead of silently keeping the previous bearer value. The TokenSource commits the refreshed token to its state before the client writes it into the interceptor slot, so a silent drop left the interceptor on the old (expiring) token with no path back to a refresh. Surfacing the error fails the call loudly instead. Adds a unit test covering a token that can't be encoded as gRPC metadata. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): refresh OIDC tokens on raw routes and harden rotation Raw gRPC access never triggered OIDC refresh: a client that only used raw_grpc/raw_inference kept sending the initial bearer until it expired, with no proactive or reactive refresh. Add raw_grpc_fresh and raw_inference_fresh accessors that refresh before returning the client, plus force_refresh for reactive recovery after an Unauthenticated raw RPC. Guard the single-flight refresh commit against a concurrent replace(). The in-flight attempt now records the generation it started from and skips its write when an external replace() has advanced it, so timer or callback driven rotation is no longer clobbered by a slower refresh. Remove TokenSource::snapshot(): it returned an empty string under write contention and had no consumer on the CLI/TUI path. Tests read committed state directly instead. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * docs(sdk): rewrite crate README for consumers Recast the openshell-sdk README as a usable crate README rather than an RFC excerpt. Drop the Responsibilities/Non-responsibilities/Consumers scope-boundary sections and the mTLS migration rationale, folding the useful facts (explicit token, no disk/name resolution, Refresh trait, SdkError mapping) into the intro, a new Auth and refresh section, and Public surface. Remove the dead relative RFC link and status-label prose so the doc renders cleanly wherever it is published. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): address review feedback on OIDC refresh and transport Apply Drew's review notes on PR #1862: - Drop unused `rustls-pemfile` dependency and move `tokio-stream` to dev-dependencies (only used by tests). - Guard OIDC `expires_at` against u64 overflow with `saturating_add`. - Fix stale `#[non_exhaustive]` rationale in `AuthConfig` (the struct `Oidc` variant it described as future already ships). - Stop double-wrapping refresh errors: store the bare refresh-error text so the single `SdkError::auth` wrap happens once at await. - Strip stale CLI porting breadcrumbs from `build_channel` docs, keeping the branch table. - Treat proactive token refresh as best-effort: a transient failure falls through to the request instead of failing an RPC whose current token is still valid, with a regression test. - Collapse `exec`'s inline auth retry into the shared `unary` helper. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> * fix(sdk): preserve transient/terminal distinction in refresh errors The refresh single-flight collapsed both `RefreshError::Transient` and `RefreshError::Terminal` into a stringified `SdkError::Auth`, so consumers (CLI, TUI, future language bindings) had no machine-readable way to tell a retryable IdP blip from a dead session that needs re-authentication. Carry the `RefreshError` through the shared outcome (kept `Clone` for `Shared`) instead of its rendered text, and map it at the await site to a new `retryable` flag on `SdkError::Auth`. Add `SdkError::auth_retryable` and a `SdkError::retryable()` accessor; transient refresh failures report `true`, every other error `false`. Add classification tests. Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> --------- Signed-off-by: Max Dubrinsky <mdubrinsky@nvidia.com> |
||
|
|
5ca39b0490 |
docs(rfc): require issues before RFCs (#1918)
* docs(rfc): require issues before RFCs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): correct RFC statuses Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): mark template accepted Signed-off-by: Drew Newberry <anewberry@nvidia.com> * Update rfc/README.md Co-authored-by: krishicks <khicks@nvidia.com> * docs(rfc): document accepted RFC project tracking Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): reframe RFC discussion guidance Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): label issues when assigning RFCs Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): link labeled RFC issues to board Signed-off-by: Drew Newberry <anewberry@nvidia.com> * docs(rfc): clarify RFC labels and board tracking Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: krishicks <khicks@nvidia.com> |
||
|
|
c5ce3ed623 |
AGENTS.md: Add more detailed signoff guidance (#1852)
* AGENTS.md: Add more detailed signoff guidance My coding agent started hallucinating my email address to be `rbryant@nvidia.com`. Add more guidance in `AGENTS.md` about how the `Signed-off-by` header should be added so that it uses the email address from git configuration instead of one it makes up. Signed-off-by: Russell Bryant <rbryant@redhat.com> * Update AGENTS.md Co-authored-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
f236279669 |
docs: document DCO commit sign-off requirement (#1811)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
188b355033 | docs(config): update gateway config reference (#1624) | ||
|
|
028763d4db |
refactor(vm): remove legacy openshell-vm crate (#1239)
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
70a0f6c547 | refactor(cli): remove gateway lifecycle management (#1221) | ||
|
|
cc2114e267 | docs(architecture): reset subsystem docs (#1184) | ||
|
|
f56c09c7df | docs: update gateway deployment architecture (#1108) | ||
|
|
5975805424 | feat(server): add bundled docker compute driver (#888) | ||
|
|
e4d6f92d9b | feat(vm): add standalone libkrun compute driver (#858) | ||
|
|
8b15ef772e |
docs(fern): move published docs into docs tree (#796)
Remove the legacy Sphinx pipeline and make docs/ the single source of truth so the published site matches the repository layout. |
||
|
|
f92923e06b | docs(fern): finalize preview workflow and nav cleanup (#784) | ||
|
|
ddb85b1704 | feat(vm): add openshell-vm crate with libkrun microVM gateway (#611) | ||
|
|
b7779bdefa |
feat(sandbox): integrate OCSF structured logging for sandbox events (#720)
* feat(sandbox): integrate OCSF structured logging for all sandbox events WIP: Replace ad-hoc tracing calls with OCSF event builders across all sandbox subsystems (network, SSH, process, filesystem, config, lifecycle). - Register ocsf_logging_enabled setting (defaults false) - Replace stdout/file fmt layers with OcsfShorthandLayer - Add conditional OcsfJsonlLayer for /var/log/openshell-ocsf.log - Update LogPushLayer to extract OCSF shorthand for gRPC push - Migrate ~106 log sites to OCSF builders (NetworkActivity, HttpActivity, SshActivity, ProcessActivity, DetectionFinding, ConfigStateChange, AppLifecycle) - Add openshell-ocsf to all Docker build contexts * fix(scripts): attach provider to all smoke test phases to avoid rate limits GitHub's unauthenticated API rate limit (60/hour) causes flaky 403s for Phases 1, 2, and 4. Fix by attaching the provider to all sandboxes and upgrading the Phase 1 policy to L7 so credential injection works. Phase 4 (tls:skip) cannot inject credentials by design, so relax the assertion to accept either 200 or 403 from upstream -- both prove the proxy forwarded the request. * fix(ocsf): remove timestamp from shorthand format to avoid double-timestamp The display layer (gateway logs, TUI, sandbox logs CLI) already prepends a timestamp. Having one in the shorthand output too produces redundant double-timestamps like: 15:49:11 sandbox INFO 15:49:11.649 I NET:OPEN ALLOWED ... Now the shorthand is just the severity + structured content: 15:49:11 sandbox INFO I NET:OPEN ALLOWED ... * refactor(ocsf): replace single-char severity with bracketed labels Replace cryptic single-character severity codes (I/L/M/H/C/F) with readable bracketed labels: [LOW], [MED], [HIGH], [CRIT], [FATAL]. Informational severity (the happy-path default) is omitted entirely to keep normal log output clean and avoid redundancy with the tracing-level INFO that the display layer already provides. Before: sandbox INFO I NET:OPEN ALLOWED ... After: sandbox INFO NET:OPEN ALLOWED ... Before: sandbox INFO M NET:OPEN DENIED ... After: sandbox INFO [MED] NET:OPEN DENIED ... * feat(sandbox): use OCSF level label for structured events in log push Set the level field to 'OCSF' instead of 'INFO' for OCSF events in the gRPC log push. This visually distinguishes structured OCSF events from plain tracing output in the TUI and CLI sandbox logs: sandbox OCSF NET:OPEN [INFO] ALLOWED python3(42) -> api.example.com:443 sandbox OCSF NET:OPEN [MED] DENIED python3(42) -> blocked.com:443 sandbox INFO Fetching sandbox policy via gRPC * fix(sandbox): convert new Landlock path-skip warning to OCSF PR #677 added a warn!() for inaccessible Landlock paths in best-effort mode. Convert to ConfigStateChangeBuilder with degraded state so it flows through the OCSF shorthand format consistently. * fix(sandbox): use rolling appender for OCSF JSONL file Match the main openshell.log rotation mechanics (daily, 3 files max) instead of a single unbounded append-only file. Prevents disk exhaustion when ocsf_logging_enabled is left on in long-running sandboxes. * fix(sandbox): address reviewer warnings for OCSF integration W1: Remove redundant 'OCSF' prefix from shorthand file layer — the class name (NET:OPEN, HTTP:GET) already identifies structured events and the LogPushLayer separately sets the level field. W2: Log a debug message when OCSF_CTX.set() is called a second time instead of silently discarding via let _. W3: Document the boundary between OCSF-migrated events and intentionally plain tracing calls (DEBUG/TRACE, transient, internal plumbing). W4: Migrate remaining iptables LOG rule failure warnings in netns.rs (IPv4 TCP/UDP, IPv6 TCP/UDP) to ConfigStateChangeBuilder for consistency with the IPv4 bypass rule failure already migrated. W5: Migrate malformed inference request warn to NetworkActivity with ActivityId::Refuse and SeverityId::Medium. W6: Use Medium severity for L7 deny decisions (both CONNECT tunnel and FORWARD proxy paths) to match the CONNECT deny severity pattern. Allows and audits remain Informational. * refactor(sandbox): rename ocsf_logging_enabled to ocsf_json_enabled The shorthand logs are already OCSF-structured events. The setting specifically controls the JSONL file export, so the name should reflect that: ocsf_json_enabled. * fix(ocsf): add timestamps to shorthand file layer output The OcsfShorthandLayer writes directly to the log file with no outer display layer to supply timestamps. Add a UTC timestamp prefix to every line so the file output matches what tracing::fmt used to provide. Before: CONFIG:VALIDATED [INFO] Validated 'sandbox' user exists in image After: 2026-04-01T15:49:11.649Z CONFIG:VALIDATED [INFO] Validated ... * fix(docker): touch openshell-ocsf source to invalidate cargo cache The supervisor-workspace stage touches sandbox and core sources to force recompilation over the rust-deps dummy stubs, but openshell-ocsf was missing. This caused the Docker cargo cache to use stale ocsf objects from the deps stage, preventing changes to the ocsf crate (like the timestamp fix) from appearing in the final binary. Also adds a shorthand layer test verifying timestamp output, and drafts the observability docs section. * fix(ocsf): add OCSF level prefix to file layer shorthand output Without a level prefix, OCSF events in the log file have no visual anchor at the position where standard tracing lines show INFO/WARN. This makes scanning the file harder since the eye has nothing consistent to lock onto after the timestamp. Before: 2026-04-01T04:04:13.065Z CONFIG:DISCOVERY [INFO] ... After: 2026-04-01T04:04:13.065Z OCSF CONFIG:DISCOVERY [INFO] ... * fix(ocsf): clean up shorthand formatting for listen and SSH events - Fix double space in NET:LISTEN, SSH:LISTEN, and other events where action is empty (e.g., 'NET:LISTEN [INFO] 10.200.0.1' -> 'NET:LISTEN [INFO] 10.200.0.1') - Add listen address to SSH:LISTEN event (was empty) - Downgrade SSH handshake intermediate steps (reading preface, verifying) from OCSF events to debug!() traces. Only the final verdict (accepted/denied) is an OCSF event now, reducing noise from 3 events to 1 per SSH connection. - Apply same spacing fix to HTTP shorthand for consistency. * docs(observability): update examples with OCSF prefix and formatting fixes Align doc examples with the deployed output: - Add OCSF level prefix to all shorthand examples in the log file - Show mixed OCSF + standard tracing in the file format section - Update listen events (no double space, SSH includes address) - Show one SSH:OPEN per connection instead of three - Update grep patterns to use 'OCSF NET:' etc. * docs(agents): add OCSF logging guidance to AGENTS.md Add a Sandbox Logging (OCSF) section to AGENTS.md so agents have in-context guidance for deciding whether new log emissions should use OCSF structured logging or plain tracing. Covers event class selection, severity guidelines, builder API usage, dual-emit pattern for security findings, and the no-secrets rule. Also adds openshell-ocsf to the Architecture Overview table. * fix: remove workflow files accidentally included during rebase These files were already merged to main in separate PRs. They got pulled into our branch during rebase conflict resolution for the deleted docs-preview-pr.yml file. * docs(observability): use sandbox connect instead of raw SSH Users access sandboxes via 'openshell sandbox connect', not direct SSH. * fix(docs): correct settings CLI syntax in OCSF JSON export page The settings CLI requires --key and --value named flags, not positional arguments. Also fix the per-sandbox form: the sandbox name is a positional argument, not a --sandbox flag. * fix(e2e): update log assertions for OCSF shorthand format The E2E tests asserted on the old tracing::fmt key=value format (action=allow, l7_decision=audit, FORWARD, L7_REQUEST, always-blocked). Update to match the new OCSF shorthand (ALLOWED/DENIED, HTTP:, NET:, engine:ssrf, policy:). * feat(sandbox): convert WebSocket upgrade log calls to OCSF PR #718 added two log calls for WebSocket upgrade handling: - 101 Switching Protocols info → NetworkActivity with Upgrade activity. This is a significant state change (L7 enforcement drops to raw relay). - Unsolicited 101 without client Upgrade header → DetectionFinding with High severity. A non-compliant upstream sending 101 without a client Upgrade request could be attempting to bypass L7 inspection. |
||
|
|
a912848217 | refactor(build): unify image build graph for cache reuse (#390) | ||
|
|
e45d415230 | chore(repo): migrate github label taxonomy (#454) | ||
|
|
a4e2c91008 |
chore: add vouch system for first-time contributors (#375)
chore: add vouch system for first-time contributors |
||
|
|
0792dcb425 | docs: unify install command in landing page, change docs skill name, update contributing guides (#355) | ||
|
|
7746c77c35 |
chore: add docs contributing guides and skills (#301)
* add docs contributing guides and skills * add section |
||
|
|
6ddb6aef83 |
chore: establish agent-first development ethos across project (#293)
* docs: enhance AGENTS.md as primary agent entrypoint Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: add opencode agent definitions for sub-agent workflows Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs: overhaul CONTRIBUTING.md for agent-first workflow Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * docs: reframe README.md with agent-first identity Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: add bug report and feature request issue templates Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: add issue template config with agent-first contact links Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: add pull request template Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: add CODEOWNERS Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: update .claude/ README to document skill symlink and consolidation Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * chore: remove github-rest-api skill Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat: add triage-issue skill for community issue assessment Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor: update create-github-issue skill for template conformance Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor: update create-github-pr skill for template conformance Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * refactor: update build-from-issue skill for template conformance and triage awareness Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * ci: add issue triage gate workflow Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * feat: add sync-agent-infra skill to detect and fix infrastructure drift Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
fbd93a4632 | refactor: rename navigator- crate prefix to openshell- (#277) |