mirror of
https://github.com/NVIDIA/OpenShell.git
synced 2026-10-03 16:11:17 +08:00
windows
209
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8f22fe84f6 |
ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forwar… (#3826)
* ci(windows): run MXC host probe, WebSocket agent, and OpenClaw forward examples checks Add separate hosted CI tasks for the MXC host probe, WebSocket agent, and OpenClaw forward examples. Use mock workloads to verify gateway, CLI, driver, and sandbox lifecycle wiring without requiring wxc-exec. Document that mock passes do not validate forwarding or MXC enforcement. * fix(tests): enhance environment isolation for WebSocket and OpenClaw mock tests * fix(mxc): enhance OpenClaw mock validation to require 'Ready' sandbox state * fix(mxc): restore native process helper in WebSocket example Restore Invoke-NativeCaptured for the PowerShell argument regression test and delegate CLI execution through it while preserving explicit gateway endpoint selection. Signed-off-by: Akber Raza <akberr@nvidia.com> --------- Signed-off-by: Akber Raza <akberr@nvidia.com> |
||
|
|
d46a814141 |
ci(windows): exercise MXC credential, audit, and aggregate E2E flows (#3787)
* ci(windows): exercise MXC provider credential example Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * ci(windows): exercise MXC OCSF audit example Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * ci(windows): enable aggregate MXC example E2E Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * ci(windows): select native MXC mock target Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * fix(windows): isolate MXC example CI harnesses Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> |
||
|
|
7ba7a39d09 |
ci(windows): exercise MXC inference demos with mock API (#3780)
* ci(windows): exercise MXC Ollama demo with mock API Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * ci(windows): cover both MXC inference demos Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * fix(mxc): preserve executable extension resolution Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(mxc): use absolute PowerShell in lifecycle checks Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(mxc): make lifecycle write probes deterministic Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(mxc): assert stable lifecycle completion Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> |
||
|
|
651ed7e03d |
NVBug 6783374: make MXC HTTPS L7 qualification authoritative (#3479)
* test(mxc): qualify HTTPS L7 enforcement (NVBug 6783374) Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * docs(mxc): clarify real qualification failures Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * docs(mxc): align real test skip semantics Signed-off-by: Shailendra Singh <shailendras@nvidia.com> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> |
||
|
|
a69d0319f2 |
test(windows): define GB300 MXC qualification contract (NVBug 6643699) (#3471)
* test(windows): define GB300 MXC qualification contract Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(mxc): harden GB300 qualification provenance Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(windows): package portable GB300 qualification Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(windows): honor external qualification checkout Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * revert: remove portable GB300 packaging Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * test(mxc): bind GB300 evidence to exact inputs Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> |
||
|
|
fb2980e077 |
fix(windows): restore MXC qualification and cold-start readiness (#3468)
* fix(windows): restore MXC qualification gates Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> * fix(mxc): poll target for full readiness budget --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> |
||
|
|
49b4f0eb7f |
feat(mxc): add UI policy, credentials, relay lifecycle, and proxy auth
Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
b799fccb8b |
fix(auth): harden OIDC trust root retrieval (#3332)
* fix(auth): harden OIDC trust root retrieval Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(e2e): pass OIDC HTTP acknowledgement value Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
26f2f96393 |
feat(mxc): add Windows host proxy for MXC sandbox network egress (#3163)
* Implement Windows host proxy integration and update dependencies for OpenShell
* Update README and gateway config to clarify egress proxy address handling and allocation
* Refactor ProxyIdentityMode to return Result for static_binary and add tests for binary path and SHA256 hash
* Enhance platform_hosts_path for Windows to use SystemRoot and improve error handling for hosts file reading
* Refactor FileFingerprint to use Option for mtime and ctime, simplifying metadata handling
* Add conditional compilation for Windows host module
* add unit tests for OPA policy evaluation and identity handling
* remove openshell-supervisor-network from unsupported driver package test exclusion list
* feat(mxc): enable host proxy TLS state generation
Generate per-sandbox TLS state for the MXC host proxy so HTTPS L7 enforcement can use the same MITM path as Linux. Grant generated CA material to the MXC process and inject standard trust env vars, while matching Linux behavior by disabling TLS termination on CA setup failure and relying on proxy fail-closed handling.
* fix(docs): remove outdated notes on governed egress from docs
* fix(tests): update TLS environment variable paths to use temporary directory
* fix(examples): make run-mxc-e2e harness correct and orphan-free
The MXC e2e harness never actually exercised the fs scenarios: it started
the gateway once and patched agent_command per scenario AFTERWARDS, so the
running gateway kept launching the default demo agent (not shipped in the
kit) and every fs scenario failed with CreateProcessW error:2. It also
scored on the `sandbox create` exit code (non-zero due to the harmless
interactive attach), wrote sandbox records to the persistent gateway DB
(leaving orphans that collided on later runs), and its deny scenarios never
proved denial.
Changes:
- Start a FRESH gateway per scenario so each scenario's agent_command is
actually loaded (root cause of CreateProcessW error:2).
- Score by on-disk artifact / expected outcome, not `sandbox create` exit.
- Real deny assertions: a control write to a granted path must succeed
(proves the agent ran) while the denied write must be absent. fs-empty
probes an ungranted out-of-share path (share_dir is mapped rw by design).
- Run the gateway on an ephemeral in-memory DB (sqlite::memory:) so the
harness never writes to the persistent store and cannot leave orphan
sandbox records; also use unique per-run sandbox names + pre-delete.
- Fix the process_container probe: use a real cwd + absolute cmd.exe
(canonical wxc-exec does not expand %TEMP% -> 0x8007010B).
- Fix summary counts (@() so a single FAIL is counted and exit is non-zero).
Verified PASS=4 FAIL=0 on 7F203-MXC-003 (no BaseContainer velocity keys)
using a canonical wxc-exec build (AppContainer fallback).
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): probe timeout is milliseconds (10ms->30000ms)
MXC process.timeout is wall-clock ms (wire.rs). The 10 value meant 10ms,
which the base-container tier (7F203-MXC-001/.181) enforced strictly and
timed the probe out. AppContainer path (.18/-003) happened to slip under
it. Bump to 30000ms so the process_container preflight probe is reliable
across both tiers.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): use native paths in real runtime probes
Signed-off-by: Shailendra Singh <shailendras@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc): make processcontainer work with mxc-latest-released wxc-exec
Three fixes to support the release wxc-exec binary (BaseContainer dispatcher)
in addition to mxc-fixes-env-vars:
1. Seed process env from host (driver.rs)
ProcessContainer starts with a completely blank environment -- no PATH,
SystemRoot, or anything. Seed the process env from the gateway host
environment so the agent binary can locate DLLs and run. Skip internal
Windows drive-letter variables (keys starting with '=') which cause
CreateProcessW to return ERROR_ENVVAR_NOT_FOUND. User agent_env entries
and TLS CA vars are applied as overrides on top of the host env.
2. Remove TLS readonly_paths grant (driver.rs)
The release wxc-exec (BaseContainer dispatcher) requires write-DAC
permission on every path in readonly_paths to set up AppContainer ACLs.
Adding the proxy's temp TLS directory caused a DACL error and exit -1.
The CA cert paths remain available to the agent via TLS env vars.
3. Remove allowedHosts from network JSON (mxc.rs)
The release wxc-exec rejects network.allowedHosts / network.blockedHosts
on Windows with "not yet supported". Removed the loopback exemption
attempt (127.0.0.1, ::1, localhost) from the network section.
Intra-container loopback works natively in the release binary without
it -- the spawner can connect to the server at 127.0.0.1:22000 directly.
Additional changes:
- mxc-ws-agent.rs: add relay-debug.txt error capture and relay-ready.txt
marker for reliable timing of host client connections.
- mxc-ws-gateway.toml: debug = true for JSON config dump during diagnosis.
- run-ws-agent-test.ps1: default port changed to 17670 (gateway default);
relay-ready.txt polling before ws-echo to avoid connecting before the
spawner has established the proxy bridge.
Signed-off-by: Prashant Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(e2e): address CodeRabbit review on run-mxc-e2e.ps1 (MR !46)
Four robustness/correctness fixes from CodeRabbit:
1. Start-Gw: kill the spawned gateway before the "did not start within 30s"
throw. If the process is alive but never binds the port, $gw is not yet
assigned in the caller, so the finally block cannot reap it -> orphan
gateway holding the port for the next run.
2. create-fail scoring: a non-zero `sandbox create` exit alone is not proof
of a policy rejection (gateway-registration/transport/fixture errors also
exit non-zero and would false-pass). PASS now requires a genuine
rejection signal (network / invalid_argument / network_policies) AND that
it is not an infrastructure failure; other non-zero exits go to FAIL with
output captured.
3. deny scenarios (ControlTarget path): snapshot the deny target AFTER
Wait-File lands the control artifact, so a late denied write (enforcement
regression racing the control write) can no longer be recorded as PASS.
4. -KeepRunning: break out of the scenario loop after the first scenario so
a later scenario does not start a second gateway on the same port
(previously a reliable port collision instead of a usable debug mode).
Re-verified PASS=4 FAIL=0 on both boxes (7F203-MXC-001 base-container and
7F203-MXC-003 AppContainer fallback); network-policy-rejected correctly
scores as "policy rejection".
Signed-off-by: Akber Raza <akberr@nvidia.com>
* feat(mxc-e2e): collect run-mxc-e2e output into a results bundle
Mirror the sibling run-*.ps1 scripts by collecting every run's logs into a
timestamped results-e2e-<stamp>\ folder and zipping it. The bundle contains the
console transcript, per-scenario gateway stdout/stderr, the exact TOML rendered
for each scenario, the policy fixture used, and a summary.txt with the verdict
table.
Per-scenario gateway logs now land in gateway.<scenario>.log/.err.log inside the
bundle instead of a single fixed gateway.e2e.log in the script directory.
Wrap pre-flight, mode setup, scenario definitions, and the scenario loop in a
single try/catch/finally so the finally always writes the summary, stops the
transcript, and zips the bundle -- even on a pre-flight failure. The existing
per-scenario gateway-cleanup try/finally stays nested inside. All scenario
logic, scoring rules, and comments are preserved.
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(mxc-e2e): address CodeRabbit review on run-mxc-e2e.ps1
- Require -Scenario when -KeepRunning: the loop breaks after the first
scenario, so a full-suite run would execute only one scenario yet still
report the suite as PASS. Fail fast so a partial run can't be mislabeled
complete.
- Start-Transcript now runs inside the guarded try block with a
$transcriptStarted flag; Stop-Transcript is only called when it actually
started, so a Start-Transcript failure still yields the results bundle.
- Wrap the -Scenario filter in @() so a single exact match stays an array
(reliable .Count and a proper array for the scenario loop on PS 5.1).
Signed-off-by: Prashant S Khodade <pkhodade@nvidia.com>
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(examples): pass gateway config via OPENSHELL_GATEWAY_CONFIG for spaced paths
Start-Process -ArgumentList does not quote array elements, so launching the
gateway with a bare --config <path> token split on any space in the install
path (e.g. C:\Users\First Last\...), and clap rejected the fragment with
'unrecognized subcommand'. Every MXC example launcher that started the gateway
hit this when the kit was unzipped under a path containing a space.
Pass the config path through the OPENSHELL_GATEWAY_CONFIG env var (which the
gateway already reads via clap) and drop the --config token. Env vars carry
spaces safely.
Affected: run-ocsf-audit, run-mxc-e2e, run-demo, run-inference-test,
run-ollama-test. run-mtls-test was not affected (its launch passes no config
path). Root-caused and fix-verified on 7F203-MXC-003 from a spaced path.
Signed-off-by: Akber Raza <akberr@nvidia.com>
* fix(run-mxc-e2e): improve scoring logic and enhance command execution handling
* fix(mxc): reconcile proxy support after rebase
Restore the proxy-enabled OCSF audit example removed by
|
||
|
|
5b9daab935 |
fix(ci): restore mise run ci on macOS (#3294)
- Replace BSD-incompatible in-place sed calls with portable temp-file rewrites. - Remove test-only shell interception and capture generated gateway config directly. - Allow parity tests to use supplied supervisor binaries without resolving a Linux target. - Normalize temporary-directory paths and use portable RPM config installation. - Set a valid setuptools-scm version for Python protobuf generation in Jujutsu checkouts. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
5b57f0d154 |
fix(ci): restore Windows test portability (#3288)
* fix(ci): restore Windows test portability Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(tasks): skip Unix lockfile check on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(tasks): use buf shim on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
b92620e838 |
ci(rust): reject stale Cargo lockfiles (#3227)
* ci(rust): reject stale Cargo lockfiles Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(ci): clarify lockfile validation policy Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(ci): structure and test Cargo lockfile validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(ci): remove lockfile validator regression tests Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(rust): complete locked validation and lint examples Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(rust): check lockfile diffs after validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(build): shorten lockfile validation notes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(rust): skip lockfile check after failures Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
02b664bb0d |
refactor(config): normalize and enforce gateway schema v2 (#2814)
* refactor(config): normalize compute driver field names Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): introduce canonical gateway fields Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * refactor(config): enforce gateway schema version 2 Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve compute driver runtime guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address schema v2 review regressions Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): complete schema v2 migration safeguards Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): expand schema v2 regression coverage Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(config): add schema v2 parity manifest Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): correct parity manifest inventory Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): record schema v2 intentional changes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): disposition schema v2 parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add dual schema parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): establish compute lifecycle parity baseline Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve gateway option compatibility Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record gateway option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * docs(config): close gateway-wide parity gaps Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(podman): apply configured pids limit Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): validate Podman option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add Kubernetes option parity harness Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record Kubernetes option parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition VM parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): add external driver parity lane Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): preserve external driver pull policy Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity artifacts and launches Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): require clean parity build sources Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): use isolated supervisor tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): qualify parity image tags Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): serve parity supervisor locally Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): isolate parity podman services Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): harden parity evidence provenance Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): pin parity sandbox artifacts Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): attest parity runtime inputs Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): bind parity runtime evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): record compute boundary parity Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(e2e): disposition cross-cutting parity lanes Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight gateway config upgrades Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): preserve rebase integration guarantees Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(ci): isolate temporary git signing config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): update remaining schema v2 consumers Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(ci): provide e2fs tools to VM tests Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): align preflight with gateway startup Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(vm): preserve rootfs tar configuration Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * chore(config): adopt duration unit constructors Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(packaging): preflight RPM gateway config Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(config): address driver review findings Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(e2e): require fresh semantic parity evidence Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * fix(docker): update tests for renamed sandbox label Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> * test(gateway): preserve selective driver coverage after rebase Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Jesse Jaggars <jjaggars@redhat.com> Signed-off-by: Drew Newberry <anewberry@nvidia.com> Co-authored-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
38f2aef930 |
feat(gateway): support selective compute driver builds (#3118)
* feat(gateway): support selective compute driver builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(gateway): support selective Windows MXC builds Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
ddc8bba967 |
ci(windows): add Windows MSVC CI jobs (#2738)
* fix(ci): preserve Windows Rust build cache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): invalidate empty Windows caches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): cache Windows builds with sccache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): restore target directory caching Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): use prebuilt Z3 on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * perf(ci): layer sccache on Windows target cache Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): split PR checks from main validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): separate checks builds and cache seeding Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): simplify Windows build dependency Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): rely on Windows job dependency status Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): use valid opt-in Windows ARM runner Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): keep ARM64 validation local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): install Clippy for Windows validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): focus platform lint coverage Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(licenses): explain bzip2 allowance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): simplify workflow name Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): allow async platform stub Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(network): make file fingerprints portable Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): lint supported deliverables Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): allow platform-gated lint Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(ci): align Windows cache action with main Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): align Windows validation with prerequisites Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): pin Rust toolchain action Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): use enterprise-approved Windows actions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): restore strict MSVC validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): run Rust tests with nextest Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): normalize nextest lock provenance Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): add native arm64 validation Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): lock nextest for Windows ARM64 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): resolve duplicate MXC authentication method Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(conformance): use native absolute paths on Windows Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(windows): address MSVC review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): simplify cache key names Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): isolate Windows Rust toolchains for stable caches Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(ci): configure Rustup home in runner setup Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): surface sccache server write diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): remove temporary cache diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * ci(windows): address review feedback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(windows): reconcile merged driver capabilities Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(deps): preserve AWS-LC-only lockfile Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(deps): allow z3 prebuilt TLS wrapper Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
ce25acca5a |
fix(build): honor Cargo target directory when staging binaries (#3262)
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
a0814443f1 |
feat(docs): publish versioned release snapshots (#3149)
* feat(docs): add version availability labels Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): use supported Python for sync Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(docs): publish versioned docs from releases Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): format dev version label Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * chore(docs): upgrade Fern CLI to 5.112.0 Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): make release publishing monotonic Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): preserve snapshot release identity Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(fern): document versioned publishing Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(docs): cover explicit snapshot rollback Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(docs): use Fern refs for versions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): bundle components for ref versions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * revert(docs): keep complete version copies Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(docs): publish latest and dev channels Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
f4dc6be4b2 |
refactor(inference): remove managed inference routes (#3195)
* refactor(inference): remove managed inference routes Closes #3172 Remove the inference route control plane, inference.local data path, built-in router crate, and SDK surface. Move inference workloads to explicitly imported provider profiles and native endpoints, with migration cleanup and updated tests and documentation. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> * fix(policy): preserve alternate upstream isolation Restore the provider policy activation guard so legacy OpenAI and Anthropic providers configured for alternate base URLs do not grant egress to the built-in public vendor endpoints. Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
8af79a7f4b |
fix(podman): resolve macOS Podman socket dynamically (#3135)
* docs(podman): document macOS socket path mismatch and dynamic lookup On macOS, Homebrew-installed Podman does not create the default socket path that the Podman driver probes. Document the OPENSHELL_PODMAN_SOCKET override and the podman machine inspect lookup in both the compute drivers reference and the debug-openshell-cluster skill. Fixes #1690 Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(podman): resolve macOS Podman socket dynamically Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: restore debug skill file Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: drop legacy debug skill path Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(podman): trim unrelated e2e changes Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * fix(e2e): harden shell array expansion Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> * chore: remove unrelated skill note * ci: retrigger checks Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> --------- Signed-off-by: Gaizka Menendez Hernandez <gmenende@redhat.com> |
||
|
|
3693b32841 |
ci(trivy): add artifact and PR configuration scans (#3185)
* ci(trivy): add artifact and PR configuration scans Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(ci): harden Trivy gate detection and finding diff Signed-off-by: Adrien Langou <alangou@nvidia.com> * feat(ci): scan released artifacts in release pipelines Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(ci): harden and simplify Trivy scans Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(ci): consolidate Trivy reports and prevent collisions Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
519e5eb35f |
feat(e2e): make e2e:kubernetes work transparently on OpenShift (#3183)
* feat(e2e): make e2e:kubernetes work transparently on OpenShift
Running `mise run e2e:kubernetes` on OpenShift required manual namespace
creation, SCC grants, Helm value overrides, and cleanup. A separate
`e2e:openshift` task existed but only checked pod readiness without
running the Rust e2e test suite, and even with the suite wired up the
SSH-relay `sandbox connect` path stalled to the ready timeout because
`kubectl port-forward` cannot carry round-trip-heavy SSH over the
internet.
The harness now auto-detects OpenShift via the `route.openshift.io` API
group and, on OpenShift, both configures the cluster and switches the
gateway transport automatically:
- Drives the gateway through a passthrough OpenShift Route secured with
mandatory mTLS instead of port-forward, so the connect suites
(live_policy_update, port_forward, sync, connect-based
sandbox_lifecycle, settings_management) actually pass. Computes the
Route host from the cluster ingress domain, extracts client mTLS
material from the openshell-client-tls secret, waits for the Route to
serve mTLS, asserts a certless caller is rejected at the TLS
handshake, and registers an mTLS CLI gateway pointing at the Route.
- Applies an SCC-compatible Helm values overlay that removes hardcoded
runAsUser/fsGroup, letting OpenShift assign UIDs from the namespace
range.
- Grants the privileged SCC to openshell-sandbox before Helm install
and removes it during cleanup.
- Grants the anyuid SCC to the PostgreSQL fixture service account in
DB scenarios and removes it during cleanup.
- All oc commands use --context to target the correct cluster.
The OpenShift e2e overlay (ci/values-openshift-e2e.yaml) turns TLS back
on, enables the Route, promotes the cert-verified caller to a dev
principal, and forces `image.pullPolicy`/`supervisor.image.pullPolicy`
to Always so runs against the `latest` upstream image use it instead of
a stale copy cached on the cluster nodes. Every OpenShift branch is
gated on OPENSHIFT_DETECTED, so the vanilla-Kubernetes port-forward path
is unchanged.
The Helm template for podSecurityContext is wrapped with {{- with }} so
null values omit the block instead of rendering invalid YAML.
The separate e2e:openshift task and e2e-openshift.sh script are removed
since e2e:kubernetes now covers OpenShift.
TESTING.md is updated with Kubernetes e2e documentation including
OpenShift auto-detection, dropping the e2e-host-gateway feature on
remote clusters, pinning IMAGE_TAG when the CLI and image versions
differ, task variants, and environment variables.
The debug-openshell-cluster skill gains an OpenShift platform row and
two SCC failure patterns (gateway rejected over hardcoded runAsUser,
sandbox missing the privileged SCC) covering the SCC handling and
podSecurityContext behavior this change introduces.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
* fix(e2e): harden OpenShift SCC cleanup, mTLS gate, and Route timeout
Track the anyuid SCC grant for the PostgreSQL fixture with a dedicated
OPENSHIFT_POSTGRES_SCC_GRANTED flag set before the fixture apply, so a
failed apply no longer leaks the binding; cleanup now revokes it whenever
the grant succeeded, independent of deploy state.
Validate the Route server cert in the certless security gate (curl
--cacert instead of -k) and classify curl's exit code so only a TLS
client-auth rejection (35/56) counts as the expected certless rejection;
an unrelated DNS/timeout/TLS failure now fails loudly instead of masking
a potential mTLS hole.
Raise the OpenShift Route timeout in the e2e overlay. The default HAProxy
Route timeout is 30s, which severed long-lived transfers (large sandbox
upload/download, SSH-relay `sandbox connect`) mid-stream and failed the
sync e2e tests. Set both haproxy.router.openshift.io/timeout and
timeout-tunnel to 300s: a passthrough Route proxies in TCP mode, so
timeout-tunnel governs the established tunnel while timeout covers the
pre-tunnel phase.
Document the OpenShift transport exception, oc prerequisites and SCC
grants, and make the skopeo tag-check example copy-safe in TESTING.md.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
---------
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
|
||
|
|
039b265096 |
feat(middleware): define HTTP response pre-return interface (#3073)
* feat(middleware): define HTTP response pre-return interface Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): clarify HTTP response interface Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): align response result actions Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): expose response reason codes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): share session end reasons Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware)!: finalize HTTP response pre-return contract Replace the separate body_end event with HttpResponseBodyUnit.end_of_stream. Every body-inspecting stage receives exactly one flagged unit, which may be empty; a zero-byte body is one empty flagged unit and OpenShell never reads ahead to set the flag. Defer response trailers from V1 and reserve their field numbers. HTTP/1.0 clients and Content-Length bodies cannot carry trailers and that behavior was undefined. Add HttpResponsePreflight.permitted_body_modes, computed once from the original upstream head so every stage sees the same list, and make an unlisted selection a failure rather than a downgrade. Add the block_delivery preflight action as a successful decision enforced regardless of on_error. Expose Content-Length, Content-Encoding, and Content-Range read-only in preflight. Cap STREAM_BYTES input units at half of max_payload_bytes and permit deferring bytes across replacements only for fail_closed bindings, surfaced as deferral_permitted. Split PEER_DISCONNECT into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT and attribute WebSocket relay failures by direction instead of a generic peer error. Compile the content-guard example in lint and branch checks so proto renames cannot break it silently. BREAKING CHANGE: WebSocketSessionEndReason and WebSocketSessionEnd are replaced by the shared MiddlewareSessionEndReason and MiddlewareSessionEnd. NORMAL_CLOSE is now NORMAL, UPSTREAM_REJECTED is now UPSTREAM_FAILURE, and PEER_DISCONNECT is split into DOWNSTREAM_DISCONNECT and UPSTREAM_DISCONNECT. Enum numbers are unchanged so binary wire compatibility is preserved; generated symbols and JSON names change. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): describe skip as opting out of inspection Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): add body-phase block_delivery and skip_remaining actions Body results may now stop delivery or opt out of inspecting the rest of the response after a prefix. One HttpResponseBlockDelivery message is shared by preflight and body results and documents the difference between blocking before and after head commitment. Drop the field reservations, since nothing in this contract has shipped, and renumber session_end to close the gap. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): share HTTP body leaf messages across directions HttpBodyUnit, HttpBodyPassThrough, HttpBodyTransform, HttpBodySkipRemaining, and HttpBodyMode carry no response-specific semantics, so name them for reuse by the streaming request hook. Envelopes, results, preflight, and block_delivery stay response-specific because commitment semantics differ. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): keep HTTP body leaf messages response-specific Reverts the shared HttpBody* naming. A direction-specific payload such as a response-only semantic mode would otherwise add unreachable variants to the other direction or force a source-breaking fork after 0.1.0. The streaming request hook defines its own HttpRequestBody* messages and copies the shape; SDKs present a direction-neutral body handler over both. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): simplify response proto comments Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): reject undispatched response bindings Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): simplify phase field comment Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(middleware): rename HTTP response preflight result Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(middleware): add HTTP response trailer results Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define response block delivery Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define stage-local response body modes Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define final response body units Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): define streaming response deferral Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): defer whole-body accumulation timeout Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): align response result diagnostics Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * test(middleware): cover upstream WebSocket disconnect Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(middleware): keep response streams unit-local Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(middleware): trim disconnect compatibility note Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
d7cb6e456d |
fix(dev): inherit non-expiring sandbox JWT in local gateway scripts (#2636)
* fix(dev): inherit non-expiring sandbox JWT in local gateway scripts The local gateway launcher scripts hardcode gateway_jwt.ttl_secs = 3600, which overrides the non-expiring default introduced in #1721. Local Docker, Podman, and VM sandboxes are still unrecoverable when the gateway is down longer than that TTL: the on-disk token expires and only the Kubernetes ServiceAccount path can rebootstrap, so the supervisor crash-loops on policy fetch and the sandbox never leaves Provisioning. Drop the override so local drivers inherit the default. gateway.sh also serves the kubernetes driver, which is a shared deployment and must keep a positive TTL, so it now emits ttl_secs only for that driver. The e2e regression test added in #1721 does not catch this because the e2e harness uses its own configs, which already set ttl_secs = 0. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(dev): expand sandbox JWT TTL in gateway config Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |
||
|
|
64a858dade |
fix(ci): restore Codex Security scan execution (#3124)
* refactor(ci): resolve Codex Security range in Python Signed-off-by: Adrien Langou <alangou@nvidia.com> * fix(ci): allow unprivileged userns for Codex sandbox Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
cc4ded2088 |
feat(helm): split gateway and workspace charts (#2643)
* feat(helm): split gateway and workspace charts Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(helm): preserve split chart upgrade compatibility Keep workspace manifests valid after value validation and default legacy reused values to the combined resource topology. * fix(ci): preserve VM runtime for E2E The Rust cache restores target/ after VM runtime artifacts are staged, overwriting target/vm-runtime-compressed before openshell-driver-vm is built. Stage the compressed runtime outside target and pass that location through OPENSHELL_VM_RUNTIME_COMPRESSED_DIR so build.rs can embed the supervisor. Also locate the Helm split-ownership test repository root from the script path rather than git rev-parse. The test runs in a container where the GitHub checkout can be owned by a different UID and rejected as dubious ownership. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> * fix(ci): install yq for Helm ownership test The split-chart ownership regression uses yq to inspect rendered YAML, but the Helm CI container installs only tools declared in mise. Declare and lock yq so mise install --locked provides the test dependency. Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> --------- Signed-off-by: Dhiraj Bokde <dbokde@nvidia.com> |
||
|
|
9ca19e6c80 |
refactor(compute): decouple gateway driver composition (#2823)
* refactor(compute): decouple gateway driver composition Move first-party composition and VM process ownership into openshell-gateway, leaving openshell-server backend-independent. Update packaging and build references with the new crate, simplify the compiled-driver boundary, and keep the driver-free gateway path buildable with bundled Z3 tooling. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(telemetry): bound compute driver categories Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(core): keep runtime transport generic Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): complete server driver decoupling Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): preserve driver integrations after rebase Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve docker tracing after decoupling Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(compute): preserve driver behavior after extraction Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): remove MXC policy side channel Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): separate policy delivery from readiness Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
5b925dd8af |
feat(build): add defaults-without-telemetry feature alias (#2843)
* feat(build): add defaults-without-telemetry feature alias Cargo cannot subtract a single default feature, so compiling telemetry out meant `--no-default-features` plus a hand-maintained keep-list of the crate's other defaults. That keep-list was already wrong for operators: telemetry is the only default on openshell-server and openshell-driver-vm, but openshell-sandbox also defaults to `bundled-ca-roots`, so a bare `--no-default-features` silently swapped the supervisor onto the platform trust store. Add a `defaults-without-telemetry` alias to each of the three telemetry- carrying binary crates, enumerating every default except `telemetry`. Telemetry-free builds become `--no-default-features --features defaults-without-telemetry` and stay correct as the default set grows. The alias is a keep-list, not a switch. Enabling it on top of the defaults would otherwise produce a telemetry-on binary that reads as telemetry-free, so each crate root carries a `compile_error!` for the `telemetry` + `defaults-without-telemetry` combination. Add `rust:verify:defaults-without-telemetry` to guard both properties: each alias still equals its crate's defaults minus `telemetry`, and the mutual-exclusion error is wired up. The additive-misuse check matches on the `compile_error!` text rather than a nonzero exit code so it cannot pass vacuously on hosts where openshell-driver-vm fails to build for unrelated reasons. `rust:verify:telemetry-off` now builds through the alias. Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix feature alias for openshell-server Signed-off-by: Russell Bryant <rbryant@redhat.com> * fix(ci): run Rust verification in Nix shell Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> --------- Signed-off-by: Russell Bryant <rbryant@redhat.com> Signed-off-by: John Myers <9696606+johntmyers@users.noreply.github.com> Co-authored-by: John Myers <9696606+johntmyers@users.noreply.github.com> |
||
|
|
a4f9c762ce |
fix(release): handle prerelease tag builds (#3094)
Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
c8f13205e3 |
ci(release): publish prerelease artifacts (#3093)
Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
f7180c0fd6 |
feat(ci): add Codex Security release qualification (#3087)
* feat(ci): add Codex Security release qualification Scan cumulative release-train diffs through NVIDIA inference and publish findings to Code Scanning. Signed-off-by: alangou <alangou@nvidia.com> * fix(ci): disable package cache for security scan Prevent cache poisoning in the tag-triggered Codex Security workflow. Signed-off-by: alangou <alangou@nvidia.com> * refactor(ci): simplify Codex Security reporting Remove custom inference cost accounting so the workflow remains focused on scanning and SARIF publication. Signed-off-by: alangou <alangou@nvidia.com> --------- Signed-off-by: alangou <alangou@nvidia.com> |
||
|
|
bb70461878 |
test(e2e): run conformance in gateway lanes (#2925)
* test(e2e): isolate VM-specific smoke assertions Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): add portable CLI conformance baseline Signed-off-by: Evan Lezar <elezar@nvidia.com> * feat(conformance): add standalone CLI runner Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(e2e): run conformance in gateway lanes Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
197b41371d |
fix(dev): harden local cluster and gateway startup (#2993)
* fix(helm): refresh kubeconfig for existing k3d clusters Docker can recreate the k3d load balancer on a new API port. Start existing clusters and prefer fresh k3d entries so create does not retain a stale endpoint. Signed-off-by: Kris Hicks <khicks@nvidia.com> * fix(dev): conditionally enable local OTLP export Probe port 4317 before adding OTLP configuration for the VM, Docker, and Podman gateway tasks. Document the startup behavior and troubleshooting for local collector availability. Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
f68867b869 |
feat(gateway): identify gateways in exported traces (#2647)
* feat(gateway): add installation name configuration Add a first-class operator-assigned gateway name with TOML, CLI, environment, and Helm configuration surfaces. Local gateways default to openshell, while Helm defaults to the chart fullname; operators sharing a collector across namespaces or clusters can set a globally distinct name. Signed-off-by: Kris Hicks <khicks@nvidia.com> * feat(gateway): identify gateways in exported traces Attach the configured gateway installation name and compute driver to the gateway OpenTelemetry resource so operators can filter traces from multiple installations that share a collector. Forward the gateway name and OTLP endpoint to managed external drivers so their distinct service resources carry the same installation identity. Keep service.name stable per process type, omit blank resource values, and leave per-span operation names and request attributes unchanged. Refs #2507 Signed-off-by: Kris Hicks <khicks@nvidia.com> --------- Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
37072ee81c |
feat(build): publish OCI SBOM and provenance attestations (#2836)
* feat(build): embed auditable Rust dependency metadata Signed-off-by: Adrien Langou <alangou@nvidia.com> * feat(build): publish OCI SBOM and provenance attestations Signed-off-by: Adrien Langou <alangou@nvidia.com> --------- Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
9f88f8ff9b |
ci: remove rootless podman e2e lane (#2981)
Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
bcd517bbe0 |
feat(driver-mxc): native Windows MXC compute driver + server wiring (#2721)
* feat(driver): add MXC compute driver for Windows isolation sessions Introduces the openshell-driver-mxc crate implementing ComputeDriver backed by Microsoft MXC isolation sessions (Windows only). Wires the new driver into the server's build_compute_runtime dispatch and adds the Mxc variant to ComputeDriverKind. Also adds a local protobuf-src stub (tools/protobuf-src-local) to unblock Windows builds that lack MSYS2/MinGW, and pins the zig Windows x64 toolchain in mise.lock. (cherry picked from commit 4f7012224efb18fbfeb47aa87e0cfd3f036f32f0) Signed-off-by: Jamie King <jamiek@nvidia.com> * wip(mxc): checkpoint hung-agent work (recon, policy_map embed, A1 wiring, demo artifacts) Safety checkpoint of uncommitted work from the background agent run that stalled mid-Step-7. Includes: mxc-driver-recon.md (Step 0.5), policy_map.rs (~876L embedded mapper), A1 policy-threading edits across driver.rs/policy.rs/mxc.rs/compute/mod.rs, and examples/ (demo.yaml + mxc-gateway.toml). Not yet verified to compile end-to-end; to be reorganized into the skill's Step 11 commit sequence. (cherry picked from commit 38e42c03870be3d10e984a54f17a3b61122ff510) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(mxc): fix lifecycle and policy unit-test compile drift - Bring futures::StreamExt into scope for the watch-stream `.next()` call in driver::lifecycle_tests so the negative policy proof test compiles. - Bind a local `mapper` and drop the unused/deprecated NetworkBinary in the embedded-mapper network-policy rejection test. Signed-off-by: Jamie King <jamiek@nvidia.com> (cherry picked from commit 039b0baf98735ca672dae52be8d3af2417dc0c1a) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): downgrade missing sandbox_token to debug log The gateway mints `sandbox_token` only when a sandbox-JWT issuer is configured. There is no in-sandbox supervisor on MXC (supervisor-removal design — D1/D4), so no component ever consumes the token; requiring it on the driver side blocks the demo's `--disable-tls` smoke gateway with a spurious `invalid_argument`. Log the absence and proceed instead. Signed-off-by: Jamie King <jamiek@nvidia.com> (cherry picked from commit cea209797d0edcb1d152251e748900b0a63cca62) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): keep sandbox Ready after a successful one-shot agent exec monitor_exec demoted Ready->Error on exit 0 (reason ExecCompleted), so the positive demo (write hello.txt + exit) landed in Error phase. Keep Ready=True (reason AgentCompleted) on success; only non-zero exits go to ExecFailed. Tighten the positive lifecycle test to assert the terminal condition stays Ready=True/AgentCompleted. Verified live via gateway mock round-trip: phase now Provisioning->Ready with no demotion. (cherry picked from commit 54ab030f03ca0f83d0050d8dd843b633651684ad) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(mxc): add processContainer backend for default-deny enforcement Add a backend selector to the MXC driver (isolation_session default | process_container). process_container drives a one-shot AppContainer that is genuinely default-deny: a write to any ungranted path is denied by the OS, unlike isolation_session which is grant-only and cannot deny. The lifecycle forks on the flag - isolation_session keeps provision/start/exec, process_container runs a single ephemeral container via run_oneshot. Also: run-demo.ps1 gains -Backend and hardens the CLI register/create calls; docs corrected to state isolation_session does NOT deny out-of-policy writes and that the negative proof requires process_container. Verified end-to-end on a real demo box (gateway -> CLI -> driver -> MXC): in-policy write succeeds, out-of-policy write denied (PermissionDenied), OVERALL: PASS. (cherry picked from commit c6cde3860bbe1b8edb3147d3e840f6bf0ece32d8) Signed-off-by: Jamie King <jamiek@nvidia.com> * refactor(driver-mxc): embed policy mapper as a module; remove standalone crate Adopt the proto-based mapper (map_to_mxc) as the single source of truth, embedded in openshell-driver-mxc as a Windows-gated `policy_map` module. Rewire EmbeddedPolicyMapper to call it directly on the typed SandboxPolicy, deleting the serde_yaml proto->YAML bridge. Move the CLI to a windows-gated example and the parity tests into the crate; delete openshell-policy-mapper. - gate policy_map + seam Windows-only (MXC is Windows-only) - drop serde_yaml; add dev-deps openshell-policy, clap, anyhow - normalize mapped paths to Windows form in the seam, in one place - docs: add driver-mxc to AGENTS.md table; correct design doc section 17 test lane Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit f22f9c7a25b9a651c5c5cc73f62fb01c4d6c1a8d) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): implement lossless split_policy for proxy-delegated egress Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 96d6afa0e2dc6a1d54edd12c34a0ceb0a30dadd0) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): implement Pattern-C governed-egress split through the policy seam - split_policy: SocketAddr proxy_redirect (replaces bare port), processcontainer containment guard naming MXC M1, version preserved in the trimmed proxy_policy, delegation reported as an info loss item - seam: MappedConfig carries trimmed_policy + proxy_addr; MapCtx.egress selects the split path; coarse path unchanged when egress is disabled - driver: [openshell.drivers.mxc] egress_proxy / egress_proxy_addr config, validated at create (isolation_session rejected until M1); lifecycle threads the redirect into provision and stores the trimmed policy per sandbox, emitting an EgressRedirect platform event - mxc: optional MxcNetwork block (defaultPolicy=block + proxy) in provision and one-shot configs; mock records configs for test assertions - tests: lossless-invariant suite over all example policies (validate + serialize round-trip), split lifecycle proof, M1 rejection; example gains --split --proxy-addr writing mxc-config.json / trimmed-policy.yaml / loss-report.json Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 34d54ad9f25dc6034c3ba15668555ff0d22cddd8) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): emit MXC network.proxy as {localhost: port} Verified against the real wxc-exec 0.6.0-alpha via --dry-run: MXC accepts only the {localhost: N} proxy shape (the form the design doc specifies) and rejects {host, port} with a parse error. Schema 0.6.0-alpha can express only a loopback port, so non-127.0.0.1 redirect addresses are now rejected: split_policy emits an error loss (no proxy block) and the driver refuses egress_proxy_addr values off 127.0.0.1. Per-sandbox attribution must use per-sandbox ports until the schema widens. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit edde8d5434571fd5398409204fcf6862672c0793) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): serialize isolation_session stop/deprovision as unit variants Empirical contract finding from the real test lane (build 26300.8553, wxc-exec 2026-06-10): the stop and deprovision experimental blocks are unit variants in the wxc-exec schema and must serialize as null; sending {} is rejected with malformed_request (invalid type: map, expected unit), while provision/start accept maps. The production invoker, the real-lane test, the probe script, and the e2e runner all sent {} - the driver could provision and run an agent but never stop or delete an isolation-session sandbox against this build. Pinned by a unit test. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 0df39ca0b22ebc21eb965b2b567a5b2cac26af32) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): add Tier-0 mapper coverage matrix with schema drift guard Three-quadrant, table-driven matrix (38 tests): mappable fields assert exact MXC output; every OpenShell field MXC cannot express asserts a loss item with the expected severity (and seam rejection on error); an empty policy asserts the restrictive default-deny posture for every MXC knob OpenShell does not control. The handled_fields_inventory drift guard serializes a fully-populated policy and compares its YAML keys against the mapper-handled field lists, so a new openshell-policy field fails the suite until consciously mapped, delegated, or reported as loss. Re-exports the policy seam types for integration tests; adds serde_yml, base64, serde_json as dev-dependencies. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 91807f984a3b16846e35d6ca0d5ec41057aafa3a) Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(driver-mxc): inject agent_env into sandbox process.env Add MxcComputeConfig.agent_env: each entry is either KEY=VALUE (verbatim) or a bare KEY resolved from the gateway host environment at launch, keeping secrets (e.g. inference API keys) out of the config file. Wire it into the agent process so gateway-launched agents can authenticate to cloud endpoints (process.env was previously hardcoded empty). Unit-tested via resolve_agent_env_passthrough_and_host_lookup. Also add a gateway-driven cloud-inference (T1) test harness: mxc-inference.toml (agent_env + curl agent), inference.yaml policy, and run-inference-test.ps1 which starts the gateway, creates an isolation_session sandbox, runs an authenticated Nemotron call, and bundles redacted results. Documented agent_env in mxc-gateway.toml. Validated end-to-end on the test box (chat HTTP 200 + completion via the gateway). (cherry picked from commit 94d9e827b8b77e0af9dc654943ea8e7d8400cc9f) Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): avoid unsafe env mutation in resolve_agent_env test Replace std::env::{set,remove}_var (unsafe + racy under parallel test execution in edition 2024) with a read-only PATH lookup. Preserves all three behaviors under test and drops the #[allow(unsafe_code)]. (cherry picked from commit ac5766eeab2db4e8cc6fcd8d8a97809edaf3df30) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(mxc): adapt MXC driver to current GitHub OpenShell API The MXC driver crate was authored on GitLab against an earlier proto/core API. Adapt it to the API on GitHub main: - build_capabilities_response no longer takes supports_interactive_session - DriverSandboxSpec.gpu (bool) is now resource_requirements; detect GPU via effective_driver_gpu_count(driver_gpu_requirements(..)) - DriverSandbox gained a `workspace` field - SandboxPolicy gained `network_middlewares`: pass it through the proxy split, emit a loss item on the coarse MXC path, and account for it in the mapper drift-guard test Verified: cargo check + 75 mock-based tests pass (lib 27, examples 10, policy_mapper_matrix 38). Signed-off-by: Jamie King <jamiek@nvidia.com> * feat(server): wire the MXC compute driver into the gateway on Windows Register openshell-driver-mxc as the Windows-only in-process compute backend so compute_driver = "mxc" resolves to a working runtime: - ComputeRuntime::new_mxc, adapted to the current 11-arg from_driver - mxc_policy_sink A1 side channel, staged in create_sandbox before dispatch - mxc_config_from_context loader and the Mxc dispatch arm (Windows constructs; other targets return an explicit "Windows-only" error) - Windows-gated openshell-driver-mxc dependency - Mxc arms for the telemetry, config-file required-fields, and CLI reserved-builtin matches to keep them exhaustive/correct Verified with cargo check --workspace --features openshell-prover/bundled-z3 on x86_64-pc-windows-msvc, stacked on PR #2496. Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): implement GetGatewayListenerRequirements for #2496 base Signed-off-by: Jamie King <jamiek@nvidia.com> * test(driver-mxc): add probe-gated real wxc-exec test lane (no mocks) - tests/wxc_exec_real.rs: ignored-by-default integration tests against a real wxc-exec. Six --dry-run contract tests run wherever the binary exists (they caught the network.proxy shape mismatch); enforcement tests (processcontainer default-deny positive/negative, isolation session lifecycle round trip with a deprovision drop-guard) probe the backend and SKIP with a recorded reason where it is not live. - examples/probe-mxc-host.ps1: classifies a host (OS build, --probe, per-backend trial) and emits a JSON capability verdict. - examples/run-mxc-e2e.ps1 + e2e-policies/: scenario runner generalizing run-demo.ps1 (fs-rw, fs-readonly, fs-default-deny-empty, network-policy-rejected) with PASS/FAIL/SKIP gating and a stale OPENSHELL_MXC_MOCK_WXC guard in real mode. - tasks/windows.toml: windows:test:mxc-real:x64, windows:e2e:mxc, windows:e2e:mxc:mock. Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> (cherry picked from commit 49afafe892caded59c4df50a9b652011ad97f41c) Signed-off-by: Jamie King <jamiek@nvidia.com> * fix(driver-mxc): address PR review feedback Signed-off-by: Shailendra Singh <shailendras@nvidia.com> * docs: defer public MXC documentation Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(driver-mxc): build Windows capabilities response Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * fix(server): gate in-tree tracing on Windows Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): fix cross-platform test assumptions Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): verify process and PEM portably Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> * test(windows): keep lifecycle command in policy Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> --------- Signed-off-by: Jamie King <jamiek@nvidia.com> Signed-off-by: Giedrius Burachas <gburachas@nvidia.com> Signed-off-by: Shailendra Singh <shailendras@nvidia.com> Signed-off-by: Drew Newberry <385+drew@users.noreply.github.com> Co-authored-by: Prashant K <pkhodade@nvidia.com> Co-authored-by: Giedrius Burachas <gburachas@nvidia.com> Co-authored-by: Shailendra Singh <shailendras@nvidia.com> Co-authored-by: Drew Newberry <385+drew@users.noreply.github.com> |
||
|
|
d0dfb22baf |
feat(kubernetes): export driver traces over OTLP (#2958)
Mirror the VM, Podman, and Docker driver tracing setup for Kubernetes. Export standalone driver spans through OTLP/gRPC as the distinct openshell-driver-kubernetes service, preserve gateway trace context, record lifecycle operations and gRPC failures, and flush spans on shutdown. Kubernetes currently runs in-process when selected as a built-in gateway driver. Use the temporary server-boundary shim shared with Podman and Docker so traces retain the shape they will have when Kubernetes moves to a separate process. Move the common ComputeDriver RPC tracing layer into openshell-otel to keep all drivers aligned. Propagate the active W3C context through the controller-reserved Sandbox annotation and enable Agent Sandbox OTLP export in the local k3s workflow. This connects asynchronous controller reconciliation spans to the originating OpenShell create trace. Expose gateway OTLP configuration through Helm and add an Aspire collector to the local k3s workflow. Extend helm:k3s:forward with OTLP ingest and trace UI forwarding for Kubernetes and local container gateway development. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
c399342649 |
feat(dev): unify local Kubernetes gateway workflow (#2914)
Make the local k3s gateway workflow match the Docker and Podman flows by registering and selecting successful plaintext Skaffold deployments with the OpenShell CLI. Derive the registration name from the worktree-specific k3d cluster name so parallel worktrees retain independent gateway metadata. Add helm:k3s:forward as the standard way to expose the Kubernetes gateway on localhost:8090, and update the development and debugging guidance to use the active registered gateway instead of one-off endpoint flags. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
18ce13b9b1 |
feat(providers): expose actionable OAuth refresh failures (#2887)
* fix(providers): classify OAuth refresh failures Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * test(providers): add Keycloak refresh e2e lane Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): harden OAuth refresh recovery Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): classify post-mint refresh failures Signed-off-by: Mrunal Patel <mrunalp@gmail.com> * fix(providers): tolerate malformed OAuth subtypes Signed-off-by: Mrunal Patel <mrunalp@gmail.com> --------- Signed-off-by: Mrunal Patel <mrunalp@gmail.com> |
||
|
|
fb6610df39 |
feat(build): embed auditable Rust dependency metadata (#2734)
Signed-off-by: Adrien Langou <alangou@nvidia.com> |
||
|
|
455883905a |
fix(python): remove CLI from wheel (#2321)
The Maturin-based wheel packaging was a historical remnant from when the local gateway launch path and OpenShell CLI were coupled in one binary. The gateway and CLI now ship as standalone artifacts, so the Python distribution should contain only the SDK. Build a single platform-independent setuptools wheel, verify that it cannot contain native code or an openshell entry point, and simplify the release jobs and documentation for SDK-only PyPI installs. Signed-off-by: Simon Scatton <sscatton@nvidia.com> |
||
|
|
0a1f246587 |
feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS (#2512)
* feat(sandbox,podman): trust corporate CA for https:// proxies and intercepted TLS The corporate proxy chaining only accepted plain http:// proxy URLs, so operators whose forward proxy terminates TLS with a private corporate CA had no way to reach it, and TLS-intercepting proxies (mitmproxy, squid ssl-bump) that re-sign tunneled server certificates broke every upstream handshake after CONNECT. The supervisor now accepts https:// proxy URLs: it wraps the connection to the proxy in TLS before the CONNECT handshake, verifying the proxy certificate against the built-in Mozilla roots, the system CA bundle, and an optional operator corporate CA bundle. The upstream dial returns a Plain/Tls stream enum consumed generically by the relay paths. The corporate CA is delivered as a driver-supplied command-line argument (--upstream-proxy-ca-bundle), never an environment variable, matching the hardened proxy-config model where a sandbox image cannot influence the operator's egress boundary. It is folded into the sandbox combined trust bundle (write_ca_files) and the L7 upstream verification store (build_upstream_client_config) at startup, so intercepted upstream handshakes succeed and sandbox workloads trust the re-signed certificates. Configuration is fail-closed: a CA bundle set without a proxy, or an unreadable or certificate-free file, is fatal rather than silently weakening the trust boundary. The shared parse_upstream_proxy_url validator accepts https:// (recording the scheme so the driver and supervisor agree), keeping the explicit-port requirement. The Podman driver gains a proxy_ca_bundle operator setting (TOML, --sandbox-proxy-ca-bundle, OPENSHELL_SANDBOX_PROXY_CA_BUNDLE) that bind-mounts the host PEM read-only into the sandbox (a CA certificate is not secret) and points --upstream-proxy-ca-bundle at it, with a create-time readability check. The standalone dev gateway task passes OPENSHELL_SANDBOX_PROXY_CA_BUNDLE through to the generated podman config, so a local gateway can be pointed at a TLS-intercepting proxy without hand-editing the regenerated TOML. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox): reject CA bundles with valid PEM framing but invalid X.509 DER The proxy CA bundle validation counted PEM blocks that base64-decoded successfully, but did not verify the decoded bytes were accepted as trust anchors by RootCertStore. A bundle with syntactically valid PEM framing but invalid DER would pass the startup check while contributing zero usable anchors, causing opaque TLS failures at runtime instead of a fail-closed startup error. Validate decoded certificates through RootCertStore::add_parsable_certificates and reject the bundle unless at least one is accepted. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(sandbox,podman): allow proxy auth without insecure acknowledgement for https:// proxies For an https:// proxy the Proxy-Authorization credential travels inside the verified TLS session, so the proxy_auth_allow_insecure acknowledgement is unnecessary. Previously both http:// and https:// proxies required it, producing a misleading cleartext-risk diagnostic for a path that is already encrypted. Skip the requirement when the proxy URL uses https://; the acknowledgement is still tolerated if set. Updated in both the supervisor and Podman driver validation paths, with docs and tests. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix: format Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(kubernetes): use truly unsupported scheme in proxy validation test https:// is now a supported proxy scheme after a13c4dce, so the unsupported-scheme test must use a genuinely unsupported scheme. Signed-off-by: Philippe Martin <phmartin@redhat.com> * fix(e2e): sign the https proxy fixture listener cert with a CA The corporate-proxy E2E fixture served a single `openssl req -x509` certificate as its TLS listener identity. OpenSSL marks that certificate `basicConstraints: critical, CA:TRUE`, and rustls refuses a CA certificate presented as an end-entity certificate (CaUsedAsEndEntity). The supervisor's TLS handshake with the proxy therefore failed, the upstream dial errored, and the workload's CONNECT was dropped without a response, so podman_corporate_proxy_trusts_ca_bundle_for_https_proxy failed on the approved destination while policy denial still worked. Generate a corporate CA and a separate listener leaf signed by it, serve the leaf chain, and publish only the CA as the bundle the supervisor trusts. This is what an intercepting proxy actually presents, and it exercises the corporate-CA trust path rather than pinning the listener certificate itself. Refs #1792 Signed-off-by: Philippe Martin <phmartin@redhat.com> --------- Signed-off-by: Philippe Martin <phmartin@redhat.com> |
||
|
|
6c38646c59 |
feat(dev): add dedicated gateway:podman task (#2880)
Previously, Podman could be selected through automatic driver detection or with `mise run gateway -- --driver podman`, but it did not have a dedicated task like the Docker and VM drivers. This adds a gateway:podman task and moves the Podman-specific setup into its own script. The generic gateway task now delegates Podman launches to that script. Additionally: Unlike Docker, which rebuilds and bind-mounts the supervisor binary, Podman uses a dev-tagged supervisor image that can become stale. The default Podman supervisor image is therefore rebuilt on each launch. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
679fe4c334 |
fix(policy): validate the applicable advisor candidate (#2850)
* fix(policy): bind reviews to applicable candidates Build and validate the exact effective-policy candidate before approval, bind review to live policy/provider/credential inputs, and preserve inspected endpoint contracts during mechanistic expansion. Closes #2821 Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(policy): canonicalize advisor review inputs Serialize nested protobuf maps in stable key order for proposal review tokens and effective-policy hashes. Narrow reused multi-port endpoint contracts to the denied port so advisor proposals cannot widen binary access. Add regressions for both cases. Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * test(e2e): keep advisor sandbox running Create the issue 2821 regression sandbox detached with a durable canonical main process so policy denial, approval, and hot-reload checks run before lifecycle exit. Signed-off-by: John Myers <johntmyers@users.noreply.github.com> * fix(policy): apply reviewed draft batches atomically Signed-off-by: John Myers <johntmyers@users.noreply.github.com> --------- Signed-off-by: John Myers <johntmyers@users.noreply.github.com> Co-authored-by: John Myers <johntmyers@users.noreply.github.com> |
||
|
|
3be2cd8a29 |
fix(helm): preflight Agent Sandbox APIs (#2867)
* fix(helm): preflight Agent Sandbox APIs Signed-off-by: Evan Lezar <elezar@nvidia.com> * test(kubernetes): share Agent Sandbox setup Signed-off-by: Evan Lezar <elezar@nvidia.com> * fix(e2e): wait for Agent Sandbox CRD status Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(canary): sparse-checkout sandbox helper Signed-off-by: Evan Lezar <elezar@nvidia.com> --------- Signed-off-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
40f822906c |
feat(compute): add standalone first-party drivers (#2822)
* feat(compute): add standalone first-party drivers Build Docker, Podman, Kubernetes, and VM drivers as external binaries and exercise each through the public compute-driver API. Keep the external E2E setup complete at introduction, including VM image selection, Kubernetes post-renderer isolation, supervisor reuse, and scoped Podman coverage. External Kubernetes endpoints support shared and managed workspace modes. Operator mode remains restricted to the in-process driver because gateway authentication and the driver must share a dynamic namespace allowlist. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * ci(e2e): run managed and external drivers independently Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(compute): cover external driver socket contract Signed-off-by: Evan Lezar <elezar@nvidia.com> * ci(e2e): install bundled Z3 build dependency Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(server): gate in-tree driver tracing Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> Signed-off-by: Evan Lezar <elezar@nvidia.com> Co-authored-by: Evan Lezar <elezar@nvidia.com> |
||
|
|
ef296806f5 |
feat(sandbox): add canonical main process (#2726)
* feat(sandbox): add canonical main process Closes #2710 Persist and supervise one canonical workload per sandbox, attach sandbox connect to its retained session, and make every unexpected main-process exit terminal. Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): simplify canonical main process contract Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve legacy VM main compatibility Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): preserve main status across driver updates Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): satisfy macOS process lint Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): gate Linux exit acknowledgement publisher Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(sandbox): simplify main process plumbing Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): make controlling tty ioctl portable Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(supervisor): initialize canonical process environment Signed-off-by: Drew Newberry <anewberry@nvidia.com> * perf(supervisor): optimize retained main session Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sandbox): detach main session on ctrl-c Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(sandbox): use explicit main detach keys Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(dev): atomically stage Docker supervisor Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(test): align Docker main environment assertion Signed-off-by: Drew Newberry <anewberry@nvidia.com> * feat(sdk): expose canonical main process fields Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com> |
||
|
|
59479f492a |
feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) (#2656)
* feat(k8s): add namespace-per-workspace support (RFC 0011 Phase 3) Implement three workspace namespace modes for the Kubernetes compute driver: shared (default, preserves current single-namespace behavior), managed (auto-creates/deletes namespaces per workspace), and operator (pre-provisioned namespaces with dynamic discovery via label selector or drop-in allowlist file). Key changes: - WorkspaceMode enum and namespace resolution in driver config - Managed namespace lifecycle with ServiceAccount and OpenShift SCC annotation propagation - Cluster-wide sandbox CR watchers for managed/operator modes - NamespaceValidator (Exact/Prefix/Allowlist) for SA token auth - Workspace-aware credential secret storage - Helm ClusterRole for multi-namespace RBAC - Gateway config, architecture, and reference docs Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add end-to-end tests for managed and operator workspace modes introduced in RFC 0011 Phase 3. The managed mode tests verify namespace creation with correct labels, ServiceAccount provisioning, sandbox CR placement, and namespace survival with remaining sandboxes. The operator mode tests verify rejection of unlabeled and nonexistent namespaces. The positive operator path (sandbox in labeled namespace) is known to fail due to an RBAC gap and will be addressed separately. Also fixes Helm 4 compatibility: move SPDX license headers inside conditional guards in 8 chart templates to prevent empty comment-only documents, and fix a trailing whitespace trimmer in clusterrole.yaml that concatenated the license header with apiVersion. Adds cleanup sweep in with-kube-gateway.sh to remove managed and operator namespaces before Helm uninstall, and mise tasks for running each mode independently. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add operator namespace label watcher Spawn a background kube::runtime::watcher in the K8s driver that watches namespaces matching the configured label selector and populates the OperatorNamespaceAllowlist at runtime. The driver owns the allowlist and exposes its Arc so the server can share the same set with the SA token authenticator. create_sandbox now gates pod creation on the allowlist in operator mode — workspaces whose namespace is not yet labeled are rejected at resource render time rather than silently proceeding. Workspace lifecycle itself is unaffected; only sandbox (resource) creation is gated. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): harden operator mode and address review findings Close the fail-open gap in operator mode when only operator_namespace_file is configured: the allowlist is now created unconditionally in operator mode (fail-closed from startup). Implement the namespace file watcher using the notify crate, following the TLS hot-reload pattern (parent-directory watch, 1s debounce, ConfigMap symlink-swap safe). The file format is a JSON array of namespace name strings. Additional fixes from the 10-reviewer audit: - Change allowlist rejection from InvalidArgument to FailedPrecondition so callers know the request may succeed later once the namespace is provisioned. - NamespaceValidator::Allowlist now holds the OperatorNamespaceAllowlist newtype instead of a raw Arc<RwLock<BTreeSet>>, eliminating silent denial on RwLock poison. - Verify LABEL_MANAGED_BY and LABEL_GATEWAY_ID ownership before deleting a managed namespace. - Replace fixed 5s sleep in operator e2e test with a 30s poll loop. - Add Helm validation for workspaceMode values. - Fix Helm README type column and description for operator fields. - Add insert/remove methods to OperatorNamespaceAllowlist; label watcher now uses them instead of reaching through shared(). - Reject configs with both operator_namespace_label and operator_namespace_file set. Signed-off-by: Derek Carr <decarr@redhat.com> * feat(k8s): add workspace-level compute driver RPCs and harden RBAC Decouple namespace lifecycle from sandbox lifecycle by adding EnsureWorkspace/DeleteWorkspace RPCs to the ComputeDriver service. Namespace creation now happens before credential storage and namespace deletion happens on workspace delete, fixing credential storage in managed workspace mode. - Add EnsureWorkspace and DeleteWorkspace proto RPCs with implementations across all compute drivers (K8s managed delegates to ensure_namespace/delete_namespace_if_empty; others no-op) - Wire ensure_workspace into provider create/update/refresh paths so the namespace exists before the credential driver writes secrets - Wire delete_workspace into workspace deletion for cleanup - Remove delete_namespace_if_empty from sandbox deletion path - Scope ClusterRole secrets access to non-shared workspace modes - Add TODO for TLS cert hot-reload in sandbox gRPC client - Harden e2e tests with control-plane sandbox resolution assertions - Fix docker image save --platform flag for OCI index manifests Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address re-review findings and add test coverage - Use server-side apply for TLS secret sync (fixes second sandbox creation failure when TLS is enabled) - Scope gateway-ID label selector unconditionally across all workspace modes (fixes operator reads/watches/deletes seeing foreign sandboxes) - Validate operator allowlist in EnsureWorkspace and DeleteWorkspace RPCs (prevents credential writes to namespaces outside the allowlist) - Extend ClusterRole secrets patch+delete to all non-shared modes with credential driver enabled (fixes operator credential storage RBAC) - Validate namespace ownership on 409 conflict in ensure_namespace (prevents adopting unowned namespaces in managed mode) - Replace delete_namespace_if_empty with unconditional delete_namespace letting Kubernetes cascade cleanup (fixes stuck terminating CRs) - Strengthen NetworkPolicy TODO to cover both managed and operator modes - Extract selector and ownership logic into testable free functions - Add unit tests for gateway-ID selectors and namespace ownership - Add Helm ClusterRole RBAC tests for operator credential driver Signed-off-by: Derek Carr <decarr@redhat.com> * ci(k8s): add workspace managed and operator mode e2e to CI Wire the existing e2e:kubernetes:workspace-managed and e2e:kubernetes:workspace-operator mise tasks into the branch-e2e workflow so they run alongside the other core Kubernetes e2e suites. Both are gated by run_core_e2e and included in the Core E2E result gate. Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): add e2e tests for workspace namespace modes Add 7 new e2e tests covering workspace namespace lifecycle, TLS secret copying, ownership conflict detection, DNS-1123 validation, operator namespace preservation, and dynamic label watcher behavior. Fix async sandbox deletion race condition in existing tests by polling sandbox list instead of asserting immediately after delete. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): grant secrets/patch unconditionally and backfill gateway-id labels Address two review findings: 1. RBAC: server-side apply (PATCH) is used for TLS secret sync in multi-namespace modes, but the ClusterRole only granted patch when the kubernetes-secrets credential driver was enabled. Grant patch unconditionally for non-shared modes since TLS sync always needs it; keep delete gated on the credential driver. 2. Upgrade safety: the new gateway-id label selector would orphan legacy Sandbox CRs that predate its introduction. Add a startup backfill in shared mode that patches any managed Sandbox CR missing the gateway-id label before the driver begins serving requests. Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address follow-up review findings Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): preserve workspace lookup after rebase Signed-off-by: Derek Carr <decarr@redhat.com> * fix(helm): allow managed secret creation Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): stop pods in workspace namespace Signed-off-by: Derek Carr <decarr@redhat.com> * test(k8s): scope pod deletion check to v1alpha1 Signed-off-by: Derek Carr <decarr@redhat.com> * fix(k8s): address workspace namespace review findings Signed-off-by: Derek Carr <decarr@redhat.com> --------- Signed-off-by: Derek Carr <decarr@redhat.com> |
||
|
|
ae40cf6744 |
fix(gateway): respect OPENSHELL_BIND_ADDRESS in dev task (#2756)
This is convenient when you want to run a local gateway pointed at a remote compute driver so that the supervisor can reach across the network to the gateway which is listening on 0.0.0.0. Signed-off-by: Kris Hicks <khicks@nvidia.com> |
||
|
|
bdabb54cb3 |
fix(security): authenticate extension services (#2638)
* fix(security): authenticate extension services Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * feat(extension-core): verify gateway JWTs Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): keep inbound verification external This should become an extension SDK package. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(security): harden the extension authentication contract Follow-up hardening on the alpha extension authentication mechanism. Claim contract: - Extension tokens carry an explicit `typ` of `openshell-ext+jwt`. They share a signing key with sandbox-to-gateway admission tokens and were otherwise separated by audience alone, so a verifier that neglects to check `aud` could accept a gateway credential. The header is a second, independent discriminator. - Publish OIDC-shaped discovery at `/.well-known/openid-configuration` so a service configured with only the gateway URL can learn the exact expected issuer and the JWKS location. It is shaped, not compliant: `issuer` is the gateway identity, not the serving URL. Audience agreement: - `MiddlewareManifest` and `InterceptorManifest` gain `expected_audience`. The audience is otherwise configured independently on each side of the boundary, where a mismatch surfaces only as an opaque authentication failure on every call. OpenShell now compares the two and fails at startup. An empty field keeps the check off for existing services. Compatibility: - Add `allow_insecure_transport` per registration. Enabling gateway JWT signing previously made any plaintext endpoint a hard startup failure, including the endpoint form used in our own documentation. The opt-out attaches no credential, is refused by the gateway if a supervisor asks for one, and warns at every startup. - Make the transport requirement kind-aware. A middleware endpoint must be reachable from every sandbox supervisor, so only interceptors may use a gateway-local Unix socket. Credential lifecycle: - Replace the process-global slot map with a supervisor-owned `ExtensionCredentialStore` shared explicitly across the gateway connections the supervisor opens, removing test-order coupling. - Rotate only when a credential is missing or has passed four fifths of its lifetime. Configuration polling ran every ten seconds against fifteen-minute credentials, so each poll re-ran gateway effective-policy resolution and re-minted the gateway token. - Bound credential minting per sandbox, since each request resolves the caller's effective policy. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs: record alpha extension authentication in RFC appendices Restore the RFC 0009 and 0010 bodies to their accepted text and move every extension-authentication update into appendices instead. An RFC records a decision at a point in time; superseding detail belongs alongside it rather than rewritten into it. RFC 0009's appendix carries the shared contract: claims, authorization, key distribution, the `allow_insecure_transport` replacement for the body's `allow_insecure`, and residual risks. RFC 0010's records only what differs for interceptors and links to it. The existing protocol-extensions appendix, which parked the phase 2 transport question, now points forward to what was built. Also document the audience handshake, the discovery endpoint, the `typ` requirement, and `jti` replay guidance in the extensibility and gateway configuration pages, and correct the middleware transport guidance: middleware endpoints must be reachable from sandbox supervisors, so Unix sockets are not an option there. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * refactor(extension-core): abstract extension server trust Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * docs(core): update middleware manifest example Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): preserve unsigned gateway compatibility Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> * fix(extension-auth): reject cross-domain token replay Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> --------- Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com> |