* test(tmachine): add K3s conformance scenario
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* refactor(tmachine): use Helm values file for K3s installer
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(tmachine): run K3s conformance in integration jobs
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(tmachine): verify installer scripts and document version baseline
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
---------
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* test(podman): move podman_preflight into driver-podman integration tests
podman_preflight verifies that openshell-driver-podman fails fast when
its Podman socket is unreachable. It only needs the standalone driver
binary, not a gateway, so it never fit the gateway-backed e2e-podman
harness it lived under and never ran anywhere in CI.
Move it into crates/openshell-driver-podman/tests/ as a plain Cargo
integration test. It now runs via the existing required workspace test
job with no special mise task, workflow step, or coverage exception.
Signed-off-by: politerealism <burdcat17@gmail.com>
* test(podman): make preflight diagnostics portable
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
* fix(policy): propose rules for unknown DNS hosts
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* docs(policy): clarify synthetic DNS use across protocols
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): harden unknown-host DNS observations
- Emit the policy_dns_ineligible denial for every unknown name and
report observation staging failures as DNS failure events.
- Refuse unknown names during fail-closed quarantine and after the
observation budget, now a quarter of each address family's pool.
- Pin transparent TCP to the mapping of the deciding policy generation
so a reload between DNS and authorization fails closed.
- Stop Docker workloads from inheriting host DNS search domains, which
let the first expanded short name claim an observation address.
- Share mechanistic draft polling in conformance, register
new-hostname-proposal in the installed suite, and update docs.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(policy): build policy DNS proxy tests on every target
The proxy tests name PolicyEndpointId, which proxy.rs imported only on
Linux, so the macOS test build failed. Import it for test builds too.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* fix(policy): name DNS queries and mapped hosts in OCSF denials
DNS denial and failure events attached port 53 to the queried name,
which read as a connection to that host. They now carry only the name.
Transparent TCP denials for a policy DNS address show the mapped
hostname and keep the synthetic address in dst_endpoint.ip.
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
---------
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(podman): run driver-podman userns suite against rootful Podman too
The driver-specific-integration job only ran the driver-podman testsuite
(default/auto/keep-id/private userns reference checks) against
fedora-podman-rootless, leaving rootful behavior for this scenario
unverified even though the compute driver auto-detects and explicitly
supports rootful Podman.
The default-userns-baseline and userns-profile playbooks hard-asserted a
rootless tmachine gateway user, so pointing them at a rootful environment
would have failed that assertion immediately rather than exercising
anything. They now detect rootful vs. rootless via the existing
tmachine_container_runtime role and branch the reference-capture user
accordingly, while keeping the captured reference file itself owned by
tmachine, since the archived test binary that reads it back always runs
unprivileged as tmachine regardless of daemon mode.
Signed-off-by: politerealism <burdcat17@gmail.com>
* test(podman): add real-daemon coverage for resource limits and daemon failure
Neither the Podman driver's resource-limit enforcement nor its behavior
when the Podman daemon is unreachable had any test coverage against a
real daemon; both were only exercised through unit tests against a
mocked Podman client.
podman_resource_limits.rs creates a sandbox with --cpu/--memory flags and
reads /sys/fs/cgroup/memory.max and cpu.max from inside the sandbox
itself, verifying the limit is actually enforced rather than just echoed
back by the template API. Expected values are cross-checked against the
driver's own parse_cpu_to_microseconds/parse_memory_to_bytes and against
a real local `podman run --cpus/--memory` container.
podman_preflight.rs spawns the standalone openshell-driver-podman binary
against a guaranteed-nonexistent Podman socket and asserts it exits
non-zero within its bounded retry window with an actionable error naming
the socket path, rather than hanging or failing silently.
Signed-off-by: politerealism <burdcat17@gmail.com>
* test(podman): make rootful userns and cgroup checks pass
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(podman): match lifecycle containers by isolation role label
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(podman): accept non-expiring bootstrap tokens
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
---------
Signed-off-by: politerealism <burdcat17@gmail.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* feat(sandbox): default to official Alpine sandbox image
default_sandbox_image() now returns docker.io/library/alpine:3.22, a generic
version-qualified official image, so a fresh install no longer depends on the
community sandbox image catalog. All compute drivers (docker, podman,
kubernetes, vm) inherit this fallback.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(deploy): default deployment configs to the official Alpine sandbox image
Update the shared gateway default_image, Helm chart values, the standalone
Kubernetes manifest, and the dev gateway task scripts to use
docker.io/library/alpine:3.22 instead of the community base image, consistent
with default_sandbox_image(). GPU e2e image-build base is left unchanged (CUDA
needs a glibc base).
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* feat(driver): default to numeric non-root identity for USER-less images
With the default sandbox image now Alpine, images that declare no OCI USER
must start instead of being rejected. When the image declares no USER and
the policy requests none, the Podman and Docker drivers now supply a numeric
non-root identity (DEFAULT_SANDBOX_UID/GID = 1000) instead of rejecting,
matching the numeric-identity behavior of the Kubernetes and VM drivers. The
supervisor's resolved-identity path runs the sandbox as a synthesized
non-root account without the account existing in the image. Images that
declare a USER keep the OCI resolution path unchanged.
Part of #3116.
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* test(conformance): use Alpine workload image
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* refactor(policy): drop community image /app path from default policy
The restrictive default policy granted read-only access to /app, a directory
that only existed in the community base image. A generic Alpine default has no
/app, so remove it. Landlock best-effort already ignores absent paths; this
just stops advertising a community-specific layout in the default.
Part of #3116.
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
* docs(config): document Alpine default images
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): report early sandbox termination
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(podman): initialize rootless workspace ownership
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* fix(sandbox): qualify NVIDIA Ubuntu default
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(podman): initialize rootful default workspace
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* feat(sftp): add native sandbox adapter
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): gate runtime helper support to Linux
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): support standard OpenSSH file operations
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(sftp): harden rename and special file handling
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* refactor(runtime): remove community image dependencies
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* test(e2e): build provider readiness tool fixture
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* fix(e2e): use a dedicated Noble fixture for Docker tests
Signed-off-by: Evan Lezar <elezar@nvidia.com>
---------
Signed-off-by: Akram
Signed-off-by: Akram <akram.benaissi@gmail.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
* feat(cli): promote profile commands to top level
Add profile discovery and management commands with shared handlers for
the existing provider entry points. List a flat catalog across scopes
and follow continuation tokens through full and short pages.
Describe metadata, credentials, endpoints, TLS inspection, and MCP access
settings while preserving complete JSON/YAML definitions. Cover parser
equivalence, scope forwarding, pagination, and inspection settings with
focused unit and compiled-CLI integration tests.
Update docs, public skills, examples, and E2E command invocations.
Refs #2588
Signed-off-by: Shiju <shiju@nvidia.com>
* fix(cli): remove redundant workspace selector qualification
Use the imported WorkspaceSelector in the provider integration helper so
the target passes Clippy with warnings denied.
Signed-off-by: Shiju <shiju@nvidia.com>
---------
Signed-off-by: Shiju <shiju@nvidia.com>
## Summary
- Add `Provider` entity for managing 3p deps from a sandbox
- Add provider CRUD API/server persistence and new CLI workflows (`nav provider create/get/list/update/delete`), including `--from-existing` laptop discovery.
- Integrate providers into sandbox create flow: infer from command (`-- claude`), support repeatable `--provider <type>`, prompt before auto-create, and allow manual in-sandbox setup.
- Add a dedicated `navigator-providers` crate with per-provider modules and mockable discovery test helpers.
## Key UX Changes
- `nav sandbox create --provider gitlab -- claude`
- Missing provider prompt now asks before creating from local state.
- `nav provider list --names` for scripting/cleanup.
## Test Plan
- `mise run cluster:deploy`
- `mise run test:e2e:sandbox`
- `mise run pre-commit`
Closes#19Closes#22Closes#11
Closes#13
## Summary
- add Python sandbox execution APIs for command and callable workflows
- consolidate sandbox policy fixtures and expand e2e test coverage for policy and Python exec paths
- update CI/build config and images for sandbox e2e execution dependencies
## Test Plan
- mise run pre-commit