* fix(server): retry runtime identity persistence on sandbox create
With multiple gateway replicas, another replica can update a new sandbox
record between the create path's read and its compare-and-swap write of the
runtime identity. The create path made one attempt, so the conflict failed
the request and deleted the backend sandbox. Persist the identity through
the same retrying helper that start and startup recovery use, which rereads
the record and retries while the sandbox stays in the same generation and a
Provisioning or Ready phase.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* ci(e2e): run Kubernetes HA tests one at a time
The two HA tests scale and roll the shared gateway Deployment. Their
in-process lock does not apply under nextest, which runs each test in its
own process, so one test could delete a gateway pod while the other was
executing through it. Put both tests in a nextest test group limited to one
thread in the e2e-kubernetes profile. The override matches test names
because a binary() filter fails in workspaces that lack the HA test binary.
Signed-off-by: Kris Hicks <khicks@nvidia.com>
---------
Signed-off-by: Kris Hicks <khicks@nvidia.com>
* feat(e2e): run kubernetes suite on cargo-nextest with JUnit/HTML reports
Switch e2e:kubernetes (and all its variants) from `cargo test` to
`cargo nextest run` for per-test process isolation and output consistent
with the other nextest-based CI runs.
- Add a dedicated `e2e-kubernetes` nextest profile with a JUnit report and a
generous slow-timeout (60s flag, 5-min terminate) suited to live-cluster
tests; kept separate from `ci` so its JUnit path and timeouts don't affect
the workspace run.
- Pin `--target-dir` for the run so the profile's relative JUnit path resolves
to the repo-root results/ regardless of any inherited CARGO_TARGET_DIR
(nextest ignores absolute JUnit paths).
- Render the JUnit XML to a standalone HTML report via xsltproc and a committed
XSLT stylesheet (best-effort; never masks the test exit code).
- Name each report via `OPENSHELL_E2E_REPORT_NAME` (default `e2e-kubernetes`),
used verbatim for both the `results/<name>.{xml,html}` filenames and the HTML
heading. Tasks that invoke the script multiple times in one run set a distinct
name per invocation so the reports no longer clobber the single fixed path:
the credential-driver runs write results/e2e-kubernetes-secrets.xml and
-vault.xml, and e2e:kubernetes:agent-sandbox-versions writes
results/e2e-kubernetes-agent-sandbox-v1beta1.xml and -v1alpha1.xml.
- Declare cargo-nextest in mise [tools] so the task runs without the Nix shell.
- Ignore the results/ output directory.
The results/ reports do not leak information. They are gitignored and no
workflow uploads them as artifacts, so they stay on the ephemeral CI runner
and are discarded when it is torn down. The HTML template renders only test
names, status, timings, and failure messages (no captured stdout/stderr).
Moving from `cargo test -- --nocapture` to nextest's captured, failure-only
output also reduces what lands in the retained, viewable console logs.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
* chore(e2e): revert per-lane report names for agent-sandbox-versions
The agent-sandbox-versions task runs two lanes sequentially against the
same cluster: v0.5.0 (v1beta1 storage version) then v0.4.6 (v1alpha1).
On a reused cluster the second lane fails when kubectl applies the older
CRD, because Kubernetes refuses to drop v1beta1 from spec.versions while
it remains in status.storedVersions (the storage-version downgrade
guardrail). This is a pre-existing issue with the v0.4.6 lane, unrelated
to the nextest reporting work.
The per-lane OPENSHELL_E2E_REPORT_NAME additions do not address that
downgrade failure, so revert them to keep this PR scoped to the nextest
change. Agent Sandbox 0.4.x is also superseded (1.0.0 is published);
dropping or bumping the v1alpha1 lane is left as a follow-up.
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
---------
Signed-off-by: Jorge Garcia Oncins <jgarciao@redhat.com>
* ci(branch-checks): run Rust checks in Nix shells
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(branch-checks): run Rust tests with nextest
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(branch-checks): cache Rust workspace artifacts
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
* ci(branch-checks): run cargo-deny in Nix shell
Signed-off-by: Simon Scatton <sscatton@nvidia.com>
---------
Signed-off-by: Simon Scatton <sscatton@nvidia.com>