* fix(test): establish a writable previous-release upgrade baseline
* fix(ci): reserve capacity for durable admin fixtures
* fix(ci): group durable IAM state fixtures by resource needs
* ci: run E2E doctests with the E2E dependency graph
* fix(ci): separate fixture startup from transport deadlines
* fix(ci): bound pagination after seeding and revisit restored copies
* fix(ci): make filesystem fixture timing deterministic
* fix(ecstore): avoid metadata lock reentry during internal mutations
* test: align recovery fixtures with durable ownership contracts
## Related Issues
N/A
## Summary of Changes
Reject system-bucket incarnation lookups before entering the pool metadata owner. Keep each selected scanner backlog publication cohort in one task so waiter cancellation cannot abandon its remaining serialized conditional writes. Bind repair and replay fixtures to persisted bucket incarnations.
## Verification
Head f6d44603fc received an approval from houseme in review 5365588011. This merge does not add a local runtime validation claim. Full main CI remains a separate publication gate.
## Impact
System metadata writes avoid recursive pool locking. Publication retains per-replica conditional writes and stops a cancelled caller from advancing to its next publication phase. Fixture changes supply the identities required by existing repair admission rules.
## Additional Notes
Reverting this change restores the previous behavior.
## Related Issues
Follow-up to #8233.
## Summary of Changes
Registration runtime fixtures timed out in the workspace CI lane while the same cases passed in the feature lanes. Reserve nextest capacity for the exact `connect_registration` binary, as already done for related inventory and drive fixtures. Keep its existing deadlines, internal concurrency, assertions and zero-retry policy; report the last watch status on failure.
Rolling upgrades could pass `ListBuckets` readiness while restarted peers still lacked write quorum. Before each mixed-version phase, probe writes through every node outside the asserted workload prefixes. A shared 30-second deadline includes requests and sleeps; only HTTP 503 with `ServiceUnavailable` is retryable, with SDK retries disabled for these probes. The actual compatibility writes, reads, multipart operations and listing assertions remain unchanged.
Add four fast regression tests to the existing PR smoke profile, with matching exclusion from the full profile. No new workflow or job is introduced.
## Verification
- `cargo nextest run --locked --profile ci -p rustfs --lib --test connect_registration --test-threads 4 -E 'binary(/^connect_registration$/) | test(=connect::diagnostics::trace_runtime::tests::local_runtime_rejects_non_private_state)' --no-tests fail --status-level pass --final-status-level fail` passed 23/23 twice in 5.088s and 5.097s: 22 macOS registration tests plus one unrelated control. JUnit intervals confirm capacity reservation; no retries or test-process leaks occurred. The Linux-only inventory case remains for CI.
- `cargo nextest run -p e2e_test --lib --profile ci -E 'test(upgrade_write_readiness_tests)' --no-fail-fast --no-tests fail` passed 4/4 twice in 1.068s and 1.072s, with zero retries. Regressions cover metadata readiness followed by write unavailability, recovery through every writer, immediate permanent-error failure despite an incoming SDK retry configuration, and deadlines for repeated 503s and stalled requests. The original failure is recorded in [the mixed-version upgrade job](https://github.com/rustfs/rustfs/actions/runs/36569700716/job/109415806473).
- `cargo fmt --all --check`, `git diff --check`, `python3 scripts/check_test_wiring.py` and compiled smoke/full membership checks passed. Smoke membership changes from 188 to 192 by adding exactly these four tests; full membership is unchanged. The expected Linux smoke digest was derived from the actual prior Linux listing plus those four platform-independent additions and still requires confirmation by this PR's CI.
Local verification covers the exact source committed in `9b2ea836317ed035a91d1fc9fcf725f70c3098e2` on main `380e98a42cb4fcd0994fed79b30c2c7605deb0bc`. An independent final-diff correctness and reliability review found no findings. Fresh Linux workspace and real mixed-version upgrade runs are required before treating the remediation as fully verified; local fake-target tests do not establish that result.
## Impact
Test scheduling and readiness only; no production behavior, API, dependency, test deadline or compatibility assertion changes. Reserving capacity serializes registration fixture processes within a nextest run. The bounded readiness probes may add startup time while peer write health converges; permanent errors still fail immediately.
## Additional Notes
Rollback by reverting this PR. Existing CI restructuring from #8233 is independent of these follow-up fixes.
* fix(ci): reduce duplicate work and preserve reliable test failures
* fix(ci): retain protocol evidence and repair stale test fixtures
* test(connect): honor parent deadline during API fixture readiness
* fix(ci): reserve IO capacity for state writer proofs
* test(connect): align RPC fixtures with service capture contracts
* test(connect): cover pinned service capture failures
* feat(connect): capture offline network performance in server
* fix(connect): use storage facade in RPC test
* fix(connect): keep RPC test body behind storage facade
A bucket sweep that ends mixed clears its position and records the finishing cycle's plan as started. The next cycle requests a new plan whenever the bucket was written in between, so a fresh sweep inherited a stale started plan and ended mixed again. On a continuously written bucket no sweep could ever certify and the census kept the old root.
Start a fresh verification sweep under the requested plan when no durable position remains. Resumed sweeps keep their started plan, so a clean tail still cannot certify an old prefix.
Refs #7108
Commit 151103a609 added the field to
StorageReadinessDetails but did not update two test initializers in
readiness.rs, causing E0063 compile errors in --all-targets builds.
Two Connect tests fail on main in every Test and Lint lane.
Heartbeat: #8155 replaced "today's capabilities minus
profile.memory.service@1" with "the pre-health set minus it". That drops
the release that advertised health.check.service@1 but not yet in-service
memory profiling, so a pending heartbeat persisted by that release now
fails validation and the runtime stops. Accept that release as its own
frozen list.
Top disk/net: execute_*_job sets max_duration_millis to the requested
duration, then the capture compared the measured sleep against it. The
timer only overshoots, so a job that ran exactly as requested was
rejected with LimitExceeded whenever the overshoot reached 1 ms. Report
the authorized window, as Top API already does; validate_capture bounds
it by the limit.
* fix(ci): restore mainline and scheduled test reliability
* fix(ci): provide GitHub CLI for CPU acceptance
* ci: provide Docker for CPU service acceptance
* ci: provision Python and Docker for OIDC validation
* ci: restore hosted runners for Docker validation
* ci: use verified MinIO release packages for interop
* ci: preserve host ownership of MinIO fixtures
* fix(ci): correct diagnostic limits and isolate startup checks
* test(connect): include object CLI failure details
* test(readiness): initialize unavailable drive diagnostics
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
The health diagnostics module imported directly from
crate::storage::storage_api, bypassing the architecture migration
guardrail. Add a connect facade module to rustfs/src/storage_api.rs
and redirect the imports. Also fix a clippy::redundant-guards lint
in the coarse-flags match arm.
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
* fix(storage): weight automatic multipart admission by part size
* fix(ci): restore filesystem runner capabilities and typos dependency
* fix(ci): make release guard portable and spell out part variables
* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* feat(connect): sample memory within the running service
* test(connect): add official memory service acceptance
* test(connect): consume the final memory job without cloning
Startup finalizes IAM before init_notification_runtime publishes the notification system, and IAM finalization is what starts the fleet capability probe. The first probe pass therefore always failed closed and then slept the full ten-second probe interval. On a single node that left every durable capability, including the durable hard quota fence, unavailable for about ten seconds after /health already reported ok, so SetBucketQuota answered 503 durable quota capability is not confirmed across the cluster during that window.
Publishing the notification system now wakes the probe immediately through a Notify permit, with a 100ms bootstrap poll as the fallback for a wakeup that races the availability check. The probe keeps failing closed while the system is absent and logs the wait once. The four probe futures no longer carry an unreachable notification-system-unavailable arm, and the seven proof slots are revoked through one helper.
Fixes#8014
Pre-release RPM versions replaced the SemVer `-` with `_`, which rpm treats as an ordinary segment separator, so `1.0.0_rc.5` compared as newer than `1.0.0` and a host with an rc RPM installed could not `dnf upgrade` to the GA package. DEB previews had the mirror-image problem: only the first `-` became `~`, so `1.0.0~rc.5-preview.2` carried `preview.2` as its Debian revision and sorted above `1.0.0~rc.5`.
Both formats now spell every pre-release separator as `~`, which dpkg and rpm (>= 4.10) both treat as "sorts before anything". Verified with rpm 4.16 (AlmaLinux 9) and dpkg 1.21: `1.0.0~rc.5 < 1.0.0`, `1.0.0~rc.5~preview.2 < 1.0.0~rc.5 < 1.0.0~rc.6`, alpha < beta < rc, and `rc.9 < rc.10`; fpm 1.18 passes `~` through into both package headers, and `dnf upgrade` from a `~rc.5` RPM to the GA RPM succeeds where the `_rc.5` one refused. The test now pins the full ordering contract under both package managers and fails against the previous spelling.
The release-checksum step already normalizes `~` to `.` for every uploaded asset, so RPM assets need no further handling. Hosts that already installed an `_rc`/`_beta` RPM need a one-time `dnf install rustfs-1.0.0` or `rpm -Uvh --oldpackage` to reach GA.
Fixes#8012
The legacy bucket metadata migration writes the `.bucket-incarnation`
sidecar before `.metadata.bin`. A crash or a lost namespace lease between
the two writes left the bucket with a sidecar and no metadata, and every
later load failed closed with "bucket incarnation sidecar exists without
bucket metadata", so the bucket answered 500 to every request on every
node with no repair path (rustfs/rustfs#8003).
Load the sidecar-only state as a legacy bucket so the migration runs
again. The migration and the force-create path re-read the sidecar under
the transaction lock and keep its incarnation when persisting the
metadata, so the retry never replaces an identity other nodes fenced on.
A retired incarnation is never adopted: it is residue of a deleted
bucket, and re-publishing it would let heal reclaim the new objects.
Consolidate the operator entry point for planning, storage classes and
EC:0 consequences, expansion, rebalance, decommission, heal, drive
replacement, restart recovery, and the rc CLI mapping. Every claim was
checked against the current handlers, storage-class validation, quorum
arithmetic, and the rustfs/cli admin reference; the runbook links the
normative contracts instead of restating them.
Refs rustfs/backlog#2607, rustfs/backlog#2608
fix(ci): align security workflow test with the performance suite's own concurrency group
#7984 moved rustfs-performance-test.yml to the rustfs-performance-suite group but left test_all_suites_hold_the_shared_lock_for_manual_and_chain_runs asserting the shared functional group, so Quick Checks (make script-tests) fails on main and on every PR. Expect the dedicated group for the performance suite and drop the workflow comment that still described the shared lock.
* fix(ecstore): hide ancestors of delete residue in prefix listings
Since #7342 the never-versioned fast path hides a directory whose children are all deleted-version data dirs, but only that one level: the date-style ancestors above a deleted key (`metrics/<stream>/2026/08/28/23/`) still surfaced as empty folders that nothing could remove, which is what #6898 keeps reporting on rc.6 and 1.0.0. The residue probe now walks down through metadata-less ancestors depth-first, ending at the first file met outside a UUID data dir, with a total read budget so a large residue tree surfaces and is hidden one level down instead of costing an unbounded walk. An empty delimiter listing of the ancestor then reclaims the whole committed residue tree in one pass.
Refs #6898
* fix(ecstore): sample before full reads in the residue probe and purge orphan subtrees independently
Adversarial pass 1 findings: the probe read a wide directory in full at every level of a genuine prefix once its 8-entry batch was full, so a leaf holding thousands of objects was materialised per listed prefix; the probe now descends into the sampled children first and only reads the remainder once every sample proved unlistable, and a test pins one bounded read per level and zero complete reads on a wide genuine prefix. Hiding ancestors also made purge_orphan_dir_object fail closed on a whole ancestor when committed residue shared it with residue an older build left without markers, which 1.0.0 users could still reclaim by browsing one level lower; the purge now blocks only the ancestor chain of unpurgeable data and reclaims purgeable sibling subtrees. A UUID directory holding a subdirectory surfaces again as before.
Refs #6898
* fix(ecstore): back off repeated orphan purge scans of unpurgeable trees
Adversarial pass 2: with ancestors hidden, every empty listing of a phantom prefix scanned the whole residue subtree on every disk even when nothing under it can be purged (no committed marker), and the scan no longer stops at the first blocked directory. Remember such prefixes per set for 60 s and skip the rescan. Build committed residue file paths like directory paths so the blocked filter matches under keys containing repeated slashes.
Refs #6898