## Related Issues
Related to rustfs/backlog#2701.
## Summary of Changes
Route scanner checkpoint cycle and leader validation through the global store while retaining the owning set for cache persistence, CAS revisions and publication admission.
## Verification
Two independent final-diff source reviews found no issues across correctness, concurrency and durability, test coverage, compatibility, performance and simplicity on head `8b8fe51d092090b053f552ae283960e2e306be33`. Root approval `5373624714` is bound to that exact head. Regression tests cover real two-pool routing, stale fences, post-save rejection and CAS conflicts; their reported local execution belongs to the PR author, not this merge operation. Current required CI remains pending, and this authorized admin squash does not establish CI or runtime acceptance.
## Impact
Restores checkpoint progress when global cycle and leader state differ from a set-scoped view. No format, retry, timeout, assertion or scanner-policy changes are introduced by this diff. The three prior main scanner failures remain unproved repaired.
## Additional Notes
Full validation must run on the resulting exact main revision. Reverting this patch restores the earlier set-scoped fence lookup and its checkpoint rejection behavior.
* fix(test): establish a writable previous-release upgrade baseline
* fix(ci): reserve capacity for durable admin fixtures
* fix(ci): group durable IAM state fixtures by resource needs
* ci: run E2E doctests with the E2E dependency graph
* fix(ci): separate fixture startup from transport deadlines
* fix(ci): bound pagination after seeding and revisit restored copies
* fix(ci): make filesystem fixture timing deterministic
* fix(ecstore): avoid metadata lock reentry during internal mutations
* test: align recovery fixtures with durable ownership contracts
## Related Issues
N/A
## Summary of Changes
Reject system-bucket incarnation lookups before entering the pool metadata owner. Keep each selected scanner backlog publication cohort in one task so waiter cancellation cannot abandon its remaining serialized conditional writes. Bind repair and replay fixtures to persisted bucket incarnations.
## Verification
Head f6d44603fc received an approval from houseme in review 5365588011. This merge does not add a local runtime validation claim. Full main CI remains a separate publication gate.
## Impact
System metadata writes avoid recursive pool locking. Publication retains per-replica conditional writes and stops a cancelled caller from advancing to its next publication phase. Fixture changes supply the identities required by existing repair admission rules.
## Additional Notes
Reverting this change restores the previous behavior.
## Related Issues
Related to rustfs/backlog#2688.
## Summary of Changes
Reclaim metadata-less orphan directory trees after a complete, first-page, empty recursive root listing in an authoritative never-versioned bucket. Both listing implementations use the same fail-closed cleanup decision.
## Verification
Head b4a9120173 received an approval from loverustfs in review 5365460152. This merge does not add a local runtime validation claim. Full main CI remains a separate publication gate.
## Impact
Empty recursive root listings can remove otherwise unaddressable physical residue. Versioned buckets and populated listings remain excluded; unreadable disks, object metadata, unknown files, and uncommitted data prevent cleanup.
## Additional Notes
Reverting this change restores the previous orphan-directory behavior.
## Related Issues
Resolves the corrected lifecycle fixture findings in PR #8254.
## Summary of Changes
Preserve separate incarnation-bound durable MRF responsibilities and their lifecycle audit records, and align replay fixtures with their checkpoint identities.
## Verification
Two independent source reviews and changed-delta reviews are complete. The final review approved 607a0b9bb0 with no remaining supported findings. Local runtime claims were not independently reproduced for this pull request.
## Impact
Keeps storage generation fences and retained responsibility semantics. Test-capacity reservations preserve existing deadlines and assertions.
## Additional Notes
Squash merge of the currently approved fix under the authorized CI-bypass exception. Main CI and release acceptance remain required.
## Related Issues
Follow-up to #8233.
## Summary of Changes
Registration runtime fixtures timed out in the workspace CI lane while the same cases passed in the feature lanes. Reserve nextest capacity for the exact `connect_registration` binary, as already done for related inventory and drive fixtures. Keep its existing deadlines, internal concurrency, assertions and zero-retry policy; report the last watch status on failure.
Rolling upgrades could pass `ListBuckets` readiness while restarted peers still lacked write quorum. Before each mixed-version phase, probe writes through every node outside the asserted workload prefixes. A shared 30-second deadline includes requests and sleeps; only HTTP 503 with `ServiceUnavailable` is retryable, with SDK retries disabled for these probes. The actual compatibility writes, reads, multipart operations and listing assertions remain unchanged.
Add four fast regression tests to the existing PR smoke profile, with matching exclusion from the full profile. No new workflow or job is introduced.
## Verification
- `cargo nextest run --locked --profile ci -p rustfs --lib --test connect_registration --test-threads 4 -E 'binary(/^connect_registration$/) | test(=connect::diagnostics::trace_runtime::tests::local_runtime_rejects_non_private_state)' --no-tests fail --status-level pass --final-status-level fail` passed 23/23 twice in 5.088s and 5.097s: 22 macOS registration tests plus one unrelated control. JUnit intervals confirm capacity reservation; no retries or test-process leaks occurred. The Linux-only inventory case remains for CI.
- `cargo nextest run -p e2e_test --lib --profile ci -E 'test(upgrade_write_readiness_tests)' --no-fail-fast --no-tests fail` passed 4/4 twice in 1.068s and 1.072s, with zero retries. Regressions cover metadata readiness followed by write unavailability, recovery through every writer, immediate permanent-error failure despite an incoming SDK retry configuration, and deadlines for repeated 503s and stalled requests. The original failure is recorded in [the mixed-version upgrade job](https://github.com/rustfs/rustfs/actions/runs/36569700716/job/109415806473).
- `cargo fmt --all --check`, `git diff --check`, `python3 scripts/check_test_wiring.py` and compiled smoke/full membership checks passed. Smoke membership changes from 188 to 192 by adding exactly these four tests; full membership is unchanged. The expected Linux smoke digest was derived from the actual prior Linux listing plus those four platform-independent additions and still requires confirmation by this PR's CI.
Local verification covers the exact source committed in `9b2ea836317ed035a91d1fc9fcf725f70c3098e2` on main `380e98a42cb4fcd0994fed79b30c2c7605deb0bc`. An independent final-diff correctness and reliability review found no findings. Fresh Linux workspace and real mixed-version upgrade runs are required before treating the remediation as fully verified; local fake-target tests do not establish that result.
## Impact
Test scheduling and readiness only; no production behavior, API, dependency, test deadline or compatibility assertion changes. Reserving capacity serializes registration fixture processes within a nextest run. The bounded readiness probes may add startup time while peer write health converges; permanent errors still fail immediately.
## Additional Notes
Rollback by reverting this PR. Existing CI restructuring from #8233 is independent of these follow-up fixes.
* fix(ci): reduce duplicate work and preserve reliable test failures
* fix(ci): retain protocol evidence and repair stale test fixtures
* test(connect): honor parent deadline during API fixture readiness
* fix(ci): reserve IO capacity for state writer proofs
* test(connect): align RPC fixtures with service capture contracts
* test(connect): cover pinned service capture failures
A bucket sweep that ends mixed clears its position and records the finishing cycle's plan as started. The next cycle requests a new plan whenever the bucket was written in between, so a fresh sweep inherited a stale started plan and ended mixed again. On a continuously written bucket no sweep could ever certify and the census kept the old root.
Start a fresh verification sweep under the requested plan when no durable position remains. Resumed sweeps keep their started plan, so a clean tail still cannot certify an old prefix.
Refs #7108
## Related Issues
rustfs/rustfs#8217 and rustfs/backlog#2683.
## Summary of Changes
Reviewed 539ed275db against d2175d1e1e. No findings. Capable callers requesting missing-path reporting recover explicit pre-output FileNotFound or VolumeNotFound errors; partial-output and other failures continue to fail the stream.
## Verification
Two independent source reviews completed across all lenses below; the frozen diff matches Git and passes diff whitespace checks. Traced the authenticated handler, bounded first-chunk preflight, cancellation ownership, HTTP status/token classification and quorum error handling. The new tests cover ordinary streaming, typed missing errors, byte preservation and mixed missing/I/O failures.
Author-reported focused tests and Docker reproduction were not rerun or independently audited in this review. The reported Docker comparison predates the final report_notfound-only gate; its scope is not distributed acceptance of the final head.
## Impact
Correctness: no findings; missing errors remain typed only before output. Security/trust: no findings; authentication, body digest and operation/status/token restrictions remain. Compatibility: no findings; legacy clean EOF remains, with the documented old-server limitation. Concurrency/durability: no findings; receiver ownership cancels an abandoned producer. Simplicity: no findings. Coverage: no blocking gap identified. Performance: no findings; preflight retains one bounded chunk and ordinary report_notfound=false streams start immediately.
## Additional Notes
This review does not establish that the patch fixes any current main CI failure. Full main CI and release acceptance remain separate gates.
## Related Issues
rustfs/backlog#2682 and rustfs/rustfs#8192.
## Summary of Changes
Reviewed ca970f55ec against d2175d1e1e. No blocking code findings. The hold requires a completed healthy legacy check and an exact incarnation, lease, object, version and scope; it preserves the durable journal and rechecks on restart.
## Verification
Two independent source reviews completed, covering the lenses below. The frozen diff matches Git and passes diff whitespace checks. Reported Cargo and Docker results were not rerun or independently audited in this review.
One factual correction to the PR description: `legacy_sigkill_replay_repairs_without_releasing_unverified_responsibility` uses the default four-member fixture, with three replicas before restoration and four afterward (`mrf_partial_write_test.rs:559,567`). That named regression exercises the production manager path, but is not EC12+4. Other tests in the file use sixteen disks.
## Impact
Correctness: no findings; proofless legacy health never becomes a verified receipt. Security/trust: no findings; exact identity fences remain. Compatibility: no findings; journal encoding remains unchanged. Concurrency/durability: no findings; replay retains one checkpointed owner and replacement generations become retryable. Simplicity: no findings. Coverage: no blocking gap identified, with the test-scope correction above. Performance: no findings; held entries leave the retry index without adding per-admission queue scans.
## Additional Notes
The current main journal failure concerns DecodeFailure, which follows the separate ECDecode task path. This review does not establish that this PR fixes that failure or the scanner-cycle failure, and does not establish a passing main CI or release gate.
* fix(iam): require an explicit permission for force-delete
A force-delete header no longer inherits s3:* or consoleAdmin. Bucket
force-delete requires s3:ForceDeleteBucket whenever the header is present,
and recursive object force-delete requires s3:ForceDeleteObject. A plain
delete keeps the existing checks.
Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
* test(e2e): keep force-delete header names static
The bucket force-delete helper must pass a static header name into the
SDK request mutator.
Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
* test(e2e): move the force-delete header into the request mutator
The SDK request customizer requires a static header name owned by the
closure.
Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
* fix(iam): keep force-delete out of NotAction grants
NotAction now uses plain wildcard matching, so NotAction "s3:*" still
excludes force-delete. An Allow statement grants s3:ForceDeleteObject or
s3:ForceDeleteBucket only when its Action list names the action; a
NotAction-only Allow never does. The rule applies to both IAM and bucket
policy statements.
Also build the invalid-header errors with S3Error::with_message to keep
the s3s footprint at its baseline, and fix a clippy single_match.
---------
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Record the upload id on the completed object so a lost-response retry
returns that object's ETag instead of NoSuchUpload, while a different
part list or a replaced object keeps the existing errors.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(ci): restore mainline and scheduled test reliability
* fix(ci): provide GitHub CLI for CPU acceptance
* ci: provide Docker for CPU service acceptance
* ci: provision Python and Docker for OIDC validation
* ci: restore hosted runners for Docker validation
* ci: use verified MinIO release packages for interop
* ci: preserve host ownership of MinIO fixtures
* fix(ci): correct diagnostic limits and isolate startup checks
* test(connect): include object CLI failure details
* test(readiness): initialize unavailable drive diagnostics
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
A restarted peer can accept a pooled connection and never send response
headers, so HttpReader::open waited past the client body timeout before
the body-stall timer or erasure hedge could run. Bound that header wait
by the stall timeout and retry the open once on a fresh connection.
After write quorum, MultiWriter still waited out the full disk stall
for a silent peer, which matches the client timeout. Give remaining
writers one second, then drop them so the caller returns.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(storage): weight automatic multipart admission by part size
* fix(ci): restore filesystem runner capabilities and typos dependency
* fix(ci): make release guard portable and spell out part variables
* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(ecstore): admit cold bucket metadata at read quorum
Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.
* fix(ecstore): fence bucket creation-commit persistence
Address review findings on the creation-commit proof.
Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.
Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.
* fix(ecstore): keep commit fence error message stable
The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.
* test(ecstore): reuse canonical Object Lock fixture in regression tests
The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>