100 Commits
Author SHA1 Message Date
Chris 50e6c907f9 fix(s3): bound multipart copy admission and HTTP write buffers (#8288) 2026-10-01 21:51:27 +08:00
Chris a33dd896e2 fix(ci): restore E2E membership and pagination timeouts (#8281) 2026-10-01 09:00:02 +08:00
overtrue 99a2af8687 docs(release): validate candidates on release branch 2026-10-01 08:37:30 +08:00
Chris 2806a80c91 fix: prevent metadata lock reentry and stabilize CI fixtures (#8257)
* fix(test): establish a writable previous-release upgrade baseline

* fix(ci): reserve capacity for durable admin fixtures

* fix(ci): group durable IAM state fixtures by resource needs

* ci: run E2E doctests with the E2E dependency graph

* fix(ci): separate fixture startup from transport deadlines

* fix(ci): bound pagination after seeding and revisit restored copies

* fix(ci): make filesystem fixture timing deterministic

* fix(ecstore): avoid metadata lock reentry during internal mutations

* test: align recovery fixtures with durable ownership contracts
2026-09-30 20:57:07 +08:00
Chris 8655c38f1d fix(rio): preserve bounded retries after peer EOF (#8272) 2026-09-30 20:18:28 +08:00
Chris 88d03d8199 fix(storage): prevent metadata deadlocks and abandoned writes (#8268)
## Related Issues

N/A

## Summary of Changes

Reject system-bucket incarnation lookups before entering the pool metadata owner. Keep each selected scanner backlog publication cohort in one task so waiter cancellation cannot abandon its remaining serialized conditional writes. Bind repair and replay fixtures to persisted bucket incarnations.

## Verification

Head f6d44603fc received an approval from houseme in review 5365588011. This merge does not add a local runtime validation claim. Full main CI remains a separate publication gate.

## Impact

System metadata writes avoid recursive pool locking. Publication retains per-replica conditional writes and stops a cancelled caller from advancing to its next publication phase. Fixture changes supply the identities required by existing repair admission rules.

## Additional Notes

Reverting this change restores the previous behavior.
2026-09-30 19:49:24 +08:00
Chris 099a408d8b test(ecstore): isolate dispatch shutdown recovery fixtures (#8265) 2026-09-30 15:56:48 +08:00
Chris 5c3a2ff277 fix(ci): reserve capacity for profile cancellation probe (#8259) 2026-09-30 15:46:20 +08:00
Chris bb04355a3d fix(heal): retry contended local metadata CAS (#8267) 2026-09-30 15:45:59 +08:00
Chris fab2d7e144 fix(scanner): serialize pause backlog replica publication (#8264) 2026-09-30 15:45:37 +08:00
Chris 6bc2e10bc1 fix(tier): keep recovery worker after startup reload failure (#8260) 2026-09-30 12:05:26 +08:00
Chris 3268c42e00 fix(ci): stabilize registration and rolling upgrade readiness (#8244)
## Related Issues

Follow-up to #8233.

## Summary of Changes

Registration runtime fixtures timed out in the workspace CI lane while the same cases passed in the feature lanes. Reserve nextest capacity for the exact `connect_registration` binary, as already done for related inventory and drive fixtures. Keep its existing deadlines, internal concurrency, assertions and zero-retry policy; report the last watch status on failure.

Rolling upgrades could pass `ListBuckets` readiness while restarted peers still lacked write quorum. Before each mixed-version phase, probe writes through every node outside the asserted workload prefixes. A shared 30-second deadline includes requests and sleeps; only HTTP 503 with `ServiceUnavailable` is retryable, with SDK retries disabled for these probes. The actual compatibility writes, reads, multipart operations and listing assertions remain unchanged.

Add four fast regression tests to the existing PR smoke profile, with matching exclusion from the full profile. No new workflow or job is introduced.

## Verification

- `cargo nextest run --locked --profile ci -p rustfs --lib --test connect_registration --test-threads 4 -E 'binary(/^connect_registration$/) | test(=connect::diagnostics::trace_runtime::tests::local_runtime_rejects_non_private_state)' --no-tests fail --status-level pass --final-status-level fail` passed 23/23 twice in 5.088s and 5.097s: 22 macOS registration tests plus one unrelated control. JUnit intervals confirm capacity reservation; no retries or test-process leaks occurred. The Linux-only inventory case remains for CI.
- `cargo nextest run -p e2e_test --lib --profile ci -E 'test(upgrade_write_readiness_tests)' --no-fail-fast --no-tests fail` passed 4/4 twice in 1.068s and 1.072s, with zero retries. Regressions cover metadata readiness followed by write unavailability, recovery through every writer, immediate permanent-error failure despite an incoming SDK retry configuration, and deadlines for repeated 503s and stalled requests. The original failure is recorded in [the mixed-version upgrade job](https://github.com/rustfs/rustfs/actions/runs/36569700716/job/109415806473).
- `cargo fmt --all --check`, `git diff --check`, `python3 scripts/check_test_wiring.py` and compiled smoke/full membership checks passed. Smoke membership changes from 188 to 192 by adding exactly these four tests; full membership is unchanged. The expected Linux smoke digest was derived from the actual prior Linux listing plus those four platform-independent additions and still requires confirmation by this PR's CI.

Local verification covers the exact source committed in `9b2ea836317ed035a91d1fc9fcf725f70c3098e2` on main `380e98a42cb4fcd0994fed79b30c2c7605deb0bc`. An independent final-diff correctness and reliability review found no findings. Fresh Linux workspace and real mixed-version upgrade runs are required before treating the remediation as fully verified; local fake-target tests do not establish that result.

## Impact

Test scheduling and readiness only; no production behavior, API, dependency, test deadline or compatibility assertion changes. Reserving capacity serializes registration fixture processes within a nextest run. The bounded readiness probes may add startup time while peer write health converges; permanent errors still fail immediately.

## Additional Notes

Rollback by reverting this PR. Existing CI restructuring from #8233 is independent of these follow-up fixes.
2026-09-30 10:06:09 +08:00
Chris 9e33d54269 fix(release): package stable preview tags without updating channels (#8256) 2026-09-30 09:49:27 +08:00
Chris 50aa482136 fix(connect): add non-Unix health runtime fallback (#8247) 2026-09-30 08:10:10 +08:00
Chris e870a6d25b fix: separate CPU profile stack lifetime from accumulator (#8250) 2026-09-30 01:36:14 +08:00
Chris 5851d9eb54 fix: filter CPU profile samples to reviewed executable symbols (#8248) 2026-09-30 01:06:24 +08:00
Chris 1e7065101d feat: Capture offline CPU profiles in serving process (#8245)
feat: capture offline CPU profiles in serving process
2026-09-29 22:37:10 +08:00
Chris 380e98a42c fix(test): synchronize top-disk fixture writes with sampling (#8243) 2026-09-29 21:07:45 +08:00
Chris a88225d208 fix(ci): reduce duplicate work and preserve reliable test failures (#8233)
* fix(ci): reduce duplicate work and preserve reliable test failures

* fix(ci): retain protocol evidence and repair stale test fixtures

* test(connect): honor parent deadline during API fixture readiness

* fix(ci): reserve IO capacity for state writer proofs

* test(connect): align RPC fixtures with service capture contracts

* test(connect): cover pinned service capture failures
2026-09-29 21:06:45 +08:00
Chris 1cca0dd25c feat(connect): select offline key for client performance (#8242) 2026-09-29 20:49:59 +08:00
Chris b0a54f0e5c feat(connect): capture offline network performance in server (#8240)
* feat(connect): capture offline network performance in server

* fix(connect): use storage facade in RPC test

* fix(connect): keep RPC test body behind storage facade
2026-09-29 19:41:43 +08:00
Chris 83a2d8b02d feat(connect): capture RPC activity in offline server runtime (#8237) 2026-09-29 18:16:54 +08:00
Chris f661651aad feat(connect): capture API activity in offline server runtime (#8234) 2026-09-29 17:18:26 +08:00
Chris 9298991d14 fix(scanner): restart a finished mixed sweep under its requested plan (#8231)
A bucket sweep that ends mixed clears its position and records the finishing cycle's plan as started. The next cycle requests a new plan whenever the bucket was written in between, so a fresh sweep inherited a stale started plan and ended mixed again. On a continuously written bucket no sweep could ever certify and the census kept the old root.

Start a fresh verification sweep under the requested plan when no durable position remains. Resumed sweeps keep their started plan, so a clean tail still cannot certify an old prefix.

Refs #7108
2026-09-29 16:48:58 +08:00
Chris 3c7d85420c fix(test): control clocks in Connect network fixtures (#8228) 2026-09-29 16:48:34 +08:00
Chris 6a7880c9fc feat(connect): capture offline lock metrics in server (#8230) 2026-09-29 16:29:07 +08:00
Chris d2175d1e1e fix(test): keep bucket disk faults across reconnects (#8224) 2026-09-29 14:07:30 +08:00
Chris 2019715d1a Fix fresh capacity probes for formatted local disks (#8222) 2026-09-29 12:23:06 +08:00
Chris 31f05a44af fix(test): use expect_err for offline state root rejection (#8220) 2026-09-29 11:22:16 +08:00
Chris a1724f3dbe feat(connect): capture signed offline service health (#8219) 2026-09-29 10:50:23 +08:00
Chris c62979a45b fix(connect): secure new offline state roots (#8218) 2026-09-29 10:13:58 +08:00
Chris e14eacc5ed fix(rpc): keep legacy format reads in their original namespace (#8216) 2026-09-29 09:24:39 +08:00
Chris 47af5565c7 fix(ci): preserve cluster startup failure evidence (#8213) 2026-09-29 08:50:09 +08:00
Chris f0ce628a7f feat(connect): capture offline native threads in service process (#8215) 2026-09-29 08:48:56 +08:00
Chris 32cc1eb76c Capture offline memory profile from the running service (#8214)
feat(connect): capture offline memory from the running service
2026-09-29 08:43:25 +08:00
Chris 84fe13989e fix(ecstore): retain read quota and control reserve test hedging (#8211) 2026-09-29 07:16:31 +08:00
Chris c8f42b8dca fix(ci): isolate scanner deadline and expose test failure details (#8210)
* fix(ci): isolate scanner deadline fixture and expose readiness errors

* test(ecstore): report unexpected capacity reservation errors
2026-09-29 07:16:17 +08:00
Chris c9acf01fd0 fix(tier): reread mutation intents after lock contention (#8209) 2026-09-29 07:16:01 +08:00
Chris 87e9a84a4e fix(ecstore): preserve RPC status when cloning storage errors (#8207)
* fix(ecstore): preserve RPC status when cloning storage errors

* test(heal): retain clone coverage without redundant ownership
2026-09-29 07:15:46 +08:00
Chris af0ae798f8 fix(test): serialize replacement_bucket_metadata heal test under parallel load (#8212) 2026-09-29 07:07:31 +08:00
Chris 5564932f38 fix(heal): retry cancelled internode RPC operations (#8205) 2026-09-29 01:58:35 +08:00
Chris 8778d55e46 fix(ci): preserve diagnostics for rio storage test failures (#8204) 2026-09-29 00:15:26 +08:00
Chris 5ebc1e480b test: stabilize heal timing and improve heartbeat replay coverage (#8150)
fix: preserve heartbeat replay and stabilize timing tests
2026-09-28 22:07:15 +08:00
Chris 4d782141d8 fix(test): add missing unavailable_drives field to readiness test initializers (#8200)
Commit 151103a609 added the  field to
StorageReadinessDetails but did not update two test initializers in
readiness.rs, causing E0063 compile errors in --all-targets builds.
2026-09-28 21:26:42 +08:00
Chris 4a91c75152 fix(ci): restore checks and multipart migration retries (#8196)
* fix(ci): restore multipart lint and full E2E membership checks

* fix(ecstore): preserve multipart completion identity during migration
2026-09-28 21:22:49 +08:00
Chris 2648776be7 fix: accept catalog files in health service artifact (#8198)
* fix: accept catalog files in health service artifact

* ci: retain configured health acceptance runner
2026-09-28 20:45:26 +08:00
Chris 64bd224b3b fix(connect): accept the health-era heartbeat and measured Top windows (#8195)
Two Connect tests fail on main in every Test and Lint lane.

Heartbeat: #8155 replaced "today's capabilities minus
profile.memory.service@1" with "the pre-health set minus it". That drops
the release that advertised health.check.service@1 but not yet in-service
memory profiling, so a pending heartbeat persisted by that release now
fails validation and the runtime stops. Accept that release as its own
frozen list.

Top disk/net: execute_*_job sets max_duration_millis to the requested
duration, then the capture compared the measured sleep against it. The
timer only overshoots, so a job that ran exactly as requested was
rejected with LimitExceeded whenever the overshoot reached 1 ms. Report
the authorized window, as Top API already does; validate_capture bounds
it by the limit.
2026-09-28 19:07:39 +08:00
Chris 71f8b607ad fix(ci): restore mainline and scheduled test reliability (#8187)
* fix(ci): restore mainline and scheduled test reliability

* fix(ci): provide GitHub CLI for CPU acceptance

* ci: provide Docker for CPU service acceptance

* ci: provision Python and Docker for OIDC validation

* ci: restore hosted runners for Docker validation

* ci: use verified MinIO release packages for interop

* ci: preserve host ownership of MinIO fixtures

* fix(ci): correct diagnostic limits and isolate startup checks

* test(connect): include object CLI failure details

* test(readiness): initialize unavailable drive diagnostics
2026-09-28 19:07:29 +08:00
Chris 151103a609 fix(ecstore): recover remote disks after transient stalls
Merge the approved fix from pull request #8149.
2026-09-28 18:40:30 +08:00
Chris 6ebc78d115 fix(ci): run profile acceptance on Docker runners (#8191) 2026-09-28 16:14:07 +08:00
Chris 5f4d5e8fe8 fix(ci): run CPU acceptance with Docker
Use the Docker-enabled runner required by the connected CPU service-job test.
2026-09-28 15:38:36 +08:00
Chris e4520f70a1 fix(ci): verify CPU acceptance identity without gh
Use authenticated curl for the pinned release run and artifact checks on sm-standard-2.
2026-09-28 15:13:03 +08:00
Chris 6523ab8d9f fix(ci): verify GNU packages with CPU catalogs
Match the archive-member check to the GNU pyroscope catalog packaging condition.
2026-09-28 13:23:46 +08:00
Chris 0609e7ce14 fix(connect): admit complete release CPU symbol catalogs
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
2026-09-28 12:09:26 +08:00
ChrisandHauser afb9ddabb4 fix(helm): route Gateway API traffic to the chart service (#8145)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-28 02:26:56 +00:00
Chris 719000460b ci(connect): match GNU build runner in acceptance checks (#8180) 2026-09-28 09:08:25 +08:00
Chris d791d7b761 ci(connect): use Node 25 for acceptance workflows (#8179) 2026-09-28 08:52:32 +08:00
Chris 1a251de4f3 feat(connect): execute signed disk diagnostics in the service (#8177) 2026-09-28 08:36:49 +08:00
Chris e634df8611 fix: collect offline Top disk from the running service (#8174)
fix(connect): capture offline disk IO in the running service
2026-09-28 01:35:40 +08:00
Chris f5fbb5f4e1 ci: verify Top disk in the serving process (#8173) 2026-09-28 01:34:29 +08:00
Chris 5d2d539423 ci: default manual builds to artifacts without publishing (#8171)
Default manual builds to artifacts without publishing
2026-09-28 01:10:52 +08:00
Chris acebc9649d feat(connect): capture service runtime profiles over local IPC (#8168)
feat(connect): export service runtime profiles over local IPC
2026-09-28 00:49:45 +08:00
ChrisandHauser e33542b0c2 fix(scanner): recover after cycle-state persistence failures (#8148)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 16:39:41 +00:00
Chris 2ee0ad1704 Collect bounded metrics from the serving Tokio runtime (#8167)
feat(connect): collect bounded service runtime profiles
2026-09-28 00:20:19 +08:00
Chris 4c5fc3c991 Use pinned release catalogs for Connect CPU acceptance (#8166) 2026-09-28 00:18:41 +08:00
Chris 6d2b1c629e Generate CPU symbol catalogs from release executables (#8165) 2026-09-28 00:12:36 +08:00
Chris 6eb14a11d0 fix(ci): run health acceptance on Docker-capable workers (#8158) 2026-09-27 14:26:28 +08:00
Chris 3667eea893 fix(ci): validate health binary without file utility (#8157) 2026-09-27 14:22:49 +08:00
Chris 8d033c85d8 fix(ci): install GitHub CLI for Connect health acceptance (#8146) 2026-09-27 08:04:54 +08:00
Chris b7420fe5f2 fix(ci): bootstrap release upload CLIs (#8144)
fix(ci): bootstrap release upload CLIs on self-hosted runners
2026-09-27 08:00:48 +08:00
Chris 5cb1c9e8cb fix(architecture): route health.rs storage imports through storage_api facade (#8143)
The health diagnostics module imported directly from
crate::storage::storage_api, bypassing the architecture migration
guardrail. Add a connect facade module to rustfs/src/storage_api.rs
and redirect the imports. Also fix a clippy::redundant-guards lint
in the coarse-flags match arm.
2026-09-27 07:53:54 +08:00
ChrisandHauser 2c5c43b0e0 test(connect): expose typed drive measurement failure context (#8099)
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:47 +08:00
ChrisandHauser cf096b3c22 fix(connect): pace diagnostic network payload sends (#8096)
* fix(connect): pace diagnostic network payload sends

* test(connect): exercise native network pacing entry

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-27 05:11:35 +08:00
Chris 82024907c5 test: add health service acceptance workflow (#8140) 2026-09-27 03:55:17 +08:00
Chris e2785846d0 ci: keep gh-dependent release jobs on GitHub-hosted runners (#8137)
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
2026-09-27 03:37:23 +08:00
Chris 31ac243865 feat(connect): add bounded health service job 2026-09-27 02:50:49 +08:00
8741bb77ea fix(storage): weight automatic multipart admission by part size (#8118)
* fix(storage): weight automatic multipart admission by part size

* fix(ci): restore filesystem runner capabilities and typos dependency

* fix(ci): make release guard portable and spell out part variables

* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs

---------

Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-26 22:35:39 +08:00
Chris de80e8a673 fix(ci): make shared checks portable across runners (#8124)
* fix(ci): isolate monitor argument checks from runner tools

* fix(ci): install the Typos action download dependency

* fix(ci): make release policy matching portable across awk variants
2026-09-26 10:55:06 +08:00
Chris 9f8acf688c feat(github): add issue forms with storage support checks (#8123) 2026-09-26 10:07:49 +08:00
Chris 96fc26cd6a fix(release): defer installation updates until artifacts are live (#8122) 2026-09-26 09:17:15 +08:00
ChrisandHauser d3b75e695e test(connect): add scheduler receipt acceptance workflow (#8078)
* test(connect): add scheduler receipt acceptance workflow

* chore(deps): upgrade crates and fix faster-hex advisory (#8098)

* test(connect): isolate scheduler acceptance workflow

---------

Co-authored-by: Hauser <housemecn@gmail.com>
2026-09-24 03:34:27 +08:00
Chris dde1aa77a1 fix(connect): preserve diagnostic schedules across restart (#8092) 2026-09-24 03:28:37 +08:00
Chris be01e513be feat(connect): sample memory within the running service (#8091)
* feat(connect): sample memory within the running service

* test(connect): add official memory service acceptance

* test(connect): consume the final memory job without cloning
2026-09-24 03:25:28 +08:00
Chris 713920dcd6 docs: update license copyright to RustFS, Inc. (#8097) 2026-09-24 00:56:08 +08:00
Chris 3f549e26f3 fix(ci): publish version-tagged preview Docker images (#8068) 2026-09-22 13:06:23 +00:00
Chris 4047ed6c8d fix(heal): restore bucket metadata before replacement completion (#8042) 2026-09-21 12:18:06 +08:00
Chris 9aeecad4a9 docs(security): add encrypted ETag advisory lesson (#8040) 2026-09-20 23:52:47 +08:00
Chris 0632a64f91 fix(ecstore): allow server-side copy at normal read quorum (#8029) 2026-09-20 12:36:53 +08:00
Chris 0fcd34c331 fix(ecstore): retry fleet capability probe while the notification system boots (#8017)
Startup finalizes IAM before init_notification_runtime publishes the notification system, and IAM finalization is what starts the fleet capability probe. The first probe pass therefore always failed closed and then slept the full ten-second probe interval. On a single node that left every durable capability, including the durable hard quota fence, unavailable for about ten seconds after /health already reported ok, so SetBucketQuota answered 503 durable quota capability is not confirmed across the cluster during that window.

Publishing the notification system now wakes the probe immediately through a Notify permit, with a 100ms bootstrap poll as the fallback for a wakeup that races the availability check. The probe keeps failing closed while the system is absent and logs the wait once. The four probe futures no longer carry an unreachable notification-system-unavailable arm, and the seven proof slots are revoked through one helper.

Fixes #8014
2026-09-19 13:41:16 +08:00
Chris ddc5f97a04 fix(release): sort pre-release packages below their final release (#8016)
Pre-release RPM versions replaced the SemVer `-` with `_`, which rpm treats as an ordinary segment separator, so `1.0.0_rc.5` compared as newer than `1.0.0` and a host with an rc RPM installed could not `dnf upgrade` to the GA package. DEB previews had the mirror-image problem: only the first `-` became `~`, so `1.0.0~rc.5-preview.2` carried `preview.2` as its Debian revision and sorted above `1.0.0~rc.5`.

Both formats now spell every pre-release separator as `~`, which dpkg and rpm (>= 4.10) both treat as "sorts before anything". Verified with rpm 4.16 (AlmaLinux 9) and dpkg 1.21: `1.0.0~rc.5 < 1.0.0`, `1.0.0~rc.5~preview.2 < 1.0.0~rc.5 < 1.0.0~rc.6`, alpha < beta < rc, and `rc.9 < rc.10`; fpm 1.18 passes `~` through into both package headers, and `dnf upgrade` from a `~rc.5` RPM to the GA RPM succeeds where the `_rc.5` one refused. The test now pins the full ordering contract under both package managers and fails against the previous spelling.

The release-checksum step already normalizes `~` to `.` for every uploaded asset, so RPM assets need no further handling. Hosts that already installed an `_rc`/`_beta` RPM need a one-time `dnf install rustfs-1.0.0` or `rpm -Uvh --oldpackage` to reach GA.

Fixes #8012
2026-09-19 03:22:44 +08:00
Chris 4e7a8fabb8 fix(ecstore): repair buckets left with only an incarnation sidecar (#8008)
The legacy bucket metadata migration writes the `.bucket-incarnation`
sidecar before `.metadata.bin`. A crash or a lost namespace lease between
the two writes left the bucket with a sidecar and no metadata, and every
later load failed closed with "bucket incarnation sidecar exists without
bucket metadata", so the bucket answered 500 to every request on every
node with no repair path (rustfs/rustfs#8003).

Load the sidecar-only state as a legacy bucket so the migration runs
again. The migration and the force-create path re-read the sidecar under
the transaction lock and keep its incarnation when persisting the
metadata, so the retry never replaces an identity other nodes fenced on.
A retired incarnation is never adopted: it is residue of a deleted
bucket, and re-publishing it would let heal reclaim the new objects.
2026-09-18 18:14:00 +08:00
Chris aa7f55946c docs(operations): add cluster and erasure-coding lifecycle runbook (#7998)
Consolidate the operator entry point for planning, storage classes and
EC:0 consequences, expansion, rebalance, decommission, heal, drive
replacement, restart recovery, and the rc CLI mapping. Every claim was
checked against the current handlers, storage-class validation, quorum
arithmetic, and the rustfs/cli admin reference; the runbook links the
normative contracts instead of restating them.

Refs rustfs/backlog#2607, rustfs/backlog#2608
2026-09-18 12:50:14 +08:00
Chris 3f469ab257 fix(helm): default the ingress backend port and refresh the nix cli hashes (#7990) 2026-09-18 09:08:24 +08:00
Chris 5886368178 ci(performance): drop stale shared-lock comment from the performance workflow (#7987)
fix(ci): align security workflow test with the performance suite's own concurrency group

#7984 moved rustfs-performance-test.yml to the rustfs-performance-suite group but left test_all_suites_hold_the_shared_lock_for_manual_and_chain_runs asserting the shared functional group, so Quick Checks (make script-tests) fails on main and on every PR. Expect the dedicated group for the performance suite and drop the workflow comment that still described the shared lock.
2026-09-18 01:26:08 +08:00
Chris d9c3eeea0c fix(ecstore): hide and reclaim ancestors of delete residue in prefix listings (#7955)
* fix(ecstore): hide ancestors of delete residue in prefix listings

Since #7342 the never-versioned fast path hides a directory whose children are all deleted-version data dirs, but only that one level: the date-style ancestors above a deleted key (`metrics/<stream>/2026/08/28/23/`) still surfaced as empty folders that nothing could remove, which is what #6898 keeps reporting on rc.6 and 1.0.0. The residue probe now walks down through metadata-less ancestors depth-first, ending at the first file met outside a UUID data dir, with a total read budget so a large residue tree surfaces and is hidden one level down instead of costing an unbounded walk. An empty delimiter listing of the ancestor then reclaims the whole committed residue tree in one pass.

Refs #6898

* fix(ecstore): sample before full reads in the residue probe and purge orphan subtrees independently

Adversarial pass 1 findings: the probe read a wide directory in full at every level of a genuine prefix once its 8-entry batch was full, so a leaf holding thousands of objects was materialised per listed prefix; the probe now descends into the sampled children first and only reads the remainder once every sample proved unlistable, and a test pins one bounded read per level and zero complete reads on a wide genuine prefix. Hiding ancestors also made purge_orphan_dir_object fail closed on a whole ancestor when committed residue shared it with residue an older build left without markers, which 1.0.0 users could still reclaim by browsing one level lower; the purge now blocks only the ancestor chain of unpurgeable data and reclaims purgeable sibling subtrees. A UUID directory holding a subdirectory surfaces again as before.

Refs #6898

* fix(ecstore): back off repeated orphan purge scans of unpurgeable trees

Adversarial pass 2: with ancestors hidden, every empty listing of a phantom prefix scanned the whole residue subtree on every disk even when nothing under it can be purged (no committed marker), and the scan no longer stops at the first blocked directory. Remember such prefixes per set for 60 s and skip the rescan. Build committed residue file paths like directory paths so the blocked filter matches under keys containing repeated slashes.

Refs #6898
2026-09-17 10:11:45 +08:00
Chris 706ffdc9ac fix(ecstore): keep the single-block bound on every inline candidate (#7954) 2026-09-17 08:57:49 +08:00
Chris f018628374 fix(ilm): keep expiry pending gauge balanced and drain overwrite tails (#7944) 2026-09-17 00:38:02 +00:00
Chris 14081c41a4 fix(ecstore): forward quota admission on CopyObject destination writes (#7947) 2026-09-17 06:38:16 +08:00
Chris 58a3c599ba chore(release): prepare 1.0.1 (#7946) 2026-09-16 19:06:40 +00:00
Chris f7431ce747 fix(ci): bound the fault-tolerance dashboard upload step (#7948) 2026-09-17 02:02:44 +08:00