Files
rustfs/docs
f54945a4c4 fix(ci): repair pool budgets and Connect test failures (#8155)
* ci: give the pool run a real wall-clock budget

The pool suite budgets up to 24h each for rebalance and decommission
(REBALANCE_TIMEOUT / DECOMMISSION_TIMEOUT in rustfs_pool_expand.sh),
but the workflow step capped it at 45 minutes. Four consecutive runs
died identically: steps 1-7 all PASS, then the rebalance wait was
killed at exactly 45:12 - chain 35680308150, standalone 35702803577,
chain 35757773373, chain 36283120750 - with rebalance at completed=1/4
(~21 minutes in), so a full pass has never been observed.

Make the budget an input (default 240 minutes: covers the observed
rebalance pace plus one decommission pass) and document why. The suite
keeps failing the job through its [POOL-STEP] marker adjudication.

* fix(ci): align pool job budget and timeout contract tests

* fix(connect): preserve legacy heartbeats and isolate I/O tests

---------

Signed-off-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: RustFS <hello@rustfs.com>
2026-09-28 14:26:21 +08:00
..

Documentation

Use the focused indexes rather than treating this directory as an unordered collection:

Operations

For the logical per-operation io_uring read cap, see io_uring read chunk size.

For bounded local read-backend startup and cancellation, see io_uring backend initialization.

For legacy protection assessment and bounded protected copies, see Object integrity inventory, audit, and migration.

Operational runbooks live under operations/. Replication operators should start with:

Runbook Use it for
Site replication operations Health fields, pending operations, outage recovery, re-pair admission, IAM/SSE boundaries, and upgrades.
Replication target check Validating an S3 destination and version fidelity before enabling replication.
Replication object size limits Multipart routing, large-object limits, and retry characteristics.
Replication outbound transport Integrity headers, generic target behavior, and transport knobs.

For the erasure-coded cluster lifecycle (planning, parity and EC:0, expansion, rebalance, decommission, heal, drive replacement, restart recovery, and the rc CLI mapping), start with Cluster and erasure-coding lifecycle operations.

For persisted administrator bucket tasks and bucket recreation, see Bucket heal recovery.

For disk replacement across VM restarts and schema 5/6 maintenance migration, see Replacement generation recovery.

For historical GET timeouts during PUT or Heal, see Object lock contention diagnostics.

Other runbooks remain grouped by filename in operations/; architecture pages link to the relevant runbook where a cross-boundary procedure is required.

For storage dashboards, see Storage metrics and observer selection: drive ownership, snapshot freshness, counter queries, and rolling upgrades.

For optional shard commitments, see Independent shard integrity rollout: activation, legacy repair results, multipart mode changes, and rollback limits.

For crates.io publication of workspace crates, see Workspace Cargo Publish: dependency ordering, dry-run, publish, and failure handling.