* ci: give the pool run a real wall-clock budget The pool suite budgets up to 24h each for rebalance and decommission (REBALANCE_TIMEOUT / DECOMMISSION_TIMEOUT in rustfs_pool_expand.sh), but the workflow step capped it at 45 minutes. Four consecutive runs died identically: steps 1-7 all PASS, then the rebalance wait was killed at exactly 45:12 - chain 35680308150, standalone 35702803577, chain 35757773373, chain 36283120750 - with rebalance at completed=1/4 (~21 minutes in), so a full pass has never been observed. Make the budget an input (default 240 minutes: covers the observed rebalance pace plus one decommission pass) and document why. The suite keeps failing the job through its [POOL-STEP] marker adjudication. * fix(ci): align pool job budget and timeout contract tests * fix(connect): preserve legacy heartbeats and isolate I/O tests --------- Signed-off-by: Hauser <housemecn@gmail.com> Co-authored-by: overtrue <anzhengchao@gmail.com> Co-authored-by: Hauser <housemecn@gmail.com> Co-authored-by: RustFS <hello@rustfs.com>
Documentation
Use the focused indexes rather than treating this directory as an unordered collection:
Operations
For the logical per-operation io_uring read cap, see io_uring read chunk size.
For bounded local read-backend startup and cancellation, see io_uring backend initialization.
For legacy protection assessment and bounded protected copies, see Object integrity inventory, audit, and migration.
Operational runbooks live under operations/. Replication
operators should start with:
| Runbook | Use it for |
|---|---|
| Site replication operations | Health fields, pending operations, outage recovery, re-pair admission, IAM/SSE boundaries, and upgrades. |
| Replication target check | Validating an S3 destination and version fidelity before enabling replication. |
| Replication object size limits | Multipart routing, large-object limits, and retry characteristics. |
| Replication outbound transport | Integrity headers, generic target behavior, and transport knobs. |
For the erasure-coded cluster lifecycle (planning, parity and EC:0,
expansion, rebalance, decommission, heal, drive replacement, restart
recovery, and the rc CLI mapping), start with
Cluster and erasure-coding lifecycle operations.
For persisted administrator bucket tasks and bucket recreation, see Bucket heal recovery.
For disk replacement across VM restarts and schema 5/6 maintenance migration, see Replacement generation recovery.
For historical GET timeouts during PUT or Heal, see Object lock contention diagnostics.
Other runbooks remain grouped by filename in operations/;
architecture pages link to the relevant runbook where a cross-boundary
procedure is required.
For storage dashboards, see Storage metrics and observer selection: drive ownership, snapshot freshness, counter queries, and rolling upgrades.
For optional shard commitments, see Independent shard integrity rollout: activation, legacy repair results, multipart mode changes, and rollback limits.
For crates.io publication of workspace crates, see Workspace Cargo Publish: dependency ordering, dry-run, publish, and failure handling.