* fix(test): establish a writable previous-release upgrade baseline
* fix(ci): reserve capacity for durable admin fixtures
* fix(ci): group durable IAM state fixtures by resource needs
* ci: run E2E doctests with the E2E dependency graph
* fix(ci): separate fixture startup from transport deadlines
* fix(ci): bound pagination after seeding and revisit restored copies
* fix(ci): make filesystem fixture timing deterministic
* fix(ecstore): avoid metadata lock reentry during internal mutations
* test: align recovery fixtures with durable ownership contracts
Include helm/rustfs/README.md in the uploaded helm-package artifact so
publish-helm-package updates the root README.md in rustfs/helm.
The build-helm-package job copies helm/README.md into helm/rustfs/
before running helm package, but upload-artifact previously uploaded
only helm/rustfs/*.tgz. Because both files share the helm/rustfs/
parent directory, actions/upload-artifact places both the chart
archive and README.md at the artifact root, and download-artifact
extracts README.md into the root of rustfs/helm alongside the tarball.
* fix(ci): reduce duplicate work and preserve reliable test failures
* fix(ci): retain protocol evidence and repair stale test fixtures
* test(connect): honor parent deadline during API fixture readiness
* fix(ci): reserve IO capacity for state writer proofs
* test(connect): align RPC fixtures with service capture contracts
* test(connect): cover pinned service capture failures
The security workflow checks the repository out into rustfs-repo/ (#7212)
but its chain evidence steps still invoke
scripts/functional_chain_evidence.py relative to the workspace root, which
on the persistent shared runner resolves to a stale checkout left by another
job. The evidence gate then compares that checkout's HEAD against the chain
workflow_sha and rejects the lane before any case runs
("lane checkout differs from chain workflow source").
Run the evidence script from the lane's own checkout so ROOT resolves to
rustfs-repo/, whose HEAD is exactly the chain-pinned workflow_sha.
* fix(ci): restore mainline and scheduled test reliability
* fix(ci): provide GitHub CLI for CPU acceptance
* ci: provide Docker for CPU service acceptance
* ci: provision Python and Docker for OIDC validation
* ci: restore hosted runners for Docker validation
* ci: use verified MinIO release packages for interop
* ci: preserve host ownership of MinIO fixtures
* fix(ci): correct diagnostic limits and isolate startup checks
* test(connect): include object CLI failure details
* test(readiness): initialize unavailable drive diagnostics
* ci: give the pool run a real wall-clock budget
The pool suite budgets up to 24h each for rebalance and decommission
(REBALANCE_TIMEOUT / DECOMMISSION_TIMEOUT in rustfs_pool_expand.sh),
but the workflow step capped it at 45 minutes. Four consecutive runs
died identically: steps 1-7 all PASS, then the rebalance wait was
killed at exactly 45:12 - chain 35680308150, standalone 35702803577,
chain 35757773373, chain 36283120750 - with rebalance at completed=1/4
(~21 minutes in), so a full pass has never been observed.
Make the budget an input (default 240 minutes: covers the observed
rebalance pace plus one decommission pass) and document why. The suite
keeps failing the job through its [POOL-STEP] marker adjudication.
* fix(ci): align pool job budget and timeout contract tests
* fix(connect): preserve legacy heartbeats and isolate I/O tests
---------
Signed-off-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: RustFS <hello@rustfs.com>
Route the two linux-aarch64-* release legs to the dedicated arm64 runner
and build them natively (cross: false) instead of cross-compiling via
cargo-zigbuild on x86_64 hosts.
With native legs now executable, extend the packaged-artifact checks that
were previously gated to x86_64-unknown-linux-gnu only:
- Every Linux leg runs rustfs --version and rustfs-cli --help from its own
package, proving the artifact matches the runner architecture.
- The packaged-console runtime smoke (server boot + console HTTP 200) also
covers aarch64-unknown-linux-gnu, so both architectures get full runtime
verification. Musl legs run the binary checks but skip server boot for
now.
* ci: run the chain orchestration jobs where the gh CLI exists
#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.
Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.
* ci: run the chain orchestration jobs where the gh CLI exists
#8112 moved these jobs to the sm-standard-2 pool, whose images ship
without the gh CLI (a known property of the build fleet, see #7572).
Last night's chain died in prepare before any suite ran:
resolve_functional_candidate.py:34 does subprocess.run(["gh", "api"])
and got FileNotFoundError; complete-chain and the hourly
functional-chain-health canary fail the same way. On Sep 22/23 the same
jobs ran green on GitHub-hosted runners.
Move prepare, complete-chain and the health canary to ubuntu-latest.
The suite lanes stay on smoke-testing.
* ci: point the chain orchestration jobs at the smoke-testing runner
Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
* ci: point the chain orchestration jobs at the smoke-testing runner
Per maintainer decision, consolidate them onto the same runner as the
test lanes. Verified on rustfs-smoke-testing as the runner user:
gh 2.45.0, jq 1.7.
The sm-standard-2 fleet has no gh CLI (see #7572), so moving the release publication jobs onto it in #8112 broke every tag release: the 1.0.1-preview.12 build failed in Create GitHub Release with 'gh: command not found'. Move create-release, upload-release-assets, publish-release, cleanup-preview-releases, and package.yml resolve/package back to ubuntu-latest.
Companion to rustfs/auto-testing#109: the 4x1 topology is removed from
MATRIX_TOPOS because the product rejects a 1-drive-per-node distributed
pool at startup (FATAL BelowMinimum). The unfiltered sweep is now 36
combinations / 300 cases; update the contract totals and comments so an
unfiltered sweep is still held to an exact expected case count.
#8112 replaced ubuntu-latest with sm-standard-2 across workflows. Both nightly
KMS Vault lanes need a working Docker daemon; the sm-standard-2 pool does not
provide one, so both lanes die in seconds on 'Docker is not available' while
the package compiles and publishes fine - and the functional chain never fires
because its gate requires a successful build run. Restore the lanes to
ubuntu-latest, as documented in the lane comment before #8112.
* fix(scanner): bound SNSD deep scans and checkpoint cloning
* ci: run mount-dependent jobs on hosted VMs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
docs(github): clarify supported storage in bug reports
State that local disks, SAN/Fibre Channel volumes, and JBOD are supported,
recommend XFS, and keep the existing unsupported remote/shared filesystem rule.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* ci: replace ubuntu-latest runner with sm-standard-2 across workflows
* ci: keep scheduled-validation monitors on hosted runners
The freshness and watchdog jobs report stalled scheduled validations.
Running them on the same sm-standard-2 pool means a pool outage stalls
the monitors too, so nothing reports it.
---------
Co-authored-by: overtrue <anzhengchao@gmail.com>
* feat(connect): sample memory within the running service
* test(connect): add official memory service acceptance
* test(connect): consume the final memory job without cloning
* ci(package): build gnu and musl DEB/RPM variants with distinct file names
The Build and Release workflow produces four Linux binaries
(x86_64-gnu, aarch64-gnu, x86_64-musl, aarch64-musl), but packaging
only consumed the two gnu artifacts. Add matrix entries for the two
musl artifacts so every release ships all four DEB/RPM variants.
The libc variant is now part of the package file names, which would
otherwise collide between gnu and musl builds of the same version:
- deb: rustfs_<version>_<libc>_<arch>.deb
- rpm: rustfs-<libc>-<version>-<release>.<arch>.rpm
The dpkg Package and rpm Name stay plain "rustfs", so gnu and musl
remain mutually exclusive upgrades of one package rather than
co-installable packages fighting over /usr/bin/rustfs.
Dependency declarations now follow the linkage: gnu binaries
dynamically link glibc and keep Depends: libc6 (>= 2.31) /
glibc >= 2.31; musl binaries are statically linked and declare no
libc dependency. The libc variant is also visible in the package
description.
scripts/release/package_versions.sh gains a LIBC argument and its
contract tests cover both variants plus the invalid-libc cases.
* ci(package): align deb/rpm file names with the zip artifact naming
Rename the package file names so every release asset of one build
shares the same stem as its binary artifact, differing only by
extension:
- before: rustfs_<deb_version>_<libc>_<deb_arch>.deb
rustfs-<libc>-<rpm_version>-<rpm_release>.<rpm_arch>.rpm
- after: rustfs-linux-<arch>-<libc>-v<version>.deb / .rpm
e.g. rustfs-linux-x86_64-gnu-v1.0.0.zip,
rustfs-linux-x86_64-gnu-v1.0.0.deb,
rustfs-linux-x86_64-gnu-v1.0.0.rpm.
Non-development builds embed the raw release tag (with 'v'), like the
zips; development builds embed dev-<full sha>. The dpkg/rpm versions
(including the '~' prerelease ordering) are unchanged - they live in
the package metadata, and a side effect is that release asset names no
longer contain '~' (which GitHub normalizes to '.').
package_versions.sh now takes the target arch (x86_64|aarch64) instead
of the deb/rpm arch pair; the deb Architecture (amd64/arm64) in the
control metadata still comes from the workflow matrix. The two test
workflows that assemble deb download URLs from a release tag
(rustfs-table-test, rustfs-upgrade-test) are updated to the new name,
which also removes their '~'-to-'.' asset name workaround.
Adds rustfs-fault-tolerance-matrix.yml, a manually dispatched workflow that
runs the --matrix topology x EC outage sweep from rustfs/auto-testing (38
combinations, 320 cases). The scenario suite hardcodes --all and a 60-minute
budget, so the exhaustive sweep had no CI entry point.
The fault-tolerance suite now registers and runs scenario E
(rustfs/auto-testing#100): --all covers A, B, C, C2, D, E with 38 + 6
= 44 expected cases. The chain-evidence assertion still hardcoded
[A, B, C, C2, D] and 38 everywhere, which would fail every green FT
run's evidence validation once E is registered.
Bump the scenario list to include E and all count assertions from 38
to 44; step name updated to match. No other lanes reference 38.
resolve_functional_candidate.py probes rustfs/auto-testing (private) to
age the pinned functional-script revision and fall back to main HEAD
after 24h. The prepare step passed github.token, which cannot see the
private repo, so every nightly chain logged
staleness probe failed (...exit status 1.); keeping pinned revision
and replayed the 09-14 harness. On the 09-20 nightly that harness died
on the dpkg conffile prompt in all 12 lanes (see rustfs/auto-testing#97)
because the --force-confold and other fixes never reached the chain.
Use PF_TESTING_GH_TOKEN - already required by the other steps in this
workflow - so the probe can actually run and the >24h fallback works.
Docker Hub's overview is a separate `full_description` field that
`docker push` never touches, so it had drifted into an 8 KB snapshot of
an old README.md that still linked to
https://docs.rustfs.com/introduction.html (now 404).
Add a `sync-dockerhub-description` job to docker.yml that runs after the
images are pushed and republishes README.md from the same commit via
peter-evans/dockerhub-description (pinned to v5.0.0). It reuses the
existing DOCKERHUB_USERNAME / DOCKERHUB_TOKEN credentials, so no new
secrets are needed. Relative links (docs/, CONTRIBUTING.md) are
rewritten to github.com URLs so they resolve on Docker Hub.
Docker Hub caps the field at 25,000 bytes and the action truncates to
fit with only a warning; README.md is at 22,879 bytes today. Read the
published overview back after the sync and fail the job if it hit the
cap, so a truncated overview cannot be published silently.
Fixes#7995
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(package): preserve service state across upgrades
* fix(package): match legacy DEB versions in tilde form
Published prerelease packages carry ~ in the dpkg control version
(package_versions.sh maps the SemVer prerelease - to ~), so the
legacy fallback list written with dots never matched 1.0.0~rc.x and
upgrades away from those DEBs still left the service stopped (#8011).
Fix the legacy glob to the tilde form and update the contract test,
which had enshrined the dot form. Verified on Ubuntu 24.04 systemd
containers: DEB upgrade 1.0.0~rc.5 -> 1.0.1 now keeps the service
running; 1.0.0 -> 1.0.1 and 1.0.1 -> 1.0.2 marker path still pass.
* docs(package): expand /etc/default/rustfs example template
Document the commonly used RUSTFS_* settings as commented examples in
the packaged conffile and point to docs.rustfs.com.