670 Commits
Author SHA1 Message Date
Joe HuberandRuslan Konviser 72604aef90 chore(git): add github issue templates (#10330)
* chore(git): add github issue templates

* Update config.yml

* Reorder community links in issue template

---------

Co-authored-by: Ruslan Konviser <evereq@gmail.com>
2026-09-29 13:41:28 +02:00
Ruslan Konviser efdba39b6b Merge remote-tracking branch 'origin/develop' into fix/unit-tests-hermetic-harness 2026-09-28 09:03:49 +02:00
Ruslan KonviserandClaude Opus 5.5 e77e3130f4 ci: reword the build-record comment (cspell)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 21:27:58 +02:00
Ruslan KonviserandClaude Opus 5.5 6c68c1fda0 ci: stop uploading Docker build records and summaries
docker/build-push-action uploads a build record (*.dockerbuild) for every image
build. The record stores every build-arg value in plain text, and on a public
repo any signed-in GitHub user can download it. Set DOCKER_BUILD_RECORD_UPLOAD
and DOCKER_BUILD_SUMMARY to false at workflow level in every workflow that uses
the action. BuildKit secrets are never recorded, so nothing else changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 21:02:42 +02:00
Ruslan Konviser d0b2a8e080 Merge remote-tracking branch 'origin/develop' into fix/unit-tests-hermetic-harness 2026-09-27 20:53:10 +02:00
Ruslan KonviserandClaude Opus 5.5 91d91ffca3 test: make the unit-test run hermetic and fix the Angular jest harness
The Unit Tests workflow has never been green (24 of 97 projects red at
develop eac7136562, run 36319732347). Causes addressed here:

- Env leak: Nx loads the committed .env.local (DEMO=true,
  WORKER_QUEUE_ENABLED=false, NODE_ENV=development) into every task.
  The test step now sets NX_LOAD_DOT_ENV_FILES=false, and the two specs
  that depend on a mode pin it themselves (role-permission counts pin
  environment.demo=false and gain a demo-mode suite asserting 209 per
  role; the worker spec clears the queue env before the constants load).
  .env.local is unchanged.
- Angular harness: nohoist gives each workspace its own @angular copy,
  so setupZoneTestEnv() initialised a different TestBed from the one the
  spec used. A workspace resolver (jest.resolver.js, wrapping Nx's) makes
  every @angular/* import in a Jest project resolve to one copy. The
  Angular projects now inherit the preset's transformIgnorePatterns,
  which gains the .mjs exception plus @datorama, @ngneat and lodash-es.
- ui-config's environment.ts is generated and gitignored; the workflow
  runs `yarn config:dev` before the tests, as the Playwright workflow does.
- Misconfigured targets: gauzy (jest.config.js -> .ts, Angular transform,
  setupFile moved into the config), integration-sim-ui (.ts -> .cts),
  integration-activepieces (config file added), mcp-auth (ran
  `node build/main.js --test`; now the jest executor like apps/mcp).
- Real failures: the openai "silence" case used a body with no `text`,
  which the shared helper rejects by contract; the toolbar spec read the
  tabIndex property, which is 0 on any button; the docs-ui linearity
  tests asserted wall-clock bounds, now a 4x-input growth ratio.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 17:39:40 +02:00
Ruslan KonviserandClaude Opus 5.5 40826df61b ci: unpack the lint cache off the RAM disk; PRs into develop start no workflow
Static Checks / lint (x64-4 lane): every cache hit since the job moved to
this lane ran out of space. The 2.47 GB archive plus the tree it unpacks to
does not fit the 16Gi RAM-backed workspace, so each hit fell back to a
1.5-4 h cold install (16 of 16 runs cold since #10303). The Restore step
now moves the archive onto the disk-backed package-cache volume (the
parent of YARN_CACHE_FOLDER) before extracting, so only the tree lives in
RAM. Where YARN_CACHE_FOLDER is unset, or the move fails, it extracts in
place exactly as before. Two report-only df lines (after the restore and
at the top of Summarize) record workspace headroom and can never fail a
step.

Triggers: apply the owner decision of 2026-09-21 (already on draft #10254,
same text) directly to develop. A pull request INTO develop no longer
starts static-checks, typos (cspell), snyk-analysis (trigger removed,
restore lines kept in a comment), or build, secrets-analysis and the
external-uptime-monitor self-test (branches-ignore: develop, so PRs into
stage/master still run them). Every push trigger is unchanged, so all of
them still run when code lands on develop. develop has no required status
checks, so no PR is blocked by the missing runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 15:11:58 +02:00
Ruslan KonviserandClaude Opus 5.5 cec254cd35 ci(desktop): release prod snaps to the Snap Store stable channel
Every Gauzy app snap (desktop, desktop-timer, server, api-server, agent,
mcp-server; amd64 and arm64) went to the Snap Store 'edge' channel only,
from both the stage lane (stage-apps) and the prod lane (apps). The
electron-builder snap target publishes with its own snapStore config.
When the build config has no snapStore entry, that config has no
channels, and the Snap Store publisher then defaults to 'edge'.

Prod builds now release to 'stable' and stage builds to 'edge'. Each
Linux build step appends
-c.snap.publish.provider=snapStore -c.snap.publish.channels=<channel>
to its yarn command. yarn adds extra arguments to the end of the script,
which is the electron-builder call. The override is in the workflows
because the root build scripts are shared by both lanes. Only
snap.publish is overridden: a -c.publish override would be merged into
the GitHub publish entry. GitHub releases, update channels and every
other target are unchanged. Stage names 'edge' explicitly so each
prod/stage pair still differs only in the channel.

The generated snaps already use grade stable and strict confinement,
and they need no store-approved plugs, so the store accepts them on
stable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 01:10:55 +02:00
Ruslan Konviser fb2184e41f Merge pull request #10305 from ever-co/fix/desktop-update-followups
fix(desktop): update-feed follow-ups: publish-channel CI guard, updater tag re-resolve, local-update note, stray channel
2026-09-25 00:53:23 +02:00
Ruslan KonviserandClaude Opus 5.5 271b43e24e ci(desktop): fail win/linux releases whose dist package.json lacks the per-arch update channel
Gauzy Server's Windows and Linux release jobs published channel-less
update manifests (latest.yml, latest-linux*.yml) from v107 to v111.44.45.
Its :isolated scripts ran `pack --arch` after the build had already copied
apps/server/src/package.json into dist. electron-builder reads
dist/apps/<x>/package.json, never saw build.publish[].channel and fell
back to `latest`. The updater only requests `latest-${process.arch}`, so
installs stopped auto-updating, and nothing caught it for 4.5 months.
#10299 fixed that app. This change makes the whole class fail loudly.

Add .scripts/electron-package-utils/assert-publish-channel.js, a Node
script with no dependencies. It takes --project <dist dir> --arch
<x64|arm64> and requires every build.publish[] entry in
<dist dir>/package.json to have channel latest-<arch>. Otherwise it prints
a GitHub `::error` annotation naming the file and the expected and actual
channels, plus a hint, and exits 1. A missing file, unreadable JSON or bad
arguments also exit 1.

Run it immediately before electron-builder in the 28 root scripts that
publish for --windows/--linux and stamp the channel with `pack --arch`.
The list was found from the scripts, not typed by hand. Each guard's
--project and --arch are asserted equal to that script's electron-builder
--project, its --x64/--arm64 flag and its pack --arch. A bad channel now
stops the job before anything is published. No other script changes.
The 24 win/linux publishing scripts without `pack --arch` have no callers
in any workflow and are left alone.

Add assert-publish-channel.test.js (node --test, no dependencies) as
`yarn test:publish-channel` and run it in build.yml next to
test:postinstall. It covers the pass, fail and misuse cases. It also
checks that every per-arch win/linux publishing script keeps the guard on
the same arch and dist dir, so a new or edited release script cannot drop
it unnoticed. That check fails against the current develop package.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 00:42:07 +02:00
Ruslan KonviserandClaude Opus 5.5 5ab57761ba ci(static-checks): derive the lint job's expect-vip from x64-4 to match its runs-on
Moving `lint` to `RUNNER_LINUX_X64_4` left its Configure Registry step's
`expect-vip` still derived from `RUNNER_LINUX_X64_8`. Flagged by both
greptile-apps and coderabbitai on line 74.

`expect-vip` tells the pinned configure-registry action (aa4ee199) that the
runner is in-network: it probes the Verdaccio VIP twice instead of once and
emits "Verdaccio VIP unreachable from an in-network runner" on failure. Either
way the job falls back to public npm and never fails - it is a diagnostic.

Every other job in the repo derives `expect-vip` from the same variable as its
`runs-on` (all nine x64-8 jobs read x64-8; the matrix release jobs use
"self-hosted or ever-k8s"). `lint` was the only mismatch, introduced by the
previous commit. This restores the invariant.

Took greptile's fix (derive from x64-4). Did NOT take coderabbit's literal
`false`: x64-4 is an in-network ARC runner, so `false` would make `lint` the one
self-hosted job with no retry and no warning, turning a registry outage on that
lane into a silent single-probe fallback.

No behaviour change today, since both variables are set org-wide; this closes
the case where x64-4 is configured and x64-8 is not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 21:09:41 +02:00
Ruslan KonviserandClaude Opus 5.5 2d6036e1b9 ci(static-checks): run the non-blocking lint job on the x64-4 lane
The `lint` job is report-only (`continue-on-error: true`) and runs
`nx run-many -t lint --parallel=2`, so it never uses more than two parallel
tasks and gains nothing from the x64-8 lane's 12-core limit. Its sibling in the
same workflow, `typecheck-configs`, already runs on x64-4.

The x64-8 lane is capped at 12 runners and each one requests 28 Gi, against
12 Gi for x64-4. It was observed saturated at 12/12 on 2026-09-22 while Docker
image builds queued behind it. Moving a job that cannot use the extra cores
frees a slot and 16 Gi of requested memory for the builds that can.

Deliberately NOT moved:
- unit-tests: jest forks workers per project, so fewer cores would slow a job
  that already takes 123-233 minutes and risk its 360-minute timeout.
- the e2e suites (playwright / cypress / currents): timing-sensitive, and a
  slower runner makes that class of test flakier, not just slower.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 20:15:54 +02:00
Ruslan KonviserandClaude Opus 5.5 f6fe23a4bf fix(ci): run snapcraft as root for arm64 host-mode snap builds
On arm64 electron-builder has no template snap, so it builds without one:
snapcraft pulls stage-packages and, in host (destructive) mode, runs a bare
`apt-get update`. As the unprivileged runner user that fails:

  E: Could not open lock file /var/lib/apt/lists/lock - open (13: Permission denied)
  Failed to refresh package list: failed to run apt update.

amd64 never hits this because it uses the template snap and runs no apt,
which is why "snap worked before" held only for amd64.

Add a PATH shim in every release-linux-arm64 job that runs ONLY snapcraft
as root, then chowns what it wrote back to the runner user. Every other
target (deb/rpm/AppImage/flatpak/tar.gz) is untouched.

Proven on ubuntu-24.04-arm with an electron@38.2.2 / electron-builder@26.0.3
harness using the apps' exact snap config (base core22):
  without the shim: 12/12 jobs fail at "pull app" (runs 35896229973,
                    35902920526, 35903381014, 35904503217)
  with the shim:     2/2 jobs produce the .snap, owned by runner (run 35905018543)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 20:52:25 +02:00
Ruslan KonviserandClaude Opus 5.5 68c4a65998 fix(ci): pre-install snapcraft's build snaps with sudo on arm64
The linux-arm64 snap target has failed intermittently since June 2026 with

  Error installing snap 'gnome-3-28-1804' from channel 'latest/stable'
  -> ERR_ELECTRON_BUILDER_CANNOT_EXECUTE

while x64 passed. In host (destructive) mode snapcraft installs any build snap it is
missing by running `snap install` as the unprivileged runner user, and on the arm64
images snapd intermittently refuses that:

  as the runner user:  error: access denied (try with sudo)
  with sudo:           installed

Proven on branch exp/arm64-snap-snapcraft-version with a minimal Electron app using
the same electron-builder 26.0.3 / Electron 38.2.2 / snap base core22:

- Round 1 (35895302909): arm64 failed on all four snapcraft sources, x64 passed on
  all four - so NOT the snapcraft version, and not the runner (the identical error
  hit ubicloud arm64 in July), nor host mode (on when arm64 last passed in May).
- Round 2 (35895798132): captured snap's own refusal above; also showed the denial is
  intermittent, and that x64 does NOT have the snaps pre-installed - snapd simply
  authorises the runner user there.
- Round 3 (35896229973): 4 baseline vs 4 pre-install replicas. Every baseline run
  made snapcraft perform its own `snap install` (the operation that gets denied);
  every pre-install run made ZERO. The failing path is therefore never reached -
  a fix by construction, not by pass rate.

Adds one step before Multipass in the release-linux-arm64 job of all 12 app
workflows (6 prod + 6 stage, identical text, so prod/stage parity is preserved),
installing exactly the list that was proven. x64 jobs are untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:40:38 +02:00
Ruslan KonviserandClaude Opus 5 0be2422e41 chore(ci): align the last step-name drift between agent-prod and agent-stage
agent-prod.yml said "Bump agent version" while agent-stage.yml said "Bump Agent
version". Display-name only, zero behavioural effect - but it was the single
remaining non-environment difference between the pair, so the two files now differ
by nothing except genuine environment values (branch, tag, prerelease flag).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 17:06:19 +02:00
Ruslan KonviserandClaude Opus 5 2147134daa fix(ci): bring every app workflow's -prod and -stage variants to parity
The `-prod.yml` and `-stage.yml` variants of the six app workflows had drifted:
fixes were applied to one variant and never mirrored to its sibling. This is FILE
drift, not branch drift - both files are identical on develop/stage/master/apps -
so a prod release simply never got work that stage had, and vice versa.

Ported stage -> prod (prod was missing all of this):

1. macOS signing. `Import Apple signing certificate into a keychain` plus the
   switch from CSC_LINK to CSC_KEYCHAIN. app-builder-lib passes the p12 password
   to `security set-key-partition-list -k`, which wants the KEYCHAIN password, so
   every prod mac build died on `SecKeychainUnlock: the user name or passphrase
   you entered is not correct`. The credentials were always correct - stage proved
   it by building green on 2026-09-21 with the same secrets.

2. Windows code signing, entirely. The Bump-step group (WINDOWS_PUBLISHER_NAME,
   AZURE_CERT_PROFILE_NAME, AZURE_CODE_SIGNING_ACCOUNT/ENDPOINT) that makes the
   bump script emit build.win.azureSignOptions, and the Build-step group
   (WIN_CSC_LINK, WIN_CSC_KEY_PASSWORD, AZURE_*). Prod Windows installers have
   been shipping UNSIGNED, and with no publisherName in package.json,
   electron-updater signature verification was off for the prod channel only.

3. `Install .NET SDK (required by Azure Trusted Signing)`. Without it
   Invoke-TrustedSigning reports sdk-not-found, SKIPS signing, and the job goes
   GREEN with an unsigned artefact - so fixing 2 without this would have looked
   successful and still shipped unsigned.

4. The ARM64 Visual Studio Build Tools guard: vswhere probe, conditional choco,
   tolerant exit. Prod had a bare `choco install` that dies on a chocolatey 504.

Ported prod -> stage (stage was missing these, so stage could not reproduce the
prod Linux packaging path):

5. snapcraft pinned to 7.x/stable. 8.0+ renamed `snap` to `pack` while
   app-builder still calls `snapcraft snap`.
6. The Python 3.11 pin for native rebuilds, with npm_config_python/PYTHON exported.

ARM64 Windows signing was deliberately NOT added: stage omits it on purpose
(Microsoft.Trusted.Signing.Client ships only bin/x64 and bin/x86), and that
reasoning is carried across in the comments rather than re-derived.

Deliberately NOT changed: docker-build-publish-prod.yml resolves the version from
a tag pointing AT the built commit with no `git describe` fallback. An audit
flagged that as missing stage's two-tier resolution, but the code documents it as
intentional - a prod image must carry a real release tag or none, never
"v111.32.3-4-gbb20466". The nine deploy/release pairs have no drift at all.

Verified: all 12 files parse under js-yaml; every pair now has identical job and
step counts (100/100, 110/110, 98/98); all 7 feature markers match across all six
pairs; no bare `CSC_LINK:` remains anywhere; zero CRLF; exactly 12 files touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 16:23:20 +02:00
Ruslan KonviserandClaude Opus 5 f2aeadc163 fix(ci): point release-linux-arm64 at a runner that exists
`release-linux-arm64` is the only app job whose matrix hardcodes a runner with no
`vars.` fallback: `os: [ubicloud-standard-8-arm]`. Every sibling job uses the
`${{ vars.X || 'default' }}` form. ever-co's ubicloud ARM capacity is gone, so the
job never gets a runner - all six app workflows on the 2026-09-22 `apps` promotion
sat queued from 23:12Z with 0/6 arm64 builds.

`timeout-minutes: 300` does NOT rescue this: that clock only starts once a job gets
a runner. An unstarted job waits until GitHub cancels it at ~24h, so the RUNS never
reach a terminal state either, and the run-level conclusion becomes meaningless.

Now `${{ vars.RUNNER_LINUX_ARM64 || 'ubuntu-24.04-arm' }}`, matching the shape of
the x64 line two jobs above it. The variable is not currently set at repo or org
level, so this resolves to GitHub's hosted `ubuntu-24.04-arm` - free for public
repositories, and ever-gauzy is public. Setting RUNNER_LINUX_ARM64 later moves every
one of these jobs to a self-hosted pool without another code change.

Applied to all six *-prod.yml and all six *-stage.yml so stage-apps cannot hit the
same wall. The job is repointed, not removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 05:08:15 +02:00
Ruslan Konviser f96b5d4ba2 ci(codeql): a pull request into develop no longer starts an analysis
The owner's instruction was that a pull request into `develop` must not start a workflow — the cost
belonged to every commit on the branch being merged, and the run belongs to the merge. Every other
workflow that fired on such a pull request was changed on `feat/platform-extensions`; this one could
not be, because `codeql.yml` exists only on `develop`.

Measured before the change, over the previous twelve CodeQL runs: **178 run-minutes in total, 137 of
them on pull requests**, at two jobs of 22–45 minutes each on the 4-vCPU ARC pool. It is paid on every
contributor's branch into develop, not on one: `feat/platform-extensions` (44.5 and 29.0 minutes),
`develop` (22.2), `fix/custom-dashboard-drag-and-drop`, `feat/9873-restrict-agent-exit-logout`.

What stays: the `push` trigger still scans `develop`, `stage` and `master`, so code that lands on any
of them is analysed — the run moves to the merge rather than disappearing. The weekly `schedule` scan
and `workflow_dispatch` are untouched, the `concurrency` block is untouched, and pull requests aimed
at `stage` and `master` — the release cascade — still run it.

`branches-ignore` rather than a `branches` list, because a list is what has to be edited when a new
protected branch appears and this is the exception to it. `build.yml` and `secrets-analysis.yml`
already spell their develop exception this way. The block carries the reversed spelling in a comment,
so restoring the trigger is one paste.

The trade-off, stated plainly: a pull request into `develop` is no longer scanned *before* merge. The
security signal moves to the push run on develop, which is after merge. If pre-merge scanning on
develop is wanted more than the runner time, this change should be closed rather than merged — the
alternative that keeps both is to leave the trigger and accept roughly one 30-minute run per push.
2026-09-22 17:51:42 +02:00
Ruslan KonviserandClaude Opus 5 58bf570f64 ci: stop concurrent CodeQL and duplicate gate runs on PRs and develop/stage (#10259)
* ci: stop concurrent CodeQL and duplicate gate runs on PRs and develop/stage

CodeQL ran through GitHub's code-scanning DEFAULT SETUP, which generates its
workflow server-side (run path `dynamic/github-code-scanning/codeql`). There is
no file to put a `concurrency:` block in, and GitHub does not supersede its own
default-setup runs, so every push to a pull request started another full
analysis alongside the ones already in flight. Measured on #10254 on
2026-09-21: runs 35547460875 and 35548451790 overlapped by 12 minutes, and
35580875040 / 35582701347 / 35582925033 were all running at 09:22. Each run is
two jobs of 25-35 minutes on the 4-vCPU ARC pool.

Converts CodeQL to advanced setup with a real workflow file that reproduces the
default-setup configuration exactly - the two analyses it actually ran
(javascript-typescript, actions; the other three listed languages are aliases),
the default query suite, the default remote threat model, the weekly schedule,
and the ever-k8s-linux-x64-4 pool - and adds the cancel-in-progress policy that
default setup could not have. Branch coverage stays at develop/stage/master
because all three are protected and default setup was scanning all three.

Also closes the remaining concurrency gaps on the PR and develop/stage surface:

- build.yml and secrets-analysis.yml both pair an unfiltered `pull_request:`
  with a `push:` on the same branches, so while a release-cascade PR is open
  (today #10257, stage -> stage-apps) one push fires two runs of the same tree
  in two groups that could never cancel each other. Keying the group on the
  head ref collapses them; qualifying it by head repository keeps a fork branch
  of the same name from claiming the group and cancelling a real build before
  its job skips.
- harvest-secrets-to-openbao.yml was the only workflow of 70 with no
  concurrency block, and two overlapping dispatches can interleave into a torn
  OpenBao snapshot.
- external-uptime-monitor.yml now cancels superseded PR self-tests while still
  never cancelling a live probe, which owns the alert issue.

Deliberately unchanged: the deploy and release workflows keep
`cancel-in-progress: false`, and test-unit.yml keeps it on the stage gate. Each
documents an incident where cancelling caused real harm.

Requires default setup to be disabled first - GitHub rejects CodeQL SARIF
uploads from a workflow while default setup is enabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: collapse duplicate CodeQL runs and teach cspell the OpenBao name

Two review follow-ups on this branch.

Greptile P2: codeql.yml had the same duplicate-run defect this PR fixes in
build.yml and secrets-analysis.yml. Both of its triggers cover
develop/stage/master, so while a cascade pull request is open (develop ->
stage), one push to develop fires a push run on refs/heads/develop and a
pull_request run on refs/pull/N/merge. Keyed on github.ref those are different
strings, so cancel-in-progress could never collapse them and the same tree was
analysed twice. Now keyed on the head ref and qualified by head repository, the
same shape used in the other two files.

Cspell: this PR is the first change to harvest-secrets-to-openbao.yml since the
spellcheck was added, and the action only scans changed files - so `openbao` had
never been checked before and surfaced as an unknown word. Added to .cspell.json
next to the other product names.

Not taken: CodeRabbit asked for the actions in codeql.yml to be pinned to commit
SHAs. Declined as inconsistent rather than wrong - this repository tag-pins 547
`uses:` references against a handful of SHA pins, and there is no dependabot
config to keep SHAs current, so pinning one new file would leave an outlier that
nothing updates. Worth doing repo-wide as a deliberate policy, not here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: use US spelling in the codeql concurrency comment

Cspell flagged `analysed` at codeql.yml:60. This repository is US English
throughout (the job is named `Analyze`), so the comment is corrected rather
than the word added to the dictionary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 11:45:19 +02:00
Ruslan KonviserandClaude Opus 5 70b1a85606 fix(security): harden the per-account login control and the Redis throttler after review
This round addresses the review findings on the brute-force and seed-credential fix.

LoginAttemptService now works in three steps: begin(), then fail(), succeed() or release(). The
counters live in EVER_REDIS_CLIENT, and each step runs as one MULTI. A 250 ms deadline sends a
slow step to an in-process store instead, and a Redis error or a disconnected client does the
same, so the control never fails open. Each account may have at most AUTH_MAX_FAILED_ATTEMPTS
checks in flight at once. A hard block needs that many failures from at least two distinct client
sources, using the same spoof-resistant tracker as the route throttler, so one client can no longer
lock a known email out. A failure that cannot be traced to an address counts against the account
(fail closed). A failure that was already in flight when a block landed does not extend it. The
new 429 carries a Retry-After header. Only a credential verdict counts as a failure: an
infrastructure error in login, workspace sign-in or magic-code sign-in releases the slot, and so
does a team-join request that carried no code at all.

RedisThrottlerStorage reads the block marker, the increment and the TTL in one MULTI snapshot. It
sets the block with NX and never asks for a PX of 0. It sends nothing to a client that is not
ready, so node-redis cannot replay those commands after the request has fallen back.

The tracker now expands IPv6 addresses to their canonical form. IPv4-mapped addresses, in hex or
dotted form, collapse to IPv4, and every other address is bucketed by its /64. The docs plugin
tracker no longer uses the client-supplied tenant-id header. parseNonNegativeInt rejects a value
that only starts with digits.

The seed guard now reads environment.demoCredentialConfig, which is what the seeder hashes,
before it reads process.env. It refuses the ever and all seeds in production because they create
the hard-coded fixture accounts. configure.ts writes a seed password into the public web bundle
only for a demo build or when the password is a published default. The Cloudflare compose
templates set THROTTLE_TRUST_CF_CONNECTING_IP. The Playwright job pins its seed passwords. The
env and README docs describe the multi-source rule and the AUTH_MAX_FAILED_ATTEMPTS=0 escape
hatch.

The verification pass renamed identifiers that cspell rejected (IPV6_TRACKER_PREFIX_HEXTETS to
IPV6_TRACKER_PREFIX_GROUPS, and zset to sortedSet in the Redis fake) and reworded one comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 05:50:09 +02:00
Ruslan KonviserandClaude Opus 5 5ada11e5ce ci: never run workflow jobs for pull requests from forks
Owner decision (2026-09-16): fork pull requests must not run on our CI.
Every job in the seven pull_request-triggered workflows (build,
static-checks, typos, snyk-analysis, secrets-analysis, mega-linter,
external-uptime-monitor) now skips when the pull request's head repository
is not ever-co/ever-gauzy. Pushes and pull requests from branches of this
repository run exactly as before. Jobs whose if used !cancelled() or
always() get the guard inside the same expression, so they do not run
after the first job is skipped; build-desktop already runs only on
develop/stage/master pushes.

This is one of three layers: the repository already requires approval for
all external contributors, and the ARC runner image now refuses fork pull
request jobs in a job-started hook (ever-co/arc-runner-image), which a fork
cannot edit away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 20:07:08 +02:00
Ruslan Konviser 625d621f19 feat(integrations): complete organization-scoped Ever Async connector and builds 2026-09-07 12:58:55 +02:00
Ruslan Konviser cb4e276e04 Merge pull request #10124 from ever-co/codex/fix-stage-uuid-migration
fix(database): unblock stage tenant migration on PostgreSQL UUIDs
2026-09-07 11:59:41 +02:00
Ruslan KonviserandClaude Opus 5 be59bce929 CI: add a jest-config typecheck gate and a non-blocking lint report (#10119)
Nothing in CI ran `tsc --noEmit` or ESLint. That is why a duplicated `transformIgnorePatterns`
key sat in `packages/core/jest.config.ts` undetected (TS1117, fixed in #10116), and why ESLint
had been broken repo-wide long enough for 93 of 93 projects to fail (fixed in #10117). Both
gaps are now covered, deliberately with very different postures.

`typecheck-configs` is a real gate. `tools/tsconfig.jest-configs.json` type-checks all 95
jest.config.ts files and is green today, so the job fails if that stops being true. It is built
to need nothing but the compiler — no `extends`, `types: []`, `noResolve`, and a small shim in
`tools/jest-config-globals.d.ts` for the 53 CommonJS configs plus the one `@nx/jest` import in
the root config. The compile is 0.5s across all 95 files; the whole job measured 68s in CI,
dominated by runner provisioning. A cold `yarn install` on this fleet is measured in hours, so
avoiding one is the entire design constraint.

Proven, not assumed: exit 0 on the current tree, and exit 2 with TS1117 when a duplicate key is
reintroduced.

`lint` is NOT a gate. It runs `nx run-many -t lint` and writes the outcome to the run summary.
The repair in #10117 left roughly 2,650 errors and 6,100 warnings, all predating the job, so a
blocking lint gate would wedge every PR on day one. Non-blocking at both the step level (so the
summary still runs) and the job level (so a cache miss or install failure on the fleet cannot
redden the check either). Verified end to end on this PR: ESLint exited 1 with real findings,
the summary step ran, and the check reported green.

TypeScript is installed pinned with `--ignore-scripts` rather than fetched through `npx --yes`,
which SonarCloud correctly flagged as a vulnerability twice (S6505/S8543) for running a
network-fetched package's lifecycle scripts in CI. It is also faster.

Neither job is wired into branch protection; `develop` has no required status checks at all, so
these are visible signals rather than hard gates. The tracking task for burning down the ESLint
backlog lives in the agent workspace at `knowledge/TASKS.md`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 16:44:16 +02:00
Ruslan Konviser fe004fe999 fix(downloads): run the R2 mirror on a GitHub-hosted runner
The self-hosted ARC runners cannot reach Cloudflare R2. Measured from inside a
running ever-k8s-linux-x64-4 pod: the R2 S3 endpoint and downloads.ever.co both
time out, while api.github.com, registry-1.docker.io and cloudflare.com are all
reachable. Forcing IPv4 does not help, so this is not the AAAA-first problem but
the ISP-level blackhole of certain Cloudflare anycast prefixes this fleet has hit
before -- the same fault that forced cdn.ever.co in-cluster.

That made the first backfill run fail its signing self-test with "fetch failed",
which is exactly what the self-test is for: it stopped before publishing an empty
manifest that would have silently reverted every download link to GitHub.

A narrow, documented exception to self-hosted-only: this job exists to talk to
Cloudflare, which a self-hosted runner structurally cannot do. ever-gauzy is
public, so GitHub-hosted minutes are free.
2026-08-30 21:00:34 +02:00
Ruslan Konviser 9a61fe7591 Merge pull request #10072 from ever-co/feat/r2-downloads-mirror
feat(downloads): mirror desktop installers to R2 behind downloads.ever.co
2026-08-30 20:43:00 +02:00
Ruslan Konviser cc4738efaa Merge pull request #10054 from ever-co/fix/configure-registry-git-stderr
fix(ci): surface git's stderr in the configure-registry reset guard
2026-08-30 20:11:06 +02:00
Ruslan Konviser f609997d04 feat(downloads): mirror desktop installers to R2 behind downloads.ever.co
Adds .github/workflows/mirror-releases-to-r2.yml.
2026-08-30 19:10:35 +02:00
Ruslan Konviser 18816d31ab test(ci): avoid eval in registry action harness 2026-08-30 18:53:16 +02:00
Ruslan Konviser c1b3ca3d9c fix(ci): encode workflow command line controls 2026-08-30 18:48:51 +02:00
Ruslan KonviserandClaude Opus 5 5aa7b27392 ci: use the real nx flag and drop the checkout credential
From review on #10055.

`--continue` is not an nx option. nx 22 spells it `nxBail`, and an unknown
option is forwarded to the executor - so `--continue` would have reached jest
rather than enabling failure continuation, and the first failing project would
still have hidden every other project's result. Now `--nxBail=false`.

Also sets `persist-credentials: false` on checkout: this job only reads the
tree and never pushes, so the credential does not need to outlive the step.

Action majors stay aligned with build.yml on purpose - this job consumes the
node_modules cache that workflow writes, so producer and consumer are kept on
the same actions rather than following the newer majors used elsewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 00:30:12 +02:00
Ruslan KonviserandClaude Opus 5 9b3ada96ea ci: run the existing jest unit tests on stage
The repository has ~330 `*.spec.ts` files and an nx `test` target on 20+
projects (`@nx/jest:jest`, already cached via nx.json targetDefaults), but no
workflow ever invoked any of them. CI ran builds plus Playwright/Cypress e2e
only, so a green pull request proved the code COMPILED and said nothing about
the unit tests - including the ~53 specs in packages/core.

This wires the existing targets up. It adds no test tooling and changes no
test code.

Scope is `stage` only, plus manual dispatch. Stage is where the cascade lands
before production and pushes to it are infrequent, so this cannot slow the
day-to-day loop while the suites' true state is still unknown.

Deliberately NOT a required status check: these suites have never run in CI so
the first runs may be red, and that is exactly the information the workflow
exists to surface - it should surface it without blocking the cascade. Promote
it to required once it has been green for a while.

`cancel-in-progress: false`, because a cancelled test run produces no verdict
and a gate that reports nothing is worse than no gate.

Dependencies come from the same lockfile-keyed cache the build workflow
writes, so a stage push that already built this tree reuses it rather than
paying for another install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 00:14:25 +02:00
Ruslan KonviserandClaude Fable 5 b1ecd32b16 fix(ci): surface git's stderr in the configure-registry reset guard
On 2026-08-27 three self-hosted Windows runners failed the reset guard with
bare '(git-exit-128)' tags and no cause - the probe discarded stderr, and
diagnosing the real error ('fatal: detected dubious ownership' after the
runner service account changed out from under the workspace) took a remote
login to the machines.

The probe now captures stderr (stderr into the substitution, stdout dropped)
and the first stderr line rides along in the ::error annotation, re-encoded
so the runner's percent-decoding cannot mangle it. A failing 'git checkout'
gets the same treatment and its full stderr is still re-emitted to the log.
Exit-status handling, per-path tracked/untracked semantics and the
fail-closed behavior are unchanged; verified under bash -eo pipefail across
tracked/untracked/absent/pnpm-shape/broken-repo/silent-failure fixtures,
with the old code as a known-dirty control.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 21:32:03 +02:00
Ruslan Konviser b6287c86b7 fix(build): use bundled Node headers for native modules 2026-08-23 23:04:54 +02:00
Ruslan Konviser 901af20ecc Merge pull request #10035 from ever-co/ci/verdaccio-registry-cypress-currents
ci(verdaccio): route the Cypress and Currents e2e installs through the internal cache

Closes the last two dependency-installing workflows in this repo that still
resolved from the public registry. Adversarially verified clean; 23/23 green.
Merged on owner instruction (2026-08-23).
2026-08-23 08:39:34 +02:00
Ruslan KonviserandClaude Opus 5 51b69031fa fix(configure-registry): restore a TRACKED .yarnrc instead of deleting it
The reset assumed ".yarnrc in particular is untracked and un-ignored". That holds
in ever-gauzy but is FALSE in ever-co/ever-teams, where .yarnrc is tracked and
carries yarn-path ".yarn/releases/yarn-1.22.22.cjs".

So `rm -f .yarnrc` destroyed the repo yarn version pin before
`yarn install --frozen-lockfile` ran, leaving the install to use whatever yarn the
runner happened to resolve. Any consumer repo that tracks .yarnrc has the same
exposure.

Fix uses the mechanism already directly below: the per-path reset loop restores a
tracked file with `git checkout --` and still `rm -f`s a genuinely untracked stale
leftover, failing closed on git exit 128. Moving .yarnrc out of the unconditional
rm -f list into that loop preserves the original behaviour everywhere .yarnrc is
untracked, and stops the data loss where it is tracked.

Found by adversarial review of ever-co/ever-teams#4433, which migrates 20 jobs onto
this action.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 22:15:40 +02:00
Ruslan KonviserandClaude Opus 5 a6ab2fc9e3 ci(verdaccio): route the Cypress and Currents e2e installs through the internal npm cache
These were the last two dependency-installing workflows still resolving packages
straight from the public npm registry. Both are covered now, using the same
placement discipline as build.yml.

The step goes AFTER the yarn cache restore and BEFORE `yarn bootstrap`, because
the action rewrites yarn.lock's resolved URLs (the load-bearing half - yarn v1
fetches the URL recorded in the lockfile and ignores the registry setting) while
hashFiles() in a cache key evaluates at step runtime. Configuring the registry
first would move the key and bifurcate the cache namespace by VIP reachability.

No save-key pinning was required. build.yml needs it because it uses the split
actions/cache/restore + actions/cache/save pair, whose save step re-derives
hashFiles() at save time. Both files here use the COMBINED actions/cache, which
records its primary key with core.saveState during the restore - before the
rewrite - and whose post-job save reads that recorded key back, so the entry is
written under exactly the key the next run's restore computes.

Scope, and what was deliberately left alone:

- test_currents.yml `e2e-tests` and test_cypress.yml `e2e-tests` get the step.
  Both resolve to the self-hosted ARC pool via vars.RUNNER_LINUX_X64_8
  (= ever-k8s-linux-x64-8), the only place the VIP answers. The action probes
  the VIP and falls back to the public registry, so the `|| 'ubuntu-latest'`
  fallback runner degrades to today's behaviour rather than hanging on a LAN IP.
- test_currents.yml `prepare` installs nothing - it only emits a build UUID.
- test_cypress.yml `e2e-tests-setup` overrides its install with
  `install-command: node -p 'os.cpus()'`, so it resolves no packages from any
  npm registry. Adding the step there would buy nothing and would rewrite
  yarn.lock ahead of cypress-io/github-action's own lockfile-keyed cache,
  splitting it in two.

`expect-vip` is derived from the runner variable rather than the
contains(matrix.os, ...) idiom the app workflows use: these jobs' only matrix
key is `containers`, so that expression would collapse to false and silently
disable the in-network VIP retry and its warning.

Additive only - no workflow, job or step was removed or reordered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 21:50:01 +02:00
Ruslan KonviserandClaude Fable 5 a81dbc073a perf(ci): route the build gate's installs through Verdaccio - with the key trap handled
build.yml was the one workflow deliberately skipped by #10025, because the
configure-registry action rewrites yarn.lock and this file keys its dependency
cache on hashFiles('yarn.lock', ...) - a naive placement moves the key and
bifurcates 2.9 GB cache entries by VIP reachability.

The deferral had a measured cost tonight: on stage run 32513912629, cleanup had
deleted the run's artifact, the cache entry was LRU-evicted, and the bootstrap
fallback pulled the whole tree from the PUBLIC registry - blowing through the
180-minute job ceiling twice (attempts 2 and 3 of build-api).

Both traps are handled explicitly:
  * Configure Registry sits AFTER every lockfile-keyed cache step - producer:
    after the restore, gated on cache-miss; consumers: after the artifact
    download and cache fallback, so only the bootstrap path pays for it.
  * The producer's save key is pinned to steps.cache.outputs.cache-primary-key
    (computed before the rewrite) instead of re-deriving hashFiles at save time,
    which would write the archive under a different key than the next run
    looks up.
  * expect-vip derives from vars.RUNNER_LINUX_X64_8 - these jobs define no
    matrix.os, so the fleet-standard contains(matrix.os, ...) is always false
    and would silently disable the VIP retry.

Pinned to aa4ee199, the action revision that survives pnpm/npm repos and treats
only git exit 1 as "untracked".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 07:21:45 +02:00
Ruslan KonviserandClaude Fable 5 2a58ca4609 fix(ci): keep the node_modules artifact when the run is red so reruns can restore it
The cleanup job deleted this run's ONLY dependency handoff unconditionally
(`if: always()`). That choice had a real rationale - red-run leftovers once
filled the org's Actions storage quota and stopped artifact uploads - but it
breaks "Re-run failed jobs": with the artifact gone, every rerun consumer falls
into the documented 1-3h bootstrap fallback.

Measured on stage run 32513912629 (2026-08-21): a seconds-long DNS blip
(getaddrinfo EAI_AGAIN on GitHub's artifact endpoints) failed build-web's
restore; cleanup deleted the artifact when the attempt ended; the rerun then
cost build-libs 178 minutes and pushed build-api into the 180-minute ceiling.
Hours of runner time to recover from a transient network error.

Deletion now happens only when no needed job failed or was cancelled. The quota
concern is bounded rather than ignored: red Build runs are rare since the gate
rework, and retention-days: 1 ages any kept artifact out within a day. skipped
results (build-desktop on PR runs) still allow deletion.

This was the greptile/cubic P1 on #10010, catalogued as "deferred - needs a
design call". Production made the call.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 04:19:22 +02:00
Ruslan Konviser aa4ee19926 fix(configure-registry): guard the tracked-file probe from errexit
Composite bash runs with -e, so the bare 'git ls-files ...; tracked=$?'
form kills the step (exit 1, no output) on any repo where the probed file
is untracked - which is every pnpm/npm consumer repo. Measured on the
first cross-repo run (git-hands/githands PR 2170, job 96838259229): the
step died 0.5s after start with no runtime output. Wrap the probe in an
if/else so the exit code is captured instead of tripping errexit; the
0 / 1 / other(fail-closed) distinction is preserved unchanged.

Verified by extracting the run script and executing it under
bash -e -o pipefail against pnpm-shaped, npm-shaped and yarn-shaped test
repos with the VIP reachable, plus the unreachable-fallback path - all
exit 0 with the expected registry outcome.
2026-08-21 18:57:35 +02:00
Ruslan Konviser 9459d29e9a chore: appease cspell (lockfiles -> lock files) 2026-08-21 18:29:35 +02:00
Ruslan Konviser 1e50a73138 fix(configure-registry): guard yarn.lock rewrite + add package-lock.json rewrite for cross-repo use
The composite step runs under bash -eo pipefail, so the unguarded
'sed ... yarn.lock' exits 2 and kills the job on any repo without a root
yarn.lock (pnpm and npm repos) the moment the Verdaccio VIP answers.
Gate it on the file existing, mirroring the reviewed local copy in
ever-works/ever-works.

Also rewrite package-lock.json where present: npm follows the lockfile's
'resolved' URLs, so for npm repos a config-only registry change is a
measured no-op (see registry-config-is-a-noop-with-lockfiles). Integrity
stays sha512-verified. package-lock.json joins the per-path reset loop so
a previous run's rewrite on a clean:false runner is reset the same way
yarn.lock already is.

Additive only - behaviour on this repo (root yarn.lock present) is unchanged.
2026-08-21 18:24:37 +02:00
Ruslan KonviserandClaude Opus 5 911b50bd2b fix(ci): fail closed when git cannot tell us whether a path is tracked
Review finding from CodeRabbit and cubic on this PR, and they are right.

`git ls-files --error-unmatch` returns 0 for tracked, 1 for genuinely absent,
and 128 when git metadata is unavailable. My loop treated every nonzero as
"untracked", so on 128 it would rm -f .npmrc and yarn.lock and carry on - which
is precisely the unknown-state hazard the reset exists to prevent. It swapped
one fail-open for another.

Verified the exit codes rather than trusting the report:
  tracked file     -> 0
  untracked file   -> 1
  outside a repo   -> 128

Only 1 is now trusted as "absent"; anything else lands in reset_failed and the
step aborts with the path and the git exit status.

Also adds "pathspec" and "unmatch" to the cspell dictionary - both are real git
terms introduced by this change, and Cspell was red on them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:45:02 +02:00
Ruslan KonviserandClaude Opus 5 53037b25f7 fix(ci): make configure-registry actually reusable from other repos
Two defects, both found while rolling this action out across the fleet. Neither
shows up in gauzy's own usage, which is why they survived.

1. An expression in action METADATA killed every cross-repo use.

   inputs.verdaccio-registry.description contained a ${...} expression naming
   `vars.VERDACCIO_REGISTRY`. GitHub evaluates action metadata, and `vars` is not
   a valid context there, so any workflow in another repo referencing this action
   died at "Set up job":

     ever-co/ever-gauzy/<sha>/.github/actions/configure-registry/action.yml
     (Line: 16, Col: 18): Unrecognized named-value: 'vars'
     -> TemplateValidationException

   Reproduced on ever-digital/ever-digital run 32475781533: the job previously
   installed dependencies successfully and afterwards executed NOTHING. The
   variable is now named in prose instead.

2. The reset hard-failed on every pnpm repo.

   `git checkout -- .npmrc yarn.lock` exits 1 when EITHER pathspec matches
   nothing, and the action then exits 1 by design. A pnpm repo tracks neither
   file, so the action turned "this repo has no cache configured" into "this
   repo's CI is red" - before it ever selected a registry. Proven with a negative
   control against a simulated pnpm checkout.

   The reset is now per-path: tracked files are restored, untracked leftovers
   from a previous run on a clean:false runner are removed, and only a genuine
   checkout failure aborts.

Behaviour for gauzy's own yarn workspace is unchanged: both files are tracked
there, so each takes the git-checkout branch exactly as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 15:30:15 +02:00
Ruslan KonviserandClaude Opus 5 fb6673c797 perf(ci): serve the Playwright deps install from Verdaccio, not the internet
`test_playwright.yml` never pointed at our internal npm cache, so the heaviest
install in the fleet pulled ~4.2 GB from registry.yarnpkg.com on every run. The
job documents itself at 88 min, 3h20m and 3h46m.

An audit of all 11 orgs found 177 workflows across 56 repos in the same state -
this is the highest-traffic one (222 runs) and the cleanest to fix.

Reuses the existing composite action rather than hand-rolling a snippet, which
matters more than it looks:

  * yarn v1 fetches the URL recorded in yarn.lock and IGNORES the registry
    setting. Measured with a negative control - a blackholed `resolved` URL with
    .yarnrc pointing at a WORKING registry still failed with ENOTFOUND. A
    config-only patch here would be a no-op that looks like a fix: green build,
    same 83 minutes. The action rewrites the lockfile URLs, which is the half
    that actually moves the bytes.
  * Verified the rewrite does NOT break --frozen-lockfile: yarn compares
    package.json patterns against lockfile entries, and the resolved host is not
    part of that comparison. Measured against the real VIP, RC=0, lockfile
    unmodified.
  * The action probes the VIP and falls back to the public registry. That is
    mandatory, not decorative: the VIP is a single MetalLB L2 announcement with a
    recurring speaker flap, so hardcoding it would convert a cache optimisation
    into a CI outage vector.

Two deliberate deviations from how the app workflows call this action:

  * `expect-vip` derives from `vars.RUNNER_LINUX_X64_8 != ''`, NOT
    `contains(matrix.os, ...)`. This workflow defines no `matrix.os` - its only
    matrix key is `shard` - so the standard expression evaluates false and
    silently disables the in-network retry and warning. A mute switch, not an
    error.
  * The step sits immediately after checkout, which is safe HERE because this
    workflow deliberately has no `actions/cache`. In build.yml the same placement
    would move the `hashFiles('yarn.lock')` cache key, so that file needs
    different handling and is not touched here.

The `cleanup` job stays on ubuntu-latest and deliberately does NOT get this step:
a GitHub-hosted runner cannot reach 192.168.1.183 and would hang.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 02:05:21 +02:00
Ruslan KonviserandClaude Opus 5 7118ab32fa fix(ci): stop every consumer downloading the dependency archive twice
The cache-fallback step carried a comment saying "Only reached if the artifact
was unavailable" - but it had no `if:` and the download step had no `id:`, so it
ran unconditionally. The producer saves that same cache key on every cold run, so
the lookup HITS and actions/cache/restore pulls the tree a second time.

Measured live on run 32418672320 (develop, the first run of this design):

  build-api       Download node_modules archive   1015s   success
                  Restore ... (cache fallback)     273s   success
  build-desktop   Download node_modules archive   1028s   success
  build-web       Download node_modules archive   1102s   success

A cache MISS returns in seconds; 273s is a download. That is ~4.5 minutes of
redundant transfer per consumer, ~18 minutes per run, on the same WAN link the
install already saturates - and it scales with every consumer job added.

Gated on `steps.artifact.outcome != 'success'`. `outcome` and not `conclusion`
on purpose: `continue-on-error: true` masks `conclusion` to success, so only
`outcome` reports a genuine download failure - which is exactly when the cache
fallback should run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 01:26:37 +02:00
Ruslan KonviserandClaude Opus 5 54ed5ebdb3 fix(ci): the install-once producer breaks on every cache HIT
Two correctness bugs from the AI review of #10010, both confirmed against the
merged code on develop.

1. build.yml - every warm run would fail.

   `Restore node_modules archive` used `lookup-only: true`, so on a cache HIT the
   archive is never written to disk. `Install`, `Run postinstall`, `Archive` and
   `Save` are all gated on `cache-hit != 'true'` and correctly skip - but
   `Upload node_modules archive` carries NO `if:` guard and sets
   `if-no-files-found: error`. On a hit it therefore uploads nothing and fails
   the job, and because the other four jobs `needs:` this one, it takes the whole
   gate with it.

   The cold run that first writes the key passes, so this is invisible until the
   SECOND run - the one the cache exists for. Dropping `lookup-only` makes the
   producer always materialise the archive: a ~2.9 GB download costs a few
   minutes against the ~90-minute install it replaces, and it is what guarantees
   the run-scoped artifact handoff actually has a file to hand off.

2. archive-node-modules.sh - a failed tar could publish a corrupt archive.

   The script set only `-o pipefail`, never `-e`, and tolerated tar exit 1 with
   `|| [ $? -eq 1 ]`. A tar exit status of 2 - a real error - left the `||`
   compound false and execution simply continued to the size check. A ~2.9 GB
   tree truncated early still clears the 100 MB floor, so the partial archive
   would be published and unpacked by four consumer jobs.

   Now `set -e`, with the tar status captured explicitly so exit 1 stays
   tolerated and anything else aborts before the archive can be published.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:24:53 +02:00
Ruslan Konviser 07784b0c9d Merge pull request #10010 from ever-co/fix/expense-null-employee-scope
fix(expenses): org-level expenses 400 since the null-where hardening; stop the CI artifact leak
2026-08-20 23:17:44 +02:00
Ruslan KonviserandClaude Opus 5 3c6e34dc00 chore(spell): American spelling in the build workflow comment
cspell is the only thing standing between this PR and merge: it fails on
"optimisation" at build.yml:113. The repo enforces the American spelling and has
corrected this exact word in this exact file before (6c78c4e92e).

No behaviour change - a comment word.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:16:00 +02:00
Ruslan KonviserandClaude Opus 5 bcdbc1ea26 fix(ci): hand the tree over by ARTIFACT — the cache was evicted mid-run
First run of the install-once design failed, and the logs are unambiguous:

  build-monorepo-root: "Cache saved with key: Linux-X64-node-modules-…2b58b651…"  (passed, 1h44m)
  build-libs (32s):    "Failed to restore cache entry. Input key: …2b58b651…"

Byte-identical keys. The producer saved it and the consumer could not read it,
because an Actions cache is a REPOSITORY resource on a 10 GB budget, not a
run-scoped one. While this run was in flight, build.yml on develop, on stage and
on PR #10014 — all still carrying `cache: yarn`, since this change exists only on
this branch — each wrote a 4.66 GB setup-node yarn cache. The repo hit 14.03 GB
and LRU took the newest entry: the tree archive, seconds after it was written.

So the cache cannot be the handoff. It is now best-effort and CROSS-run only.
The handoff is an artifact, which is scoped to the run and cannot be evicted by
another workflow. Consumers try the artifact, fall back to the cache, and only
then to a full install — and `fail-on-cache-miss` is gone, since a cache miss is
no longer the fault it was.

A cleanup job deletes the artifact when the run ends, so it costs storage only
while the run needs it — the lesson from the e2e workflow, where an artifact
left behind on every run filled the org's quota and stopped reports uploading.

Also pruned the three duplicate yarn caches by hand (9.34 GB): they are the same
key on three refs, and this change stops build.yml producing them at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 21:58:11 +02:00