524 Commits
Author SHA1 Message Date
Ruslan Konviser efdba39b6b Merge remote-tracking branch 'origin/develop' into fix/unit-tests-hermetic-harness 2026-09-28 09:03:49 +02:00
Ruslan KonviserandClaude Opus 5.5 e77e3130f4 ci: reword the build-record comment (cspell)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 21:27:58 +02:00
Ruslan KonviserandClaude Opus 5.5 6c68c1fda0 ci: stop uploading Docker build records and summaries
docker/build-push-action uploads a build record (*.dockerbuild) for every image
build. The record stores every build-arg value in plain text, and on a public
repo any signed-in GitHub user can download it. Set DOCKER_BUILD_RECORD_UPLOAD
and DOCKER_BUILD_SUMMARY to false at workflow level in every workflow that uses
the action. BuildKit secrets are never recorded, so nothing else changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 21:02:42 +02:00
Ruslan Konviser d0b2a8e080 Merge remote-tracking branch 'origin/develop' into fix/unit-tests-hermetic-harness 2026-09-27 20:53:10 +02:00
Ruslan KonviserandClaude Opus 5.5 91d91ffca3 test: make the unit-test run hermetic and fix the Angular jest harness
The Unit Tests workflow has never been green (24 of 97 projects red at
develop eac7136562, run 36319732347). Causes addressed here:

- Env leak: Nx loads the committed .env.local (DEMO=true,
  WORKER_QUEUE_ENABLED=false, NODE_ENV=development) into every task.
  The test step now sets NX_LOAD_DOT_ENV_FILES=false, and the two specs
  that depend on a mode pin it themselves (role-permission counts pin
  environment.demo=false and gain a demo-mode suite asserting 209 per
  role; the worker spec clears the queue env before the constants load).
  .env.local is unchanged.
- Angular harness: nohoist gives each workspace its own @angular copy,
  so setupZoneTestEnv() initialised a different TestBed from the one the
  spec used. A workspace resolver (jest.resolver.js, wrapping Nx's) makes
  every @angular/* import in a Jest project resolve to one copy. The
  Angular projects now inherit the preset's transformIgnorePatterns,
  which gains the .mjs exception plus @datorama, @ngneat and lodash-es.
- ui-config's environment.ts is generated and gitignored; the workflow
  runs `yarn config:dev` before the tests, as the Playwright workflow does.
- Misconfigured targets: gauzy (jest.config.js -> .ts, Angular transform,
  setupFile moved into the config), integration-sim-ui (.ts -> .cts),
  integration-activepieces (config file added), mcp-auth (ran
  `node build/main.js --test`; now the jest executor like apps/mcp).
- Real failures: the openai "silence" case used a body with no `text`,
  which the shared helper rejects by contract; the toolbar spec read the
  tabIndex property, which is 0 on any button; the docs-ui linearity
  tests asserted wall-clock bounds, now a 4x-input growth ratio.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 17:39:40 +02:00
Ruslan KonviserandClaude Opus 5.5 40826df61b ci: unpack the lint cache off the RAM disk; PRs into develop start no workflow
Static Checks / lint (x64-4 lane): every cache hit since the job moved to
this lane ran out of space. The 2.47 GB archive plus the tree it unpacks to
does not fit the 16Gi RAM-backed workspace, so each hit fell back to a
1.5-4 h cold install (16 of 16 runs cold since #10303). The Restore step
now moves the archive onto the disk-backed package-cache volume (the
parent of YARN_CACHE_FOLDER) before extracting, so only the tree lives in
RAM. Where YARN_CACHE_FOLDER is unset, or the move fails, it extracts in
place exactly as before. Two report-only df lines (after the restore and
at the top of Summarize) record workspace headroom and can never fail a
step.

Triggers: apply the owner decision of 2026-09-21 (already on draft #10254,
same text) directly to develop. A pull request INTO develop no longer
starts static-checks, typos (cspell), snyk-analysis (trigger removed,
restore lines kept in a comment), or build, secrets-analysis and the
external-uptime-monitor self-test (branches-ignore: develop, so PRs into
stage/master still run them). Every push trigger is unchanged, so all of
them still run when code lands on develop. develop has no required status
checks, so no PR is blocked by the missing runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 15:11:58 +02:00
Ruslan KonviserandClaude Opus 5.5 cec254cd35 ci(desktop): release prod snaps to the Snap Store stable channel
Every Gauzy app snap (desktop, desktop-timer, server, api-server, agent,
mcp-server; amd64 and arm64) went to the Snap Store 'edge' channel only,
from both the stage lane (stage-apps) and the prod lane (apps). The
electron-builder snap target publishes with its own snapStore config.
When the build config has no snapStore entry, that config has no
channels, and the Snap Store publisher then defaults to 'edge'.

Prod builds now release to 'stable' and stage builds to 'edge'. Each
Linux build step appends
-c.snap.publish.provider=snapStore -c.snap.publish.channels=<channel>
to its yarn command. yarn adds extra arguments to the end of the script,
which is the electron-builder call. The override is in the workflows
because the root build scripts are shared by both lanes. Only
snap.publish is overridden: a -c.publish override would be merged into
the GitHub publish entry. GitHub releases, update channels and every
other target are unchanged. Stage names 'edge' explicitly so each
prod/stage pair still differs only in the channel.

The generated snaps already use grade stable and strict confinement,
and they need no store-approved plugs, so the store accepts them on
stable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 01:10:55 +02:00
Ruslan Konviser fb2184e41f Merge pull request #10305 from ever-co/fix/desktop-update-followups
fix(desktop): update-feed follow-ups: publish-channel CI guard, updater tag re-resolve, local-update note, stray channel
2026-09-25 00:53:23 +02:00
Ruslan KonviserandClaude Opus 5.5 271b43e24e ci(desktop): fail win/linux releases whose dist package.json lacks the per-arch update channel
Gauzy Server's Windows and Linux release jobs published channel-less
update manifests (latest.yml, latest-linux*.yml) from v107 to v111.44.45.
Its :isolated scripts ran `pack --arch` after the build had already copied
apps/server/src/package.json into dist. electron-builder reads
dist/apps/<x>/package.json, never saw build.publish[].channel and fell
back to `latest`. The updater only requests `latest-${process.arch}`, so
installs stopped auto-updating, and nothing caught it for 4.5 months.
#10299 fixed that app. This change makes the whole class fail loudly.

Add .scripts/electron-package-utils/assert-publish-channel.js, a Node
script with no dependencies. It takes --project <dist dir> --arch
<x64|arm64> and requires every build.publish[] entry in
<dist dir>/package.json to have channel latest-<arch>. Otherwise it prints
a GitHub `::error` annotation naming the file and the expected and actual
channels, plus a hint, and exits 1. A missing file, unreadable JSON or bad
arguments also exit 1.

Run it immediately before electron-builder in the 28 root scripts that
publish for --windows/--linux and stamp the channel with `pack --arch`.
The list was found from the scripts, not typed by hand. Each guard's
--project and --arch are asserted equal to that script's electron-builder
--project, its --x64/--arm64 flag and its pack --arch. A bad channel now
stops the job before anything is published. No other script changes.
The 24 win/linux publishing scripts without `pack --arch` have no callers
in any workflow and are left alone.

Add assert-publish-channel.test.js (node --test, no dependencies) as
`yarn test:publish-channel` and run it in build.yml next to
test:postinstall. It covers the pass, fail and misuse cases. It also
checks that every per-arch win/linux publishing script keeps the guard on
the same arch and dist dir, so a new or edited release script cannot drop
it unnoticed. That check fails against the current develop package.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-25 00:42:07 +02:00
Ruslan KonviserandClaude Opus 5.5 5ab57761ba ci(static-checks): derive the lint job's expect-vip from x64-4 to match its runs-on
Moving `lint` to `RUNNER_LINUX_X64_4` left its Configure Registry step's
`expect-vip` still derived from `RUNNER_LINUX_X64_8`. Flagged by both
greptile-apps and coderabbitai on line 74.

`expect-vip` tells the pinned configure-registry action (aa4ee199) that the
runner is in-network: it probes the Verdaccio VIP twice instead of once and
emits "Verdaccio VIP unreachable from an in-network runner" on failure. Either
way the job falls back to public npm and never fails - it is a diagnostic.

Every other job in the repo derives `expect-vip` from the same variable as its
`runs-on` (all nine x64-8 jobs read x64-8; the matrix release jobs use
"self-hosted or ever-k8s"). `lint` was the only mismatch, introduced by the
previous commit. This restores the invariant.

Took greptile's fix (derive from x64-4). Did NOT take coderabbit's literal
`false`: x64-4 is an in-network ARC runner, so `false` would make `lint` the one
self-hosted job with no retry and no warning, turning a registry outage on that
lane into a silent single-probe fallback.

No behaviour change today, since both variables are set org-wide; this closes
the case where x64-4 is configured and x64-8 is not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 21:09:41 +02:00
Ruslan KonviserandClaude Opus 5.5 2d6036e1b9 ci(static-checks): run the non-blocking lint job on the x64-4 lane
The `lint` job is report-only (`continue-on-error: true`) and runs
`nx run-many -t lint --parallel=2`, so it never uses more than two parallel
tasks and gains nothing from the x64-8 lane's 12-core limit. Its sibling in the
same workflow, `typecheck-configs`, already runs on x64-4.

The x64-8 lane is capped at 12 runners and each one requests 28 Gi, against
12 Gi for x64-4. It was observed saturated at 12/12 on 2026-09-22 while Docker
image builds queued behind it. Moving a job that cannot use the extra cores
frees a slot and 16 Gi of requested memory for the builds that can.

Deliberately NOT moved:
- unit-tests: jest forks workers per project, so fewer cores would slow a job
  that already takes 123-233 minutes and risk its 360-minute timeout.
- the e2e suites (playwright / cypress / currents): timing-sensitive, and a
  slower runner makes that class of test flakier, not just slower.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 20:15:54 +02:00
Ruslan KonviserandClaude Opus 5.5 f6fe23a4bf fix(ci): run snapcraft as root for arm64 host-mode snap builds
On arm64 electron-builder has no template snap, so it builds without one:
snapcraft pulls stage-packages and, in host (destructive) mode, runs a bare
`apt-get update`. As the unprivileged runner user that fails:

  E: Could not open lock file /var/lib/apt/lists/lock - open (13: Permission denied)
  Failed to refresh package list: failed to run apt update.

amd64 never hits this because it uses the template snap and runs no apt,
which is why "snap worked before" held only for amd64.

Add a PATH shim in every release-linux-arm64 job that runs ONLY snapcraft
as root, then chowns what it wrote back to the runner user. Every other
target (deb/rpm/AppImage/flatpak/tar.gz) is untouched.

Proven on ubuntu-24.04-arm with an electron@38.2.2 / electron-builder@26.0.3
harness using the apps' exact snap config (base core22):
  without the shim: 12/12 jobs fail at "pull app" (runs 35896229973,
                    35902920526, 35903381014, 35904503217)
  with the shim:     2/2 jobs produce the .snap, owned by runner (run 35905018543)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 20:52:25 +02:00
Ruslan KonviserandClaude Opus 5.5 68c4a65998 fix(ci): pre-install snapcraft's build snaps with sudo on arm64
The linux-arm64 snap target has failed intermittently since June 2026 with

  Error installing snap 'gnome-3-28-1804' from channel 'latest/stable'
  -> ERR_ELECTRON_BUILDER_CANNOT_EXECUTE

while x64 passed. In host (destructive) mode snapcraft installs any build snap it is
missing by running `snap install` as the unprivileged runner user, and on the arm64
images snapd intermittently refuses that:

  as the runner user:  error: access denied (try with sudo)
  with sudo:           installed

Proven on branch exp/arm64-snap-snapcraft-version with a minimal Electron app using
the same electron-builder 26.0.3 / Electron 38.2.2 / snap base core22:

- Round 1 (35895302909): arm64 failed on all four snapcraft sources, x64 passed on
  all four - so NOT the snapcraft version, and not the runner (the identical error
  hit ubicloud arm64 in July), nor host mode (on when arm64 last passed in May).
- Round 2 (35895798132): captured snap's own refusal above; also showed the denial is
  intermittent, and that x64 does NOT have the snaps pre-installed - snapd simply
  authorises the runner user there.
- Round 3 (35896229973): 4 baseline vs 4 pre-install replicas. Every baseline run
  made snapcraft perform its own `snap install` (the operation that gets denied);
  every pre-install run made ZERO. The failing path is therefore never reached -
  a fix by construction, not by pass rate.

Adds one step before Multipass in the release-linux-arm64 job of all 12 app
workflows (6 prod + 6 stage, identical text, so prod/stage parity is preserved),
installing exactly the list that was proven. x64 jobs are untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 19:40:38 +02:00
Ruslan KonviserandClaude Opus 5 0be2422e41 chore(ci): align the last step-name drift between agent-prod and agent-stage
agent-prod.yml said "Bump agent version" while agent-stage.yml said "Bump Agent
version". Display-name only, zero behavioural effect - but it was the single
remaining non-environment difference between the pair, so the two files now differ
by nothing except genuine environment values (branch, tag, prerelease flag).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 17:06:19 +02:00
Ruslan KonviserandClaude Opus 5 2147134daa fix(ci): bring every app workflow's -prod and -stage variants to parity
The `-prod.yml` and `-stage.yml` variants of the six app workflows had drifted:
fixes were applied to one variant and never mirrored to its sibling. This is FILE
drift, not branch drift - both files are identical on develop/stage/master/apps -
so a prod release simply never got work that stage had, and vice versa.

Ported stage -> prod (prod was missing all of this):

1. macOS signing. `Import Apple signing certificate into a keychain` plus the
   switch from CSC_LINK to CSC_KEYCHAIN. app-builder-lib passes the p12 password
   to `security set-key-partition-list -k`, which wants the KEYCHAIN password, so
   every prod mac build died on `SecKeychainUnlock: the user name or passphrase
   you entered is not correct`. The credentials were always correct - stage proved
   it by building green on 2026-09-21 with the same secrets.

2. Windows code signing, entirely. The Bump-step group (WINDOWS_PUBLISHER_NAME,
   AZURE_CERT_PROFILE_NAME, AZURE_CODE_SIGNING_ACCOUNT/ENDPOINT) that makes the
   bump script emit build.win.azureSignOptions, and the Build-step group
   (WIN_CSC_LINK, WIN_CSC_KEY_PASSWORD, AZURE_*). Prod Windows installers have
   been shipping UNSIGNED, and with no publisherName in package.json,
   electron-updater signature verification was off for the prod channel only.

3. `Install .NET SDK (required by Azure Trusted Signing)`. Without it
   Invoke-TrustedSigning reports sdk-not-found, SKIPS signing, and the job goes
   GREEN with an unsigned artefact - so fixing 2 without this would have looked
   successful and still shipped unsigned.

4. The ARM64 Visual Studio Build Tools guard: vswhere probe, conditional choco,
   tolerant exit. Prod had a bare `choco install` that dies on a chocolatey 504.

Ported prod -> stage (stage was missing these, so stage could not reproduce the
prod Linux packaging path):

5. snapcraft pinned to 7.x/stable. 8.0+ renamed `snap` to `pack` while
   app-builder still calls `snapcraft snap`.
6. The Python 3.11 pin for native rebuilds, with npm_config_python/PYTHON exported.

ARM64 Windows signing was deliberately NOT added: stage omits it on purpose
(Microsoft.Trusted.Signing.Client ships only bin/x64 and bin/x86), and that
reasoning is carried across in the comments rather than re-derived.

Deliberately NOT changed: docker-build-publish-prod.yml resolves the version from
a tag pointing AT the built commit with no `git describe` fallback. An audit
flagged that as missing stage's two-tier resolution, but the code documents it as
intentional - a prod image must carry a real release tag or none, never
"v111.32.3-4-gbb20466". The nine deploy/release pairs have no drift at all.

Verified: all 12 files parse under js-yaml; every pair now has identical job and
step counts (100/100, 110/110, 98/98); all 7 feature markers match across all six
pairs; no bare `CSC_LINK:` remains anywhere; zero CRLF; exactly 12 files touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 16:23:20 +02:00
Ruslan KonviserandClaude Opus 5 f2aeadc163 fix(ci): point release-linux-arm64 at a runner that exists
`release-linux-arm64` is the only app job whose matrix hardcodes a runner with no
`vars.` fallback: `os: [ubicloud-standard-8-arm]`. Every sibling job uses the
`${{ vars.X || 'default' }}` form. ever-co's ubicloud ARM capacity is gone, so the
job never gets a runner - all six app workflows on the 2026-09-22 `apps` promotion
sat queued from 23:12Z with 0/6 arm64 builds.

`timeout-minutes: 300` does NOT rescue this: that clock only starts once a job gets
a runner. An unstarted job waits until GitHub cancels it at ~24h, so the RUNS never
reach a terminal state either, and the run-level conclusion becomes meaningless.

Now `${{ vars.RUNNER_LINUX_ARM64 || 'ubuntu-24.04-arm' }}`, matching the shape of
the x64 line two jobs above it. The variable is not currently set at repo or org
level, so this resolves to GitHub's hosted `ubuntu-24.04-arm` - free for public
repositories, and ever-gauzy is public. Setting RUNNER_LINUX_ARM64 later moves every
one of these jobs to a self-hosted pool without another code change.

Applied to all six *-prod.yml and all six *-stage.yml so stage-apps cannot hit the
same wall. The job is repointed, not removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 05:08:15 +02:00
Ruslan Konviser f96b5d4ba2 ci(codeql): a pull request into develop no longer starts an analysis
The owner's instruction was that a pull request into `develop` must not start a workflow — the cost
belonged to every commit on the branch being merged, and the run belongs to the merge. Every other
workflow that fired on such a pull request was changed on `feat/platform-extensions`; this one could
not be, because `codeql.yml` exists only on `develop`.

Measured before the change, over the previous twelve CodeQL runs: **178 run-minutes in total, 137 of
them on pull requests**, at two jobs of 22–45 minutes each on the 4-vCPU ARC pool. It is paid on every
contributor's branch into develop, not on one: `feat/platform-extensions` (44.5 and 29.0 minutes),
`develop` (22.2), `fix/custom-dashboard-drag-and-drop`, `feat/9873-restrict-agent-exit-logout`.

What stays: the `push` trigger still scans `develop`, `stage` and `master`, so code that lands on any
of them is analysed — the run moves to the merge rather than disappearing. The weekly `schedule` scan
and `workflow_dispatch` are untouched, the `concurrency` block is untouched, and pull requests aimed
at `stage` and `master` — the release cascade — still run it.

`branches-ignore` rather than a `branches` list, because a list is what has to be edited when a new
protected branch appears and this is the exception to it. `build.yml` and `secrets-analysis.yml`
already spell their develop exception this way. The block carries the reversed spelling in a comment,
so restoring the trigger is one paste.

The trade-off, stated plainly: a pull request into `develop` is no longer scanned *before* merge. The
security signal moves to the push run on develop, which is after merge. If pre-merge scanning on
develop is wanted more than the runner time, this change should be closed rather than merged — the
alternative that keeps both is to leave the trigger and accept roughly one 30-minute run per push.
2026-09-22 17:51:42 +02:00
Ruslan KonviserandClaude Opus 5 58bf570f64 ci: stop concurrent CodeQL and duplicate gate runs on PRs and develop/stage (#10259)
* ci: stop concurrent CodeQL and duplicate gate runs on PRs and develop/stage

CodeQL ran through GitHub's code-scanning DEFAULT SETUP, which generates its
workflow server-side (run path `dynamic/github-code-scanning/codeql`). There is
no file to put a `concurrency:` block in, and GitHub does not supersede its own
default-setup runs, so every push to a pull request started another full
analysis alongside the ones already in flight. Measured on #10254 on
2026-09-21: runs 35547460875 and 35548451790 overlapped by 12 minutes, and
35580875040 / 35582701347 / 35582925033 were all running at 09:22. Each run is
two jobs of 25-35 minutes on the 4-vCPU ARC pool.

Converts CodeQL to advanced setup with a real workflow file that reproduces the
default-setup configuration exactly - the two analyses it actually ran
(javascript-typescript, actions; the other three listed languages are aliases),
the default query suite, the default remote threat model, the weekly schedule,
and the ever-k8s-linux-x64-4 pool - and adds the cancel-in-progress policy that
default setup could not have. Branch coverage stays at develop/stage/master
because all three are protected and default setup was scanning all three.

Also closes the remaining concurrency gaps on the PR and develop/stage surface:

- build.yml and secrets-analysis.yml both pair an unfiltered `pull_request:`
  with a `push:` on the same branches, so while a release-cascade PR is open
  (today #10257, stage -> stage-apps) one push fires two runs of the same tree
  in two groups that could never cancel each other. Keying the group on the
  head ref collapses them; qualifying it by head repository keeps a fork branch
  of the same name from claiming the group and cancelling a real build before
  its job skips.
- harvest-secrets-to-openbao.yml was the only workflow of 70 with no
  concurrency block, and two overlapping dispatches can interleave into a torn
  OpenBao snapshot.
- external-uptime-monitor.yml now cancels superseded PR self-tests while still
  never cancelling a live probe, which owns the alert issue.

Deliberately unchanged: the deploy and release workflows keep
`cancel-in-progress: false`, and test-unit.yml keeps it on the stage gate. Each
documents an incident where cancelling caused real harm.

Requires default setup to be disabled first - GitHub rejects CodeQL SARIF
uploads from a workflow while default setup is enabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: collapse duplicate CodeQL runs and teach cspell the OpenBao name

Two review follow-ups on this branch.

Greptile P2: codeql.yml had the same duplicate-run defect this PR fixes in
build.yml and secrets-analysis.yml. Both of its triggers cover
develop/stage/master, so while a cascade pull request is open (develop ->
stage), one push to develop fires a push run on refs/heads/develop and a
pull_request run on refs/pull/N/merge. Keyed on github.ref those are different
strings, so cancel-in-progress could never collapse them and the same tree was
analysed twice. Now keyed on the head ref and qualified by head repository, the
same shape used in the other two files.

Cspell: this PR is the first change to harvest-secrets-to-openbao.yml since the
spellcheck was added, and the action only scans changed files - so `openbao` had
never been checked before and surfaced as an unknown word. Added to .cspell.json
next to the other product names.

Not taken: CodeRabbit asked for the actions in codeql.yml to be pinned to commit
SHAs. Declined as inconsistent rather than wrong - this repository tag-pins 547
`uses:` references against a handful of SHA pins, and there is no dependabot
config to keep SHAs current, so pinning one new file would leave an outlier that
nothing updates. Worth doing repo-wide as a deliberate policy, not here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: use US spelling in the codeql concurrency comment

Cspell flagged `analysed` at codeql.yml:60. This repository is US English
throughout (the job is named `Analyze`), so the comment is corrected rather
than the word added to the dictionary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 11:45:19 +02:00
Ruslan KonviserandClaude Opus 5 70b1a85606 fix(security): harden the per-account login control and the Redis throttler after review
This round addresses the review findings on the brute-force and seed-credential fix.

LoginAttemptService now works in three steps: begin(), then fail(), succeed() or release(). The
counters live in EVER_REDIS_CLIENT, and each step runs as one MULTI. A 250 ms deadline sends a
slow step to an in-process store instead, and a Redis error or a disconnected client does the
same, so the control never fails open. Each account may have at most AUTH_MAX_FAILED_ATTEMPTS
checks in flight at once. A hard block needs that many failures from at least two distinct client
sources, using the same spoof-resistant tracker as the route throttler, so one client can no longer
lock a known email out. A failure that cannot be traced to an address counts against the account
(fail closed). A failure that was already in flight when a block landed does not extend it. The
new 429 carries a Retry-After header. Only a credential verdict counts as a failure: an
infrastructure error in login, workspace sign-in or magic-code sign-in releases the slot, and so
does a team-join request that carried no code at all.

RedisThrottlerStorage reads the block marker, the increment and the TTL in one MULTI snapshot. It
sets the block with NX and never asks for a PX of 0. It sends nothing to a client that is not
ready, so node-redis cannot replay those commands after the request has fallen back.

The tracker now expands IPv6 addresses to their canonical form. IPv4-mapped addresses, in hex or
dotted form, collapse to IPv4, and every other address is bucketed by its /64. The docs plugin
tracker no longer uses the client-supplied tenant-id header. parseNonNegativeInt rejects a value
that only starts with digits.

The seed guard now reads environment.demoCredentialConfig, which is what the seeder hashes,
before it reads process.env. It refuses the ever and all seeds in production because they create
the hard-coded fixture accounts. configure.ts writes a seed password into the public web bundle
only for a demo build or when the password is a published default. The Cloudflare compose
templates set THROTTLE_TRUST_CF_CONNECTING_IP. The Playwright job pins its seed passwords. The
env and README docs describe the multi-source rule and the AUTH_MAX_FAILED_ATTEMPTS=0 escape
hatch.

The verification pass renamed identifiers that cspell rejected (IPV6_TRACKER_PREFIX_HEXTETS to
IPV6_TRACKER_PREFIX_GROUPS, and zset to sortedSet in the Redis fake) and reworded one comment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 05:50:09 +02:00
Ruslan KonviserandClaude Opus 5 5ada11e5ce ci: never run workflow jobs for pull requests from forks
Owner decision (2026-09-16): fork pull requests must not run on our CI.
Every job in the seven pull_request-triggered workflows (build,
static-checks, typos, snyk-analysis, secrets-analysis, mega-linter,
external-uptime-monitor) now skips when the pull request's head repository
is not ever-co/ever-gauzy. Pushes and pull requests from branches of this
repository run exactly as before. Jobs whose if used !cancelled() or
always() get the guard inside the same expression, so they do not run
after the first job is skipped; build-desktop already runs only on
develop/stage/master pushes.

This is one of three layers: the repository already requires approval for
all external contributors, and the ARC runner image now refuses fork pull
request jobs in a job-started hook (ever-co/arc-runner-image), which a fork
cannot edit away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 20:07:08 +02:00
Ruslan Konviser 625d621f19 feat(integrations): complete organization-scoped Ever Async connector and builds 2026-09-07 12:58:55 +02:00
Ruslan Konviser cb4e276e04 Merge pull request #10124 from ever-co/codex/fix-stage-uuid-migration
fix(database): unblock stage tenant migration on PostgreSQL UUIDs
2026-09-07 11:59:41 +02:00
Ruslan KonviserandClaude Opus 5 be59bce929 CI: add a jest-config typecheck gate and a non-blocking lint report (#10119)
Nothing in CI ran `tsc --noEmit` or ESLint. That is why a duplicated `transformIgnorePatterns`
key sat in `packages/core/jest.config.ts` undetected (TS1117, fixed in #10116), and why ESLint
had been broken repo-wide long enough for 93 of 93 projects to fail (fixed in #10117). Both
gaps are now covered, deliberately with very different postures.

`typecheck-configs` is a real gate. `tools/tsconfig.jest-configs.json` type-checks all 95
jest.config.ts files and is green today, so the job fails if that stops being true. It is built
to need nothing but the compiler — no `extends`, `types: []`, `noResolve`, and a small shim in
`tools/jest-config-globals.d.ts` for the 53 CommonJS configs plus the one `@nx/jest` import in
the root config. The compile is 0.5s across all 95 files; the whole job measured 68s in CI,
dominated by runner provisioning. A cold `yarn install` on this fleet is measured in hours, so
avoiding one is the entire design constraint.

Proven, not assumed: exit 0 on the current tree, and exit 2 with TS1117 when a duplicate key is
reintroduced.

`lint` is NOT a gate. It runs `nx run-many -t lint` and writes the outcome to the run summary.
The repair in #10117 left roughly 2,650 errors and 6,100 warnings, all predating the job, so a
blocking lint gate would wedge every PR on day one. Non-blocking at both the step level (so the
summary still runs) and the job level (so a cache miss or install failure on the fleet cannot
redden the check either). Verified end to end on this PR: ESLint exited 1 with real findings,
the summary step ran, and the check reported green.

TypeScript is installed pinned with `--ignore-scripts` rather than fetched through `npx --yes`,
which SonarCloud correctly flagged as a vulnerability twice (S6505/S8543) for running a
network-fetched package's lifecycle scripts in CI. It is also faster.

Neither job is wired into branch protection; `develop` has no required status checks at all, so
these are visible signals rather than hard gates. The tracking task for burning down the ESLint
backlog lives in the agent workspace at `knowledge/TASKS.md`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 16:44:16 +02:00
Ruslan Konviser fe004fe999 fix(downloads): run the R2 mirror on a GitHub-hosted runner
The self-hosted ARC runners cannot reach Cloudflare R2. Measured from inside a
running ever-k8s-linux-x64-4 pod: the R2 S3 endpoint and downloads.ever.co both
time out, while api.github.com, registry-1.docker.io and cloudflare.com are all
reachable. Forcing IPv4 does not help, so this is not the AAAA-first problem but
the ISP-level blackhole of certain Cloudflare anycast prefixes this fleet has hit
before -- the same fault that forced cdn.ever.co in-cluster.

That made the first backfill run fail its signing self-test with "fetch failed",
which is exactly what the self-test is for: it stopped before publishing an empty
manifest that would have silently reverted every download link to GitHub.

A narrow, documented exception to self-hosted-only: this job exists to talk to
Cloudflare, which a self-hosted runner structurally cannot do. ever-gauzy is
public, so GitHub-hosted minutes are free.
2026-08-30 21:00:34 +02:00
Ruslan Konviser f609997d04 feat(downloads): mirror desktop installers to R2 behind downloads.ever.co
Adds .github/workflows/mirror-releases-to-r2.yml.
2026-08-30 19:10:35 +02:00
Ruslan KonviserandClaude Opus 5 5aa7b27392 ci: use the real nx flag and drop the checkout credential
From review on #10055.

`--continue` is not an nx option. nx 22 spells it `nxBail`, and an unknown
option is forwarded to the executor - so `--continue` would have reached jest
rather than enabling failure continuation, and the first failing project would
still have hidden every other project's result. Now `--nxBail=false`.

Also sets `persist-credentials: false` on checkout: this job only reads the
tree and never pushes, so the credential does not need to outlive the step.

Action majors stay aligned with build.yml on purpose - this job consumes the
node_modules cache that workflow writes, so producer and consumer are kept on
the same actions rather than following the newer majors used elsewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 00:30:12 +02:00
Ruslan KonviserandClaude Opus 5 9b3ada96ea ci: run the existing jest unit tests on stage
The repository has ~330 `*.spec.ts` files and an nx `test` target on 20+
projects (`@nx/jest:jest`, already cached via nx.json targetDefaults), but no
workflow ever invoked any of them. CI ran builds plus Playwright/Cypress e2e
only, so a green pull request proved the code COMPILED and said nothing about
the unit tests - including the ~53 specs in packages/core.

This wires the existing targets up. It adds no test tooling and changes no
test code.

Scope is `stage` only, plus manual dispatch. Stage is where the cascade lands
before production and pushes to it are infrequent, so this cannot slow the
day-to-day loop while the suites' true state is still unknown.

Deliberately NOT a required status check: these suites have never run in CI so
the first runs may be red, and that is exactly the information the workflow
exists to surface - it should surface it without blocking the cascade. Promote
it to required once it has been green for a while.

`cancel-in-progress: false`, because a cancelled test run produces no verdict
and a gate that reports nothing is worse than no gate.

Dependencies come from the same lockfile-keyed cache the build workflow
writes, so a stage push that already built this tree reuses it rather than
paying for another install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 00:14:25 +02:00
Ruslan Konviser b6287c86b7 fix(build): use bundled Node headers for native modules 2026-08-23 23:04:54 +02:00
Ruslan KonviserandClaude Opus 5 a6ab2fc9e3 ci(verdaccio): route the Cypress and Currents e2e installs through the internal npm cache
These were the last two dependency-installing workflows still resolving packages
straight from the public npm registry. Both are covered now, using the same
placement discipline as build.yml.

The step goes AFTER the yarn cache restore and BEFORE `yarn bootstrap`, because
the action rewrites yarn.lock's resolved URLs (the load-bearing half - yarn v1
fetches the URL recorded in the lockfile and ignores the registry setting) while
hashFiles() in a cache key evaluates at step runtime. Configuring the registry
first would move the key and bifurcate the cache namespace by VIP reachability.

No save-key pinning was required. build.yml needs it because it uses the split
actions/cache/restore + actions/cache/save pair, whose save step re-derives
hashFiles() at save time. Both files here use the COMBINED actions/cache, which
records its primary key with core.saveState during the restore - before the
rewrite - and whose post-job save reads that recorded key back, so the entry is
written under exactly the key the next run's restore computes.

Scope, and what was deliberately left alone:

- test_currents.yml `e2e-tests` and test_cypress.yml `e2e-tests` get the step.
  Both resolve to the self-hosted ARC pool via vars.RUNNER_LINUX_X64_8
  (= ever-k8s-linux-x64-8), the only place the VIP answers. The action probes
  the VIP and falls back to the public registry, so the `|| 'ubuntu-latest'`
  fallback runner degrades to today's behaviour rather than hanging on a LAN IP.
- test_currents.yml `prepare` installs nothing - it only emits a build UUID.
- test_cypress.yml `e2e-tests-setup` overrides its install with
  `install-command: node -p 'os.cpus()'`, so it resolves no packages from any
  npm registry. Adding the step there would buy nothing and would rewrite
  yarn.lock ahead of cypress-io/github-action's own lockfile-keyed cache,
  splitting it in two.

`expect-vip` is derived from the runner variable rather than the
contains(matrix.os, ...) idiom the app workflows use: these jobs' only matrix
key is `containers`, so that expression would collapse to false and silently
disable the in-network VIP retry and its warning.

Additive only - no workflow, job or step was removed or reordered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 21:50:01 +02:00
Ruslan KonviserandClaude Fable 5 a81dbc073a perf(ci): route the build gate's installs through Verdaccio - with the key trap handled
build.yml was the one workflow deliberately skipped by #10025, because the
configure-registry action rewrites yarn.lock and this file keys its dependency
cache on hashFiles('yarn.lock', ...) - a naive placement moves the key and
bifurcates 2.9 GB cache entries by VIP reachability.

The deferral had a measured cost tonight: on stage run 32513912629, cleanup had
deleted the run's artifact, the cache entry was LRU-evicted, and the bootstrap
fallback pulled the whole tree from the PUBLIC registry - blowing through the
180-minute job ceiling twice (attempts 2 and 3 of build-api).

Both traps are handled explicitly:
  * Configure Registry sits AFTER every lockfile-keyed cache step - producer:
    after the restore, gated on cache-miss; consumers: after the artifact
    download and cache fallback, so only the bootstrap path pays for it.
  * The producer's save key is pinned to steps.cache.outputs.cache-primary-key
    (computed before the rewrite) instead of re-deriving hashFiles at save time,
    which would write the archive under a different key than the next run
    looks up.
  * expect-vip derives from vars.RUNNER_LINUX_X64_8 - these jobs define no
    matrix.os, so the fleet-standard contains(matrix.os, ...) is always false
    and would silently disable the VIP retry.

Pinned to aa4ee199, the action revision that survives pnpm/npm repos and treats
only git exit 1 as "untracked".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 07:21:45 +02:00
Ruslan KonviserandClaude Fable 5 2a58ca4609 fix(ci): keep the node_modules artifact when the run is red so reruns can restore it
The cleanup job deleted this run's ONLY dependency handoff unconditionally
(`if: always()`). That choice had a real rationale - red-run leftovers once
filled the org's Actions storage quota and stopped artifact uploads - but it
breaks "Re-run failed jobs": with the artifact gone, every rerun consumer falls
into the documented 1-3h bootstrap fallback.

Measured on stage run 32513912629 (2026-08-21): a seconds-long DNS blip
(getaddrinfo EAI_AGAIN on GitHub's artifact endpoints) failed build-web's
restore; cleanup deleted the artifact when the attempt ended; the rerun then
cost build-libs 178 minutes and pushed build-api into the 180-minute ceiling.
Hours of runner time to recover from a transient network error.

Deletion now happens only when no needed job failed or was cancelled. The quota
concern is bounded rather than ignored: red Build runs are rare since the gate
rework, and retention-days: 1 ages any kept artifact out within a day. skipped
results (build-desktop on PR runs) still allow deletion.

This was the greptile/cubic P1 on #10010, catalogued as "deferred - needs a
design call". Production made the call.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 04:19:22 +02:00
Ruslan KonviserandClaude Opus 5 fb6673c797 perf(ci): serve the Playwright deps install from Verdaccio, not the internet
`test_playwright.yml` never pointed at our internal npm cache, so the heaviest
install in the fleet pulled ~4.2 GB from registry.yarnpkg.com on every run. The
job documents itself at 88 min, 3h20m and 3h46m.

An audit of all 11 orgs found 177 workflows across 56 repos in the same state -
this is the highest-traffic one (222 runs) and the cleanest to fix.

Reuses the existing composite action rather than hand-rolling a snippet, which
matters more than it looks:

  * yarn v1 fetches the URL recorded in yarn.lock and IGNORES the registry
    setting. Measured with a negative control - a blackholed `resolved` URL with
    .yarnrc pointing at a WORKING registry still failed with ENOTFOUND. A
    config-only patch here would be a no-op that looks like a fix: green build,
    same 83 minutes. The action rewrites the lockfile URLs, which is the half
    that actually moves the bytes.
  * Verified the rewrite does NOT break --frozen-lockfile: yarn compares
    package.json patterns against lockfile entries, and the resolved host is not
    part of that comparison. Measured against the real VIP, RC=0, lockfile
    unmodified.
  * The action probes the VIP and falls back to the public registry. That is
    mandatory, not decorative: the VIP is a single MetalLB L2 announcement with a
    recurring speaker flap, so hardcoding it would convert a cache optimisation
    into a CI outage vector.

Two deliberate deviations from how the app workflows call this action:

  * `expect-vip` derives from `vars.RUNNER_LINUX_X64_8 != ''`, NOT
    `contains(matrix.os, ...)`. This workflow defines no `matrix.os` - its only
    matrix key is `shard` - so the standard expression evaluates false and
    silently disables the in-network retry and warning. A mute switch, not an
    error.
  * The step sits immediately after checkout, which is safe HERE because this
    workflow deliberately has no `actions/cache`. In build.yml the same placement
    would move the `hashFiles('yarn.lock')` cache key, so that file needs
    different handling and is not touched here.

The `cleanup` job stays on ubuntu-latest and deliberately does NOT get this step:
a GitHub-hosted runner cannot reach 192.168.1.183 and would hang.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 02:05:21 +02:00
Ruslan KonviserandClaude Opus 5 7118ab32fa fix(ci): stop every consumer downloading the dependency archive twice
The cache-fallback step carried a comment saying "Only reached if the artifact
was unavailable" - but it had no `if:` and the download step had no `id:`, so it
ran unconditionally. The producer saves that same cache key on every cold run, so
the lookup HITS and actions/cache/restore pulls the tree a second time.

Measured live on run 32418672320 (develop, the first run of this design):

  build-api       Download node_modules archive   1015s   success
                  Restore ... (cache fallback)     273s   success
  build-desktop   Download node_modules archive   1028s   success
  build-web       Download node_modules archive   1102s   success

A cache MISS returns in seconds; 273s is a download. That is ~4.5 minutes of
redundant transfer per consumer, ~18 minutes per run, on the same WAN link the
install already saturates - and it scales with every consumer job added.

Gated on `steps.artifact.outcome != 'success'`. `outcome` and not `conclusion`
on purpose: `continue-on-error: true` masks `conclusion` to success, so only
`outcome` reports a genuine download failure - which is exactly when the cache
fallback should run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 01:26:37 +02:00
Ruslan KonviserandClaude Opus 5 54ed5ebdb3 fix(ci): the install-once producer breaks on every cache HIT
Two correctness bugs from the AI review of #10010, both confirmed against the
merged code on develop.

1. build.yml - every warm run would fail.

   `Restore node_modules archive` used `lookup-only: true`, so on a cache HIT the
   archive is never written to disk. `Install`, `Run postinstall`, `Archive` and
   `Save` are all gated on `cache-hit != 'true'` and correctly skip - but
   `Upload node_modules archive` carries NO `if:` guard and sets
   `if-no-files-found: error`. On a hit it therefore uploads nothing and fails
   the job, and because the other four jobs `needs:` this one, it takes the whole
   gate with it.

   The cold run that first writes the key passes, so this is invisible until the
   SECOND run - the one the cache exists for. Dropping `lookup-only` makes the
   producer always materialise the archive: a ~2.9 GB download costs a few
   minutes against the ~90-minute install it replaces, and it is what guarantees
   the run-scoped artifact handoff actually has a file to hand off.

2. archive-node-modules.sh - a failed tar could publish a corrupt archive.

   The script set only `-o pipefail`, never `-e`, and tolerated tar exit 1 with
   `|| [ $? -eq 1 ]`. A tar exit status of 2 - a real error - left the `||`
   compound false and execution simply continued to the size check. A ~2.9 GB
   tree truncated early still clears the 100 MB floor, so the partial archive
   would be published and unpacked by four consumer jobs.

   Now `set -e`, with the tar status captured explicitly so exit 1 stays
   tolerated and anything else aborts before the archive can be published.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:24:53 +02:00
Ruslan Konviser 07784b0c9d Merge pull request #10010 from ever-co/fix/expense-null-employee-scope
fix(expenses): org-level expenses 400 since the null-where hardening; stop the CI artifact leak
2026-08-20 23:17:44 +02:00
Ruslan KonviserandClaude Opus 5 3c6e34dc00 chore(spell): American spelling in the build workflow comment
cspell is the only thing standing between this PR and merge: it fails on
"optimisation" at build.yml:113. The repo enforces the American spelling and has
corrected this exact word in this exact file before (6c78c4e92e).

No behaviour change - a comment word.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:16:00 +02:00
Ruslan KonviserandClaude Opus 5 bcdbc1ea26 fix(ci): hand the tree over by ARTIFACT — the cache was evicted mid-run
First run of the install-once design failed, and the logs are unambiguous:

  build-monorepo-root: "Cache saved with key: Linux-X64-node-modules-…2b58b651…"  (passed, 1h44m)
  build-libs (32s):    "Failed to restore cache entry. Input key: …2b58b651…"

Byte-identical keys. The producer saved it and the consumer could not read it,
because an Actions cache is a REPOSITORY resource on a 10 GB budget, not a
run-scoped one. While this run was in flight, build.yml on develop, on stage and
on PR #10014 — all still carrying `cache: yarn`, since this change exists only on
this branch — each wrote a 4.66 GB setup-node yarn cache. The repo hit 14.03 GB
and LRU took the newest entry: the tree archive, seconds after it was written.

So the cache cannot be the handoff. It is now best-effort and CROSS-run only.
The handoff is an artifact, which is scoped to the run and cannot be evicted by
another workflow. Consumers try the artifact, fall back to the cache, and only
then to a full install — and `fail-on-cache-miss` is gone, since a cache miss is
no longer the fault it was.

A cleanup job deletes the artifact when the run ends, so it costs storage only
while the run needs it — the lesson from the e2e workflow, where an artifact
left behind on every run filled the org's quota and stopped reports uploading.

Also pruned the three duplicate yarn caches by hand (9.34 GB): they are the same
key on three refs, and this change stops build.yml producing them at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 21:58:11 +02:00
Ruslan KonviserandClaude Fable 5 baa3c0eaaa ci: never cancel a release run — a cancelled release publishes no image
#10015 gave release-{demo,stage,prod} cancel-in-progress: true. That is the one
place where superseding is harmful:

  Release Demo  --workflow_run: [completed]-->  Build and Publish Docker Images Demo
                                                if: workflow_run.conclusion == 'success'

A cancelled release still fires the workflow_run event, but with conclusion
'cancelled', so its image build SKIPS. When that happens to the LAST release of a
burst, no image is published for the new code at all and the environment keeps
serving the previous image until the next push — demo sat on the 03:22 image all
day on 2026-08-20 for exactly this reason (compounded by hand-cancelling the
superseded image builds).

Releases take about a minute, so superseding them saved nothing. They are now
serialized per ref but never cancelled; the expensive jobs (build.yml, the image
builds) keep superseding, so a burst of merges still yields ONE compile and ONE
published image — the newest. The atomic tag+release from #10015 stays: it is
what makes an interrupted release harmless if one ever is.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 16:15:33 +02:00
Ruslan KonviserandClaude Opus 5 06569082de fix(ci): close the review findings on the install-once restructure
Adversarial review of d7c56524d2 found one blocker I had walked straight into,
plus several things that would have made the gate quietly worse.

BLOCKER — `needs:` without `if:` erases four verdicts. A job whose dependency
fails is recorded `skipped`, not `failure`, and a skipped required check does
NOT block a merge. So any producer hiccup would have silently removed the
build-libs / build-api / build-web / build-desktop verdicts from the gate — the
exact "no failure just means never ran" hazard this file's own header warns
about, and the shape of the 2026-08-13 template-error incident it cites. All
four now carry `if: !cancelled()`: the ordering still holds, but a dead producer
yields a real red instead of four silent skips. build-desktop folds it into its
existing branch filter, because two `if:` keys is a duplicate mapping key and
the file will not parse.

Moving the archive shell into a script silently dropped errexit — an inline
`run:` gets `bash -eo pipefail` from GitHub, a script with its own shebang does
not. `tar ... || [ $? -eq 1 ]` therefore stopped aborting on a real tar failure,
so a truncated archive could be published under a good key. Cache entries are
immutable per key, so that poison would have persisted until the lockfile
changed, with every consumer falling back to a multi-hour install. tar's status
is now checked explicitly, and the size floor moves from 100 MB (~3% of reality)
to 1 GB.

The fallback called `yarn bootstrap`, which drops `--frozen-lockfile` and
`--network-timeout` — so the degraded path could build a tree that does not match
the committed lockfile and report green, while the Docker builds that DO install
frozen fail later on develop. Added `bootstrap:ci` carrying the CI flags, used by
both workflows' fallbacks so they cannot drift apart again.

Also from the review:
- `package.json` joins the cache key. Without it a PR that edits package.json
  without regenerating yarn.lock hashed identically to develop, hit the cache,
  and skipped the only `--frozen-lockfile` execution left in the workflow.
- `lookup-only: true` on the producer: on a hit it did nothing with the archive
  (every step is gated on the miss), so it was downloading ~2.9 GB purely to
  compute a boolean, in front of four jobs waiting on it.
- `fail-on-cache-miss: true` on the consumers: the producer's success means the
  entry exists, so a miss is an eviction, not a normal path. Fail in seconds
  rather than letting four jobs re-create the contention this change removes.
- the producer moves to the 8-core pool and 360 minutes. It is now the serial
  critical path, its work is CPU-bound, and a cold install here has no yarn
  tarball cache behind it — the equivalent deps job measured 88 min, 3h20m and
  3h46m. At 180 a cold miss would time out and take the whole gate with it.
- the header comment said the `needs:` edges were gone; it now describes why
  they are back and what they cost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 15:58:34 +02:00
Ruslan KonviserandClaude Opus 5 d7c56524d2 perf(ci): build.yml installs the monorepo ONCE instead of five times in parallel
Five jobs each ran their own full `yarn install` of the same tree, concurrently,
on the same self-hosted pool — so they competed with each other for I/O. Same
commit, same run, measured on PR #10010 (run 32358690732):

  build-api             install 1h49m  -> passed
  build-monorepo-root   install 1h51m  -> passed
  build-libs            install 2h40m  -> KILLED at the 180-minute ceiling
  build-web             install 2h40m  -> KILLED at the 180-minute ceiling

The two that died never compiled a line: `Build packages` and `Build web` are
recorded as SKIPPED. Which two survive is a coin flip, and this gate has failed
that way four times during this work on code that builds fine.

`build-monorepo-root` already did install + postinstall and nothing else, so it
becomes the producer: it publishes the tree as one compressed archive, and the
other four `needs:` it and unpack in minutes instead of installing.

A CACHE, not an artifact — the opposite choice to test_playwright.yml, for two
reasons that only apply here. build.yml runs on every pull_request, so a ~2.9 GB
artifact per run would meter against the org's Actions storage (the quota that
just stopped Playwright reports uploading); caches are unbilled and live in a
separate 10 GB per-repo budget. And a cache persists ACROSS runs keyed on the
lockfile, so once develop populates it every PR restores rather than installing
— which an artifact, scoped to its own run, cannot do.

Net cache usage goes DOWN: `cache: yarn` is removed from all five jobs, which
frees the 4.66 GB setup-node yarn cache it kept per ref, and the ~2.9 GB tree
archive replaces it. The archive is only re-saved when yarn.lock, patches/ or
.scripts/postinstall.js change, so most PRs restore and never write.

The archive/restore shell is now shared with test_playwright.yml as
.github/scripts/archive-node-modules.sh and restore-node-modules.sh — one copy
behind eight call sites, rather than the drift that already bit the e2e workflow
once.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 15:42:18 +02:00
Ruslan KonviserandClaude Fable 5 0caa9ae72c ci: only the newest commit builds — supersede policy across all 67 workflows
Four merges to develop this morning left five superseded runs building dead
commits while the branch head waited. Applied one policy, per class:

- build.yml: cancel-in-progress was `github.event_name == 'pull_request'`, so
  branch pushes never superseded each other — every merge started its own 3-h
  compile. Now true for all events; it writes nothing outside the run.
- release-{demo,stage,prod}: made the release ATOMIC so it becomes safe to
  supersede. `github-tag-action` now runs with dry_run (calculate only) and
  `ncipollo/release-action` creates the tag AND the release in one call with
  `commit: github.sha`. Previously the tag was pushed by an earlier step, so a
  cancelled run stranded a tag with no release forever (the next run computes the
  NEXT version and never backfills). With that window gone: cancel-in-progress true.
- 6 non-prod image builds: now supersede within their own workflow+ref, matching
  what the prod image builds already did. Group stays per-workflow — the shared
  `gauzy-docker-build` group is what starved them before. Trade-off documented in
  the file: an intermediate version tag may end up without an image, which is inert
  because demo/stage deploy from :latest.
- 26 deploy workflows had NO concurrency key at all, so every merge queued its own
  deploy. They now share a group per workflow+ref with cancel-in-progress FALSE:
  superseded PENDING runs collapse to the newest, but a rollout already applying is
  never interrupted.

Left alone deliberately: cache-janitor and external-uptime-monitor (cron probes,
must not cancel each other) and harvest-secrets-to-openbao (manual only).

Every file parses as YAML; the classification came from reading each workflow's
steps for external side effects, not from its name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 14:24:27 +02:00
Ruslan KonviserandClaude Opus 5 e78ac94a0d ci(e2e): gate on stage only, and keep reports for 1 day
Two changes, both narrowing what this workflow costs.

TRIGGER: push to `stage`, not `develop`. The suite takes hours and is shared-DB
flaky, and running it on every develop push spent that on code that was not yet
a release candidate — while still not gating anything, since a red develop run
does not stop the cascade. Contributors run it locally before pushing; this is
the last gate before production, so that is where it belongs. `pull_request` was
already excluded and stays excluded. `workflow_dispatch` remains, so any branch
can still be checked on demand — which is how this branch was verified.

REPORT RETENTION: 14 -> 1 day, matching every other artifact here. Four reports
per run shared the org's Actions storage with the ~2.87 GB dependency archive;
at 14 days they had grown to 3.72 GB across 100 copies and tipped the quota,
which stops NEW reports uploading — the accumulated diagnostics evict the
diagnostic you actually need, exactly when a run has gone red. Triage downloads
the report during the run's life and a re-run regenerates it, so a week of
history was buying nothing.

Together with the cleanup job these compound: far fewer runs, ~2.87 GB freed at
the end of each, and a standing report cost measured in hundreds of MB instead
of gigabytes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:23:06 +02:00
Ruslan KonviserandClaude Opus 5 889b1b5894 fix(ci): halve report retention so the reports stop evicting each other
The cleanup job bounds the per-run dependency archive, but the standing cost is
the reports: 4 shards x every run x 14 days had grown to 3.72 GB across 100
copies. Added to a ~2.87 GB archive in flight, that is what tipped the org's
Actions storage over — and the failure mode is perverse, because hitting the
quota stops NEW reports uploading. The accumulated diagnostics evict the
diagnostic you actually need, exactly when a run has gone red.

7 days is still far longer than any triage in this work has taken (hours), and
it roughly halves the standing figure. Peak becomes ~2.87 GB in flight plus
~1.8 GB standing, instead of ~2.87 + 3.72.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:11:07 +02:00
Ruslan KonviserandClaude Opus 5 6a3b2982df fix(ci): scope the token per job — only cleanup needs actions: write
SonarCloud flagged both halves of the workflow-level permissions block
(S8233 write, S8264 read), and it is right: putting `actions: write` at the
workflow level handed artifact-delete rights to the deps, build and shard jobs,
none of which need them. Only `cleanup` deletes anything.

Each job now declares its own: `contents: read` everywhere, plus
`actions: write` on `cleanup` alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:17:05 +02:00
Ruslan KonviserandClaude Opus 5 cf1a3cdd42 fix(ci): delete the node_modules artifact when the run is done
"Artifact storage quota has been hit. Unable to upload any new artifacts."
stopped the Playwright REPORTS uploading on all four shards of develop run
32285142151 — including the two whose tests fully passed, which is why that run
showed red jobs next to "21 passed". Worse, it blinded the diagnosis: a red run
with no report is a red run you cannot debug.

The cause is mine. Moving the dependency tree from the Actions CACHE to an
ARTIFACT fixed the eviction that was costing entire runs, but the two budgets
are not alike: self-hosted runners provide COMPUTE, not storage, so
upload-artifact still ships to GitHub's own storage and bills the org's quota
regardless of where the job ran. Caches are a separate, unbilled 10 GB
per-repo budget. At ~2.87 GB per run this filled it quickly — measured at 6.87 GB
live across 1,089 artifacts, of which that single object was 42%.

Retention is already the 1-day minimum, so the only real fix is deleting it once
the last consumer is finished. The job runs on `always()` because the tree is
just as useless after a failed shard, and leaving it behind on red runs is how
the quota filled to begin with.

Storage peaks at one run's tree instead of accumulating. The reports themselves
(14-day retention, ~3.7 GB across 100 copies) are untouched — they are the thing
worth keeping, and they now have room.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 10:49:09 +02:00
Ruslan KonviserandClaude Opus 5 b6268ead7e fix(ci,docs-ui): SonarCloud gate — a11y bug, optional chain, and my own duplication
The build gate went green but SonarCloud failed on two conditions.

new_reliability_rating 2 (needs 1). Adding `(click)` to the table container gave
the grid a pointer-only way to open a row: the rows are plain `<tr>`s with no
interactive role, so there was no keyboard path to the detail panel at all.
`onRowKeydown` opens the focused row on Enter/Space, reusing the same
control-guard so the kebab and Retry keep their own keys, and preventing Space
from scrolling the table out from under the panel it just opened. The two
helpers widen from MouseEvent to Event since they only ever read `event.target`.

new_duplicated_lines_density 4.8% (needs <= 3%), and it was mine: the unpack +
fallback blocks were copied verbatim into the build job and the shard job. They
had ALREADY drifted once — a fix landed in one copy and missed the other, which
is why the build job needed a second patch. Both now call
`.github/scripts/restore-node-modules.sh`.

Extracting it also removes the trap that fix was for. It is deliberately ONE
step rather than unpack-then-conditional-fallback: a step `if:` with no
status-check function is implicitly ANDed with `success()`, so a fallback gated
on a previous step's outcome cannot run after that step fails hard. A single
script cannot half-fire, so the hazard is gone by construction rather than by
remembering `continue-on-error`.

Also fixes the S6582 optional-chain smell in `resolveRowIndex`, and the header
comment that still described the (now removed) cache handoff.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 10:52:31 +02:00
Ruslan Konviser 6ec0bfd3ac Merge remote-tracking branch 'origin/develop' into fix/e2e-green-post-views
# Conflicts:
#	packages/plugins/docs-ui/src/lib/pages/page-editor/document-page.component.ts
#	packages/plugins/docs/src/lib/services/document.service.ts
2026-08-18 06:42:43 +02:00
Ruslan Konviser 9fa61a55b1 Merge pull request #9997 from ever-co/refactor/registry-composite-action
refactor(ci): extract the 66x-duplicated registry blocks into composite actions
2026-08-17 15:08:36 +02:00
Ruslan KonviserandClaude Opus 5 a71e220509 fix(ci): keep the credential scrub INLINE — a local action cannot load if checkout fails
cubic raised this eight times on #9997 and it is right; worse, my PR body
explicitly dismissed it with "there's no workspace to scrub anyway", which is
false for the runners this step exists to protect.

A local composite action is resolved FROM the workspace. If `actions/checkout`
fails, `uses: ./.github/actions/scrub-registry-credentials` cannot be located, so
the `if: always()` step ERRORS instead of running. The previously inlined `run:`
had no such dependency and would still have scrubbed.

That matters because 30 of these 66 jobs check out with `clean: false`:

  jobs with the scrub: 66  of which checkout clean:false = 30

On those, a failed checkout leaves the PREVIOUS run's workspace in place — and any
live `//packages.ever.co/:_authToken=` in its .npmrc — readable by the next job on
that runner. Exactly the leak the step was added to close.

So the split is now drawn on the actual dependency:

  Configure Registry  -> composite action. It edits .npmrc/.yarnrc/yarn.lock in the
                         repo, so it genuinely cannot run without a checkout.
  Scrub credentials   -> stays INLINE, with a comment saying why, so nobody
                         "finishes the refactor" later and silently reopens this.

`.github/actions/scrub-registry-credentials/` is deleted rather than left unused —
an unused action that looks like the obvious next step is a trap.

Verified the scrub is unchanged, not merely restored:
  develop distinct scrub bodies: 1   now: 1
  scrub body identical to develop: true

Net effect vs develop: 19 files changed, 1192 insertions(+), 7110 deletions(-) —
about 5,900 lines of duplication still removed, without trading a credential leak
for it.

Also from this round (cubic P3, 2 comments): the `ever-k8s` retention rationale is
back beside the predicate at all 66 call sites. I had dropped it when moving the
predicate, which is the same comment-drift the extraction was meant to stop.

Not addressed, and pre-existing on develop rather than introduced here: the scrub's
symlink and staged-index edge cases (cubic D1-D3) and `always-auth=` without an
`_authToken=` line (D7). They apply equally to the merged develop implementation
and are worth their own change rather than widening a refactor.

Guards: 67 files parse; empty-${{ }} grep clean; cspell clean (reworded my own
coinage rather than adding it to the dictionary).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 15:06:31 +02:00
Ruslan KonviserandClaude Opus 5 adbc9bcbf7 refactor(ci): extract the 66x-duplicated registry blocks into composite actions
Asked for by CodeRabbit (Major) and cubic (P3) during PR #9993 review, and earned:
the registry-selection block and the credential scrub were duplicated verbatim in
66 jobs across 18 workflows, and over that PR the SAME lockstep edit had to be
applied to all 66 copies on nine separate occasions. On three of them the code
was updated while the adjacent comment was not - the exact drift they warned
about, arriving inside the PR that was arguing it would not.

  .github/actions/configure-registry/            (4 inputs)
  .github/actions/scrub-registry-credentials/    (no inputs)

  18 files changed, 528 insertions(+), 12786 deletions(-)

BEHAVIOUR IS UNCHANGED - verified, not asserted. Compared against origin/develop:

  distinct inlined Configure bodies on develop : 1
  distinct inlined Scrub bodies on develop     : 1
  Configure shell identical                    : true   (byte-for-byte)
  Scrub shell identical                        : true   (byte-for-byte)
  develop env keys / action step env keys      : identical set
  call sites passing the SAME expressions      : 66/66

The shell is moved, not rewritten: the action's run body is byte-identical to what
was inlined, and each call site passes exactly the expressions that job's env
block used before - including the per-job
`contains(matrix.os,'self-hosted') || contains(matrix.os,'ever-k8s')`, which stays
at the call site because it is the one genuinely per-job value.

ONE REAL BEHAVIOURAL DIFFERENCE, checked rather than assumed: a local composite
action only resolves once the repo is checked out, which an inline `run:` did not
require. Verified checkout precedes both call sites in all 66 jobs (66/66, 66/66).
The residual case is a job whose checkout itself fails: the `if: always()` scrub
would then error on a missing action instead of running and finding nothing to do.
That is noisier but not harmful - with no checkout there is no workspace .npmrc to
scrub.

Guards: 68 files parse (66 workflows + 2 actions); `grep -rn '${{ *}}'` over
.github/workflows and .github/actions is empty; cspell clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 14:35:13 +02:00