mirror of
https://github.com/ever-co/ever-gauzy.git
synced 2026-10-02 01:54:50 +08:00
BullMQ and Redis were already live — `apps/worker` registers a scheduler root, loads plugins
via `PluginModule.init()`, and runs jobs today. But `GAUZY_DOCS_QUEUE_ENABLED` defaulted to
false everywhere, so the docs queue was never registered in either process: the worker idled
and the API did every extract / classify / embed / OCR / thumbnail stage INLINE. That is the
exact outcome the worker module's own comment warns about.
The API could not simply turn it on: it had no `SchedulerModule.forRoot()` anywhere in its
graph, and a `@Processor` without a connection throws `Worker requires a connection` during
`onModuleInit` — which is what crash-looped demo and stage earlier today.
What this adds:
* `isSchedulerQueueRootEnabled()` in `@gauzy/scheduler` — ONE predicate every plugin-hosting
process evaluates, so "does a Bull root exist in THIS process" stays a derived answer rather
than an operator keeping two variables in sync.
* A producer-only root in the API (`enabled: false, enableQueueing: true`). `enableQueueing`
is what registers `BullModule.forRoot`, and `enabled: false` stops `registerSchedules()` and
short-circuits the job runner — so the API can enqueue but can NEVER fire the cron jobs the
worker already owns. Duplicated scheduled jobs would be the obvious way to break this.
* The same root in `SeederModule`, because `SeederModule.forPlugins()` loads the same plugin
list — `yarn seed` would otherwise have been the one plugin-hosting process where the derived
gate lied.
* `isDocsQueueEnabled()` derives from the shared predicate; a new `isDocsQueueWorkerEnabled()`
gates the consumer as a strict subset, so a `@Processor` can never outlive its queue.
The consumer defaults ON deliberately. A deployment with no worker must still process its own
documents; producing without consuming would leave them in UPLOADED forever, which is strictly
worse than inline. Set `GAUZY_DOCS_QUEUE_WORKER_ENABLED=false` on the API deployment to make it
a pure producer once a worker is running.
Verified — a green build proves nothing for this defect class, so each state was booted:
* no Redis -> no root, dispatch mode INLINE, boots clean
* Redis up -> queue REGISTERED, consumer REGISTERED, dispatch QUEUED
* producer-only override -> queue REGISTERED, consumer absent (by design)
* apps/worker with Redis -> queue + consumer REGISTERED; `bull:docs-processing:meta`
visible in Redis
* 🛑 Redis ENABLED but UNREACHABLE (nothing on the port) -> Nest bootstrap completes in 24ms
and `SchedulerQueueService` is present. This is the case nobody had covered and the one
that matters: it is every deployment where Redis is configured but momentarily down. The
root does not block boot, and `DocsQueueService.enqueue()` already degrades to the inline
runner when the enqueue itself fails, so an outage cannot strand a document either.
A real defect the boots caught: this workstation's `.env.local` sets `WORKER_QUEUE_ENABLED=false`,
so the worker had no root while a Redis-only gate still said "queued" — it died on
`Job "docs-reconcile-schedule" targets queue "docs-processing" but queueing is disabled`. Rather
than documenting that two variables must agree, `WORKER_QUEUE_ENABLED` now feeds the shared
predicate, where it can only ever REMOVE a root and degrade to inline.
Tests: plugin-docs 570 (43 suites), scheduler 12 (new project), worker 6. 🛑 `nx test worker`
reported EXIT=0 while its suite was failing to LOAD — the worker numbers above come from running
jest directly, and the spec's `@gauzy/scheduler` mock was fixed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>