- Replace process-global healthState.consecutiveNavFailures with
Map<userId, UserNavHealth> for per-user failure tracking (#9321)
- Wire recordNavSuccess(userId)/recordNavFailure(userId) into all
navigation routes: POST /tabs, POST /tabs/:tabId/navigate,
POST /tabs/open, POST /navigate, and handleRouteError for /act (#9323)
- Add recoverUserSession(userId) for session-scoped recovery instead
of restartBrowser() killing all sessions on nav failures (#9322)
- Remove manual browserLaunchPromise = null in restartBrowser() —
ensureBrowser() owns the single-flight primitive (#8554)
- Clean up per-user nav health on destroySession()
- Aggregate per-user failures in /health endpoint
- Add tests/unit/navHealthRecovery.test.js (8 tests)
Fixes#9321, #9322, #9323
Relates to #8554
#7266 fixed the "Found property <root>.viewport.isMobile ... not
described in this scheme" juggler rejection by switching context
creation to viewport: null in probeGoogleSearch() and getSession(), but
missed a third call site: the active health probe's setInterval calls
browser.newContext() with no arguments, which still hits Playwright's
implicit default viewport and the same isMobile rejection.
In practice this means the health probe itself can trip the very
failure it's meant to detect, log it as "health probe failed", and
call restartBrowser('health probe failed') -- restarting a browser
that was otherwise fine, repeatedly, on whatever interval the probe
runs at.
Extends tests/unit/launchCompat.test.js (added in #7266, same
source-contract idiom as tests/unit/noSecrets.test.js) with a fourth
case covering this call site. Confirmed the new test fails against the
unpatched line and passes with the fix.
Refines the session-reaping approach from #8319 without adding admission queues or HTTP 429 behavior.\n\nFixes #8555\nRefs #8319\n\nCo-authored-by: batumilove <batumilove@users.noreply.github.com>
Both the global express.json() parser and the /evaluate endpoint now
read from CAMOFOX_MAX_BODY_SIZE (default: 100kb) via CONFIG.maxBodySize.
The setting is defined in lib/config.js alongside all other env vars.
Closes#8397
The upload route had inline millisecond literals (4000, 12000, 3000,
10000, 500, 1500) scattered through its two attach strategies. Replace
them with named UPLOAD_*_MS constants declared next to the route, and
expose the overall wait budget as an optional `timeout` request field.
- UPLOAD_UI_TIMEOUT_MS (default 12000) backs the request's `timeout`:
the budget to wait for an upload UI (panel input or native chooser).
Non-numeric / <= 0 values fall back to the default.
- The panel-poll window derives from that budget minus
UPLOAD_PANEL_MARGIN_MS, preserving the original 10000/12000 split so a
late native chooser is still caught after polling stops.
- UPLOAD_INPUT/FOCUS/CLICK/REFS/POLL/SETTLE_MS name the per-call bounds.
Defaults reproduce the previous behavior exactly. OpenAPI documents the
new `timeout` field; tests cover timeout resolution and assert the route
carries no bare millisecond literals.
Attach a file to an upload control without going through the native OS
file dialog. Two strategies are tried in order:
1. If an <input type="file"> is already present, call Playwright
setInputFiles on it directly (works for hidden inputs).
2. Otherwise arm a filechooser listener, activate the trigger element
(ref or selector) via keyboard (focus + Enter) with a forced click
as fallback, and setFiles on the resulting chooser. Also polls for
an in-app panel <input type=file> that mounts after activation.
The panel-input path is preferred over the native chooser so a control
that surfaces both (e.g. LinkedIn's media picker) attaches the file
exactly once rather than producing a duplicate.
Paths must be visible inside the container (e.g. a bind-mounted dir);
the route guards with fs.existsSync and returns 400 file_not_found
otherwise. Runs under the same per-user and per-tab locks as the other
interaction routes.
Includes OpenAPI documentation and unit tests (request validation +
source-contract assertions).
Bound new-page creation with a configurable 10-second deadline, replace only the affected user context on timeout/dead-context failures, and retry once.
Co-authored-by: Pradeep Elankumaran <pradeep@askjo.ai>
Add an authenticated storage-state reset endpoint that closes the live browser context without checkpointing, waits for in-flight persistence writes, and removes the saved state so the next session starts fresh.
Co-authored-by: PluginsKers <ikers@foxmail.com>
Add optional CAMOFOX_BIND_HOST configuration, forward it to launched server subprocesses, document the setting, and report the effective listener address at startup.
Co-authored-by: SnapsNoCaps <139733767+HODL-Community@users.noreply.github.com>
The previous two commits fixed the ordering of our own shutdown
sequence (server:shutdown now blocks on emitAsync, forceTimeout/close
run before it), but production verification showed the storage-state
race still reproduced 2/2 times even with both fixes deployed.
Root cause: Playwright's launcher defaults handleSIGTERM/SIGINT/SIGHUP
to true, registering its own process signal handlers independent of
ours. On SIGTERM, that handler sends Browser.close straight to Firefox
over its debug transport -- outside pluginEvents entirely, racing
ahead of gracefulShutdown() and the persistence plugin's checkpoint
regardless of how carefully our own code is sequenced.
Set all three handleSIG* options to false so gracefulShutdown() is the
sole authority over shutdown; closeBrowserFully() already does an
explicit browser.close() plus a force-kill fallback, so nothing is
lost by disabling Playwright's own signal-driven close path.
Verified against a real deployment: reproduced the original failure
2/2 times with only the prior two commits applied, then 2/2 successful
"storage state persisted" after adding this fix, via docker restart
immediately (and with a 5s settle) after creating a real session.
No new unit test added for this specific change -- firefox.launch()
isn't currently unit-testable in isolation without substantial mocking
scaffolding disproportionate to a 3-line options change, and the
production verification above is stronger evidence than a mocked unit
test would provide.
The previous commit made server:shutdown blocking (emitAsync) to fix
the storage-state persist race, but that reordering had three side
effects: forceTimeout was armed only after the awaited emit resolved
(a hung listener disables the force-exit watchdog entirely), a
rejecting listener could abort the rest of gracefulShutdown via
Promise.all's fail-fast semantics with no top-level catch, and
server.close() ran later, widening the window where new sessions
could still be created during shutdown.
Move forceTimeout and server.close() back to running immediately, and
catch a rejecting server:shutdown listener so it can't skip the rest
of cleanup. The checkpoint-before-close ordering from the previous
commit is unchanged.
gracefulShutdown() fired the server:shutdown event with plain
EventEmitter#emit, which does not wait for async listeners. It then
called closeAllSessions() immediately after, racing the persistence
plugin's in-flight context.storageState() checkpoint against
context.close() for the same session -- intermittently failing with
"Target page, context or browser has been closed" and silently
dropping unsaved storage state on restart.
closeSession() already gets this right for session:destroying via
emitAsync(); server:shutdown just wasn't using the same helper.
Fixes#7162
When a normal click attempt fails because the element detached after a page
change (SPA re-render between snapshot and click), the mouse-sequence
fallback called locator.boundingBox() with no timeout. Playwright waited its
default 30s for the element to resolve, blowing the entire HANDLER_TIMEOUT_MS
budget and surfacing as:
click failed: action timed out after 30000ms -> 500 internal error
Worse, that generic 'timed out after' error is classified by isTimeoutError()
as a navigation timeout, so handleRouteError() destroyed the whole user
session (fresh proxy/context) over a stale ref.
Fix:
- boundingBox() is now bounded to min(3s, remaining handler budget),
floored at 500ms.
- On timeout it throws a 422 'Element not actionable' error whose message
deliberately avoids the 'timed out after' phrase, so the session survives
and the client gets an actionable hint (snapshot + retry) in ~3s
instead of a 500 after 30s.
Adds tests/unit/clickBoundingBoxTimeout.test.js covering the source
contract (no bare boundingBox() calls), the budget math, and the error
classification.
buildRefs() traverses child frames and runs ariaSnapshot on each one not matched
by IFRAME_SKIP_PATTERNS. Some sites inject a bot-detection iframe (PerimeterX
px-iframe-*/px-captcha, hCaptcha, Arkose/FunCaptcha, DataDome) that is short-lived
and detaches mid-ariaSnapshot. The ariaSnapshot call then hangs until the handler
timeout (30s), and the click/snapshot that triggered the ref rebuild returns a
spurious 500 even though the underlying action succeeded.
Concretely on LinkedIn: PerimeterX is injected on authenticated actions such as
the connection-invite 'Add a note' modal, so every such click 500s and write
flows silently fail.
Fix: add px-iframe/px-captcha/perimeterx/captcha/hcaptcha/arkose/funcaptcha/datadome
to IFRAME_SKIP_PATTERNS so buildRefs skips them (same treatment recaptcha already
gets). Both iframe loops reference this constant, so the single edit covers ref
building and snapshot YAML.
Tested: an automated LinkedIn connection-invite flow that previously 500'd on
profiles showing the intermediate 'Add a note' step now completes and verifies
Pending.
parseInt(...) || 300000 treated an explicit 0 as falsy and fell back to
the 5-minute default, and scheduleBrowserIdleShutdown armed a timer
regardless. Parse with Number.isFinite so 0 is preserved, and skip
scheduling when the timeout is <= 0, matching the README's documented
0 = never contract.
Closes#4942
Adds safePageUrl() that catches destroyed page access. Applied to popup
handler where popupPage.url() can throw if popup closes immediately.
Addresses JO-BROWSER-18 TypeError: Cannot read properties of undefined
(reading 'url').
Tab IDs on Fly are '{machineId}_{uuid}'. The UUID regex check was
testing the full prefixed string, always failing, so stale tabs after
browser restart got 404 instead of 410. Now extracts the UUID portion
before validation.
Tracks _lastBrowserStopReason to distinguish intentional idle/admin
stops (200) from unexpected deaths like browser_disconnected,
memory_pressure, browser_rss_pressure (503 + warm retry).
Fly health checks will now detect and route around machines with
unexpectedly dead browsers, while intentional idle shutdown still
passes health checks as before.