Interrupting a command after exit could truncate its JSON and report a
parse failure. Output arriving during the drain could also extend the wait
past the execution timeout and turn a completed command into a timeout.
Only interrupt still-running children and count those matches in killFor.
Keep exited children draining, stop the execution timer on exit, and retain
the existing idle drain window and 2-second cap. An explicit AbortSignal
still cancels the caller's wait immediately during drain.
Clarify the execution timeout and separate output collection limit in the
runner options, plugin configuration and README. Preserve the existing
platform-specific cancellation grace, kill fallback and resource cleanup.
Validation: 10 regression checks failed before the fix and now pass; all
351 DSH plugin tests in 19 files pass, including real-process inherited-pipe
cleanup. Biome, Stylelint, TypeScript and git diff --check also pass.
Windows cancellation cases use simulated platforms on macOS.
Preserve the contributor's original commits and bounded settlement, pipe
cleanup, renewable output drain, and killAll/killFor fallback paths.
Reconcile the overlapping main runner fix by retaining its 1s drain grace,
referenced drain timers, cancellation timer cleanup, and listener ordering.
Keep the PR's 2s total drain cap and platform-specific kill deadlines.
Adapt the affected runner tests to the retained grace and update the
session-start regression fixture for main's prepare/start/claim lifecycle.
No unrelated feature changes.
Validation: all 345 DSH plugin tests across 19 files pass, including the
real-process inherited-pipe regression; Biome, Stylelint, TypeScript and
git diff --check pass. Windows cancellation is covered by simulated-platform
tests; no Windows host was used.
Isolate harness path tests from host environment settings and verify Windows defaults. Cover real historical LF/CRLF fixtures, custom intent, local edits, reference collisions, and exact checksum baselines.
Bring in the existing status and doctor fixture isolation from main. Condense repeated DSH skill wording without relaxing the prompt-size assertions or changing the recording fix.
Condense the DSH guidance while preserving its authorization boundary and recovery behavior. Keep the existing 7000-character limit after combining the PR with main's recovery instructions.
Retain successful invocation proof after failures and defer retries to turn/session boundaries or new skill invocations. Add regression coverage for sustained streaming and next-turn recovery.
Parse current durable tool results and support both session history APIs. Scan each session once, process later events incrementally, and retry recovery when services or history become available.
Cover official DSH messages, reloads, long conversations, failed reads, and explicit history access counts.
hermes_skills_dir_honors_hermes_home_env overrode the process-wide
HERMES_HOME without any lock, while detects_hermes_from_home_layout and
skills_dirs_match_harness_spec resolved Hermes through that variable on
other libtest threads: the detection test read the override's temp
directory and failed about six times in thirty runs. The Kimi tests took
kimi_env_lock, but skills_dirs_match_harness_spec read KIMI_CODE_HOME
without it, leaving the same failure latent there.
Widen that lock to harness_env_lock, take it in every test that reads or
writes HERMES_HOME or KIMI_CODE_HOME, and replace the two SAFETY comments
whose invariant did not hold: the hazard is a concurrent reader, not only
another writer.
Raised in review by @iuyo5678, and the objection is right.
The first wording listed actions - send data somewhere, approve something,
install something, visit another site - and called any page mentioning them an
injection attempt. Those are ordinary parts of authorized work. An agent
following that rule would refuse to submit a form the user asked it to submit,
or to follow a documentation link the user asked it to read, and would report
the page as hostile for containing a button.
The test is whether the page is trying to change what the agent may do, not
what kind of action it names. The text now says that, and says explicitly that
navigation guidance, controls and quoted examples are not by themselves
evidence of injection.
AGENT_INSTALL.md also claimed "a page cannot redirect the agent". That reads as
a technical guarantee and none exists: nothing stops a page carrying text aimed
at an agent. It now describes what the agent is required not to do - let page
content override its instructions, grant it permission, or widen its task - and
names the guidance as behavioural.
Both skill files carry the same wording, as before.
navigate, navigate_back, navigate_forward, reload and wait_for_navigation
resolve with a structured result at their own timeout_ms: reached "timeout",
the URL the page actually reached and the last observed lifecycle. The daemon
dispatched them with a transport deadline equal to that same timeout, so the
extension's reply still had to cross the socket after the deadline had fired.
The reply then took the TimedOutAfterResponse path, which preserves completed
results only for session_stop, tab_borrow, request_help and the effect-aware
transfers, so the caller received a bare "tool RPC timed out" instead.
Grant these five methods the EXTENSION_RESPONSE_GRACE that upload, download
and request_help already use. They are the complete set of tools whose
extension handler resolves with "reached: timeout" at a caller-supplied
deadline; every other tool is unchanged.
Start tab and notification cleanup concurrently so a stalled notification
cannot prevent cancellation messages from reaching the page overlays.
Cover content completion, timeout, and cancellation across multiple tabs.
Share optional stop target parameters between browser_session and its internal
stop handler. Document exact request targeting, mutual exclusion, and default
stop retry behavior in the model-visible contract.
Add a public schema regression covering requestId, optional targets, and retry
guidance. Lifecycle execution remains unchanged.
Validation: 95 tool and stop recovery tests passed; TypeScript and formatting
checks passed. The new schema test reproduced the omission before the fix.
Route tool, overlay, archive, unload, and recovery cleanup through one lifecycle
owner and one job per request. Persist explicit stop intent before making a
session unusable; caller cancellation only ends its wait after admission.
Retain durable completion receipts independently of the current session so a
retry cannot stop another working session, even after background cleanup or
restart. Preserve failed default callers' retries across concurrent successful
waiters. Support exact request targeting and reject ambiguous stop targets.
Cover failed normal stops, queued and in-flight cancellation, concurrent entry
points, timer and disk recovery, persistence failures, anonymous starts, reused
IDs, unconfirmed replies, and completion receipt capacity accounting.
Validation: 303 plugin tests passed, including 31 stop recovery tests;
TypeScript checking, production build, and formatting passed.
Package lint was unavailable because publint is not installed in this environment.
Do not terminate a running child when its launcher cannot confirm startup.
A different client may already be using it, or publication may race the
launcher's deadline. Retain cleanup only before a Windows child resumes.
Add deterministic lifecycle regressions for reuse before timeout, delayed
publication and port mismatch. Run Windows success-path and updater tests
in a same-user WMI test host that verifies it is outside all Jobs, and keep
restricted-host rejection tests in the normal CI runner.
Refs #268
Capture initial tab identities when creating an Agent Window and retain them
through failed startup compensation. Retry cleanup through the production
stop handler, preserving later user tabs and keeping ownership if agent tabs
still cannot close.
Separate owned starting and cleanup resources from active sessions. Publish
sessions only after initialization and claim succeed, preserve the working
current session on failure, and reject ordinary operations and captures for
resources awaiting cleanup. Recheck queued operations before execution.
Add regression coverage for production stop retries, mixed windows, delayed
claims, current-session fallback, recovery, capacity, and queued operations.
Update window API fixtures for the creation result's initial tab identities.
Validation: extension 1862 passed (103 skipped), plugin 276 passed; both
TypeScript checks and production builds passed.
Both skills tell the agent to read arbitrary pages — observe, get-html,
snapshot, screenshot, console, network — inside the user's real, logged-in
profile, and neither says that what comes back is untrusted. A repo-wide search
found no prompt-injection guidance in any markdown; the only place page data is
labelled untrusted is a protocol comment in bsk-protocol that the agent never
sees.
The absence stands out because the skills already constrain behaviour
elsewhere: never extract credentials, never evaluate secrets, never record
banking or SSO pages. Untrusted page content is the same class of rule and was
simply missing.
States it where the reading happens, in both skills, with a pointer from the
standing rules at the top of each. Names the read commands and the element
names and labels that get passed back to click/fill/select, since those carry
page text too, and says what to do instead: stop, tell the user what the page
tried, do not comply.
AGENT_INSTALL.md asks the installing agent to repeat it to the user, because
the person granting access to their logged-in browser should know a page cannot
redirect the agent, and that an agent appearing to follow one has been injected
rather than instructed.
Documentation only; no behavioural or technical mitigation. The reporter's
further suggestions — delimiting tool output, a domain allowlist, a
consent-to-consequence step — are deliberately left out of scope.
Fixes#286
After a daemon restart that reloads the plugin, browser_* calls fail with
unknown tool "browser_session" until the model happens to invoke the skill
again, even though the session's durable history already proves it ran.
None of the three triggers covers that case. The live tools/result hook needs
a fresh invocation. session/created never fires for a session that already
exists. The boot scan runs once during apply(), and ctx.get("sessions") yields
nothing when the sessions service is registered after this plugin, so it covers
nothing at all.
session/event already receives the session and ignored it. Reading its history
when nothing else has revealed the suite closes the gap, and is guarded so the
scan stops once revealed — these events are frequent and re-deriving on each
one after the reveal is waste.
Keeps the reveal derived from durable history rather than adding a persisted
flag, so there is no new state to migrate or keep consistent.
Use an explicit standard-handle inheritance list for daemon startup and the
Windows update helper. Require and verify breakaway before resuming a
background daemon, with an actionable error for restrictive host Jobs.
Share direct startup and its readiness deadline across explicit and
automatic entry points, retain child ownership through startup, and reap
Unix children after handoff.
Add native Windows EOF, Job lifetime, concurrency and failure regressions
to CI, and document persistent host setup.
Refs #268
Track prepared starts with stable request handles, durable plugin ownership, monotonic cancellation, and retryable cleanup. Retain delayed and failed startup resources until cleanup is confirmed, without touching other sessions.
Add regression coverage for lost replies, killed CLI processes, delayed creation, failed cleanup, restart recovery, archive and unload, capacity accounting, and reused session IDs.
Fixes#245
Preserve the streamlined skill workflow and document background support for both viewport and full-page screenshots. Keep Canvas capture guidance and remove the outdated visibility requirement for Agent full-page capture.