3702 Commits
Author SHA1 Message Date
lileiandWillem Jiang f840e843d3 fix: track converted upload ownership for outlines and listings (#6101)
* fix: track ownership of converted upload markdown

* fix: invalidate overwritten companions and restore branch ownership

* docs: explain legacy upload companion behavior on upgrade

* docs: shorten upload companion guidance to fit instruction budget

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 20:36:35 +08:00
Daoyuan LiandWillem Jiang 7c52d5a187 fix(middleware): isolate temporary files for concurrent tool outputs (#6150)
* fix(middleware): isolate temporary files for concurrent tool outputs

* docs: trim tool output middleware guidance

* fix(middleware): preserve output permissions and clarify crash cleanup

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 20:17:13 +08:00
xiaodu55andxiaodu55 fc27c94b67 test(extensions): make the extension-manager suite runnable on Windows hosts and under host uv mirrors (#6142)
* test(extensions): make the extension-manager suite runnable on Windows hosts and under host uv mirrors

- use the file's existing platform-aware .venv interpreter idiom in the two
  rollback tests (Scripts/python.exe on nt);
- compare the uv add source argument as a Path so the assertion is
  separator-neutral;
- pin UV_DEFAULT_INDEX in an autouse fixture so host mirror configuration
  cannot break the source-build tests; tests installing their own local
  index keep precedence.

Fixes #6141

* test(extensions): drop the extra-index env channels from the suite environment

UV_INDEX and UV_EXTRA_INDEX_URL entries take priority over the pinned
default index and could still break the source-build tests. UV_NO_CONFIG
is intentionally not used: it would also drop the fixture's own
[tool.uv] pyproject settings the command-shape assertions rely on.

---------

Co-authored-by: xiaodu55 <xiaodu55@users.noreply.github.com>
2026-10-01 20:14:59 +08:00
yetuge 7eb41cfc93 refactor(gateway): hoist the persistent-write drain into one shared helper (#6151)
* refactor(gateway): hoist the persistent-write drain into one shared helper

Each persistent-write surface forked its own per-router drain wrapper:
_drained_write in routers/agents.py (action label, expected_errors
pass-through, lost-failure type-only logging), _run_store_mutation in
routers/subagents.py (bare drain), and the inline await_drained call in
routers/managed_models.py. Hoist the superset into
app.gateway.persistent_writes.run_drained_write and point all three
routers at it, per the review suggestion on #6087 (its family
predecessors #6078/#6092/#6093 have since landed).

The subagents and managed-models call sites now also emit the type-only
lost-failure log line on a drained-then-failed write (previously
silent), matching the tradeoff acknowledged in the #6087 review.
artifacts.py and personal_mcp.py keep their direct await_drained forms:
the former drains a coroutine with its own rollback flow, the latter
needs no label or expected-error handling.

Router tests move their monkeypatch targets to the shared name; the
lost-failure assertion pins the helper's logger. AGENTS.md points the
next persistent-write surface at the canonical implementation.

Signed-off-by: yetuge <2219677952@qq.com>

* fix(gateway): quiet expected domain errors and scope the AGENTS.md drain claim

Review follow-ups on #6151:

- The subagents create/update call sites pass (ManagedSubagentExistsError,)
  and (FileNotFoundError,) so the shared helper stays silent on routine
  4xx outcomes (duplicate-name 409, stale-store 404 race), matching
  create_agent's (AgentExistsError,).
- The managed-models save passes (HTTPException,), so the helper stops
  double-logging failures that _save_with_failure_logging already
  reports and stays silent on its routine 4xx outcomes.
- The AGENTS.md sentence now names the three converted routers instead
  of implying every persistent-write surface uses the helper; the
  shortened form also keeps the file under the guidance byte budget
  that agent-guidance enforces.

Signed-off-by: yetuge <2219677952@qq.com>

---------

Signed-off-by: yetuge <2219677952@qq.com>
2026-10-01 20:13:57 +08:00
creed 71f109275b feat(serper): support relative search time ranges (#6113)
* feat(serper): support relative search time ranges

Signed-off-by: 97three <2212371308@qq.com>

* docs(serper): address review and trim inherited guidance

---------

Signed-off-by: 97three <2212371308@qq.com>
2026-10-01 20:12:43 +08:00
Inference_ a8dc1da205 fix(runtime): honor both cursors when reading thread messages (#6136)
* fix(runtime): honor both cursors when reading thread messages

* docs(runtime): clarify run-scoped combined cursor bounds
2026-10-01 20:11:10 +08:00
Daoyuan Li 59743568ba fix(frontend): ignore inherited artifact language keys (#6152) 2026-10-01 19:42:06 +08:00
Guo-YixinandWillem Jiang d2cfe4a986 fix(models): honor Codex SSE terminal events (#6155)
* fix(models): honor Codex SSE terminal events

* fix(models): preserve malformed Codex error diagnostics

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 19:32:49 +08:00
Wenchao An 9f5897cef2 feat(frontend): reserve slash menu for built-in commands (#6154)
* feat(frontend): reserve slash menu for built-in commands

* chore(frontend): remove unused skill removal translations
2026-10-01 19:31:31 +08:00
Zeren Wang d3e8e78f3f feat(authz): middleware-declared tool path Layer 1 now covers tools declared by middlewares (#6105)
* feat(authz): extension tool authorization PR2 — middleware-declared tool path

Layer 1 now covers tools declared by middlewares, which LangChain merges
into the bound tool set after the host's explicit-list filter:

- new agents/middlewares/tool_declarations.py: record the ordinary pass's
  verdict, collect declarations from the built stack, decide only the
  delta seeded by the first pass (a name denied for this build stays
  denied), narrow the stack build-locally on independent state-preserving
  copies (direct __dict__ write, refused up front for tools data
  descriptors, never mutating caller-owned instances), and verify the
  exact bound stack after the final normalization copy.
- IsolatedMiddleware ships a state-preserving __copy__ so narrowing
  survives full/delta normalization and copies of copies.
- TodoMiddleware degrades when write_todos is denied: no system-prompt
  injection, no completion/context-loss reminders, no jump_to.
- Layer 2 forwards host-resolved tool provenance into AuthzRequest.context
  and binds the infrastructure exemption to the concrete host-created
  tool object instead of a name set (tool=None goes to the provider).
- Wired into all four assembly paths (lead x2, native subagent with the
  decision offloaded via asyncio.to_thread, DeerFlowClient).

Tests: new test_middleware_declared_tools.py (35) and
test_layer2_tool_provenance.py (7); blocking-IO anchor pinning the
subagent offload (teeth verified red->green); placement-guarantee and
enforcement suites extended; executor test doubles converted for the
async _create_agent. Behavior change called out in CHANGELOG: with
authorization enabled and a tools policy that excludes write_todos,
plan-mode builds no longer bind it (built-in RBAC default unchanged).

* fix(authz): close residual gaps from review — narrow non-collectable declarations, dedupe candidates, warn on skipped pass

- apply_declared_tool_view now removes non-BaseTool declarations from the
  bound view when authorization is enabled (they would otherwise bind via
  LangChain's create_tool auto-conversion with no provider decision);
  verify_declared_tool_view fails loudly if one survives normalization
- decide_declared_tools dedupes never-submitted names so the provider sees
  each candidate once
- subagent executor logs a warning when the provider is set but the seeded
  Layer-1 state is missing, instead of skipping the declaration pass silently
- docs: middleware-declared tools must be BaseTool instances with usable names

* fix(authz): fail loudly on non-list/tuple middleware tools declarations

_iter_declared warned and returned () for a 'tools' attribute that is not
a list/tuple, but LangChain's factory iterates the attribute as-is — so a
set/generator/dict_values declaration bound with no Layer-1 decision.
Raise DeclaredToolViewError instead, matching the module's fail-loud
standard for every other un-narrowable shape. Disabled path remains a
strict no-op. Docs updated to match.
2026-10-01 18:26:58 +08:00
b57fa6fb9e feat(scheduler): push scheduled task results to IM channels via a notification outbox (#6135)
* feat(persistence): add notification delivery outbox for scheduled task IM push

New notification_deliveries table and repository (migration 0012) backing
the outbox pattern for issue #4254: idempotent enqueue keyed on task_run_id
plus event plus provider plus target, atomic claim of due rows, and
exponential-backoff retry state moving pending to sending to sent or failed.
Execution status stays in scheduled_task_runs; delivery status lives here.

* feat(scheduler): enqueue and deliver scheduled-task notifications over IM channels

Closes the loop for issue #4254. On run completion the scheduler writes
outbox rows for the task owner's connected channel identities (success and
failure events; interrupts stay silent), and a new NotificationDeliveryWorker
polls due rows and pushes them through the channel's proactive send path.
Channel gains send_notification; WeCom implements it via the WS client's
send_message with errcode checking, the base class raises for providers
without proactive push.

Delivery is best-effort and never shadows execution status: enqueue and send
failures are logged and recorded in the outbox, which retries with backoff.
Channel liveness is re-checked at delivery time, and the worker stops before
the channel service on shutdown. Feature activates only when channel
connections are enabled; no new dependencies.

Verified end to end against a live WeCom bot: scheduled run completion
produced an outbox row that was delivered in under one second.

* feat(scheduler): include scheduled run result summary in IM notifications

The v1 notification only carried the run outcome skeleton (status, task id,
run id). Completed runs now include the agent's final answer so the IM push
is actually useful headlessly.

- The delivery worker accepts an optional resolve_run_summary(run_id,
  user_id) resolver; for run_completed events it enriches the rendered text
  with a "Result:" line (bounded to 1000 chars). The lookup is best-effort:
  resolver failures or missing summaries degrade to the skeleton text and
  never block or fail the delivery.
- The summary is read at delivery time from runs.last_ai_message (the
  convenience column written on run completion), not snapshotted into the
  outbox payload; retried deliveries re-read the freshest value and the
  outbox row payload stays skeleton-only. The claimed outbox row is never
  mutated in place.
- Gateway wiring scopes the read to the outbox row's owner_user_id so a
  stale row cannot pull another user's run content.
- Failed runs keep the error-only skeleton: a partial answer on a failed
  run would be misleading.

Tests: 9 new worker cases (rendering incl. truncation, sync/async
resolvers, failure/missing fallbacks, no resolver call for failed runs).
Full scheduled + outbox + channels regression green (441 passed), ruff
check + format clean.

* fix(scheduler): address review feedback on notification delivery reliability

Review follow-ups for the scheduled-task IM notification outbox:

1. Crash recovery for "sending" rows: claim now stamps updated_at and the
   worker reconciles rows stuck in "sending" past a 10-minute lease back to
   "pending" before every claim. Per-row try/except isolates a crashing
   delivery from the rest of the claimed batch.
2. Double-send race under multiple workers: claim returns exactly the rows
   its own UPDATE...RETURNING flipped (guarded by status='pending'), so a
   rival worker's concurrent claim can no longer be handed out twice.
3. Restore non-ASCII comments and drop the BOM accidentally introduced in
   test_scheduled_task_service.py.
4. Manual "run now" completions no longer enqueue IM notifications; the
   user is already watching the UI. Missing trigger metadata fails safe
   towards notifying.
5. Channel-outage failures no longer consume the retry budget: they are
   parked with count_attempt=False on a flat 15-minute backoff so an
   hours-long outage cannot exhaust retries and drop the notification.

Tests: repository 16, worker 21, scheduled-task service 31; targeted
regression across notification/scheduler/channel suites: 511 passed.

* fix(scheduler): address round-2 review on IM notification delivery

Park WeCom transport/outage failures without burning retries, bound worker stop, include task titles, align enqueue with delivery, and document the IM push behavior.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(scheduler): complete IM notification docs and drop unused list API

Add backend/channels/migrations AGENTS notes, CHANGELOG and config.example hints for the outbox push path, and remove the unused list_by_task_run repository method flagged in review.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(scheduler): address round-3 review on IM notification delivery

Cap channel-down parking by row age and parked attempts, redact IM egress
content, narrow WeCom transport exceptions, add lifespan wiring tests, and
run ruff format so lint-backend passes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(scheduler): redact entire PEM blocks in IM egress text

The previous BEGIN-only pattern left key material and the END marker in notifications.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(persistence): align 0014 parked_attempts with bootstrap schema tests

Backfill existing rows with a temporary server default, then drop it so create_all matches Alembic, and move HEAD anchors to 0014.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(persistence): format notification tests and shorten alembic revision id

ruff format --check failed on test_notification_delivery_worker.py.
0015_notification_delivery_parked_attempts exceeds alembic_version.version_num
(VARCHAR 32) and truncates on Postgres CI; use 0015_parked_attempts instead.

* test: fix post-merge CI failures after migration renumber

Update migration 0015 test HEAD expectation to 0017_parked_attempts and add queue_timeout_seconds to the lifespan notification worker mock scheduler config.

* fix(persistence): chain the notification outbox migrations after 0026

Main's head moved from 0016 to 0026_mcp_task_lease_tokens, so the two
outbox revisions are renumbered and re-parented:

- 0017_notification_deliveries -> 0027_notification_deliveries
- 0018_parked_attempts -> 0028_parked_attempts

Only the revision ids, parents and the migration guide change.

* fix(scheduler): adapt the notification hook and its docs to main

- tests: the completion tests assert on the atomic complete_run record;
  new cases cover a completion that was not recorded, a task deleted
  meanwhile, a failed task read, and the outbox being off.
- AGENTS.md: drop the root note and shorten the backend note to a pointer;
  the two notes pushed two effective guidance chains past the 96 KiB limit.
- CHANGELOG: move the entry from the released 2.1.0 section to Unreleased
  and add its link references.

* test(persistence): move the chain-head pin to the outbox migrations

test_migration_0026 pinned 0026 as the head, so it failed once the outbox
revisions chained after it. It now checks chain membership like the older
migration tests, and a new test file owns the head pin, checks the parents
and the 32-character revision id limit, and runs the upgrade and downgrade.

* fix(gateway): wire scheduled-task notifications after the channel service starts

The outbox wiring sat in the scheduler block and asked for the running
channel service there. #5035 moved the channel service start to after the
scheduler, and the two changes merged without a conflict, so on every
startup the lookup returned None and the feature disabled itself with the
"write-only outbox" warning. The lifespan tests patched the lookup to a
constant and kept passing.

The wiring now runs in its own step after the channel service start. The
scheduler is built as on main, and ScheduledTaskService gains
attach_notification_outbox() so the enqueue side is enabled once the
delivery worker is running. The guard, the worker and the shutdown order
are unchanged.

Tests: the lifespan helper can model the real accessor (None until the
channel service has started); a new case fails without this fix.

* fix(scheduler): enqueue from scheduler start and keep the push when the title read fails

Follow-up to the startup-order fix, from review.

- The enqueue side is wired with the scheduler again (it only needs the
  durable table), and only the delivery worker waits for the channel
  service. Attaching it later left a window after the scheduler started in
  which a finished run wrote no outbox row. If the channel service is
  missing or the worker cannot start, the enqueue side is detached, so the
  outbox still never becomes write-only.
- A failed task read after complete_run no longer drops the notification:
  the title is optional and the message falls back to the task id. A task
  deleted meanwhile still stays silent.

* fix(channels): park WeCom notifications while the SDK socket is closed

The WeCom SDK raises a plain RuntimeError when its socket is not open, and
send_notification only mapped ConnectionError, TimeoutError and OSError to
ChannelUnavailable. A disconnect therefore spent the counted retry budget
and settled the row as failed after about 15 minutes, although the docs
say transport outages park without consuming retries.

send_notification now asks the client's is_connected first. The new test
uses a real, never-connected SDK client: it raised RuntimeError after three
retries before, and parks now.

* docs(scheduler): describe which outcomes notify; mirror the changelog entry

- README (en/zh) and CONFIGURATION.md: only runs that finish as success or
  failed notify; occurrences that end without a finished run (launch error,
  queue timeout, restart recovery) do not. Parking lasts up to about a day.
- CHANGELOG: cite the issue as text like the rest of the file, and add the
  entry and link reference to CHANGELOG_zh.md.
- AGENTS.md: drop the backend pointer so this branch adds nothing to the
  guidance chains closest to their size limit; the scheduler-side rules
  and the lifespan order now live in app/channels/AGENTS.md.

* docs(changelog): cite #6135 next to #4843 for the notification outbox entry

* fix(scheduler,channels): re-check the IM binding at delivery time; park WeCom transport errors, count its rejections

Review round 1 on #6135.

The delivery worker now takes resolve_connections (the Gateway wires
ChannelConnectionRepository.list_connections) and, before the channel
liveness check and the summary read, requires the row's owner to still
hold a connected (provider, external_account_id). A disconnected target
is finalized with mark_failed(terminal=True); a lookup failure or an
unexpected result shape parks the row. Terminal outbox rows ignore
later status writes.

WeCom: the SDK raises RuntimeError from send_message for a socket that
is not open and for a dropped connection, but also for a platform
rejection ("Reply ack error: errcode=..."). Only that prefix keeps
counting; other RuntimeErrors, ConnectionError/TimeoutError/OSError and
websockets' ConnectionClosed park the row as ChannelUnavailable.

* fix(channels): treat a WeCom ACK rejection as proof the transport is up; pin the rejection prefix through the retry branch and the real connection-row shape

---------

Co-authored-by: GaoTianyuanhr <75788039+GaoTianyuanhr@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: GaoTianyuanhr <GaoTianyuanhr@users.noreply.github.com>
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
2026-10-01 17:59:30 +08:00
Dan CaldrandWillem Jiang 67db3d883c fix(middlewares): include file line count in read-before-write block message and clarify read_file range error (#6020)
* fix(middlewares): include file line count in read-before-write block message and clarify read_file range error

* fix(middlewares): address review feedback on line count pluralization and readability

* fix(middlewares): address round-2 review feedback for PR #6020

Stamping bug (_attach_read_mark, ggnnggez item 1):
Originally planned as a standalone issue; folded here since it is the exact mechanism behind the #6019 gate bypass. _attach_read_mark now returns early when message.content is in READ_FILE_NO_CONTENT_RESULTS - {READ_FILE_EMPTY}. In-band errors (invalid range, start_line exceeds file length, etc.) are returned with status='success' by LangChain so the existing message.status == 'error' guard did not catch them.

HTTP code collision [P3] (willem-bd):
Pre-stamp TOOL_META_KEY on the blocked ToolMessage before normalize_tool_result so files with 401, 403, 404, or 500 lines are not misclassified as auth/server errors by _classify_error_text. normalize_tool_message returns early when an existing TOOL_META_KEY entry is found.

Line-count parity (ggnnggez, inline L317):
Replace splitlines() with count_file_lines() in read_file_contract.py using count('\n'), matching LocalSandbox text-mode iteration and the _truncate_read_file_output formula. splitlines() over-counts for files with \f or \u2028 separators (e.g. pdftotext output).

Read hint (ggnnggez, inline _BLOCK_MESSAGE):
For empty files the ranged-read suggestion is omitted. For non-empty files the hint gives a concrete start_line=max(1, N-29), end_line=N.

Constant rename (ggnnggez, read_file_contract.py):
READ_FILE_INVALID_RANGE is the new canonical name; READ_FILE_EMPTY_RANGE kept as a backward-compat alias. _LEGACY_READ_FILE_EMPTY_RANGE removed (unreachable after the #6019 contract update).

Tests:
- TestBlockMessageLineCount parametrized with \f and \u2028 parity cases
- TestStampingGateBugFix: 4 no-content variants, gate-closed integration probe, empty-file-still-stamps assertion
- test_gate_block_on_numeric_line_count_is_recoverable: 401/403/404/500
- test_contract_constants_membership replaced with test_contract_constants_parity_with_localsandbox

---
NOT INCLUDED: ggnnggez item 2 (file length in read_file range error strings). READ_FILE_NO_CONTENT_RESULTS uses exact string membership and is also consumed by build_skill_usage for skill-read detection; making those strings dynamic requires a dedicated contract refactor. Deferred to a follow-up PR.

Note: test_web_fetch_relative_links failures in the full suite are pre-existing on upstream/main and unrelated to this change.

* test: clarify blank ranged reads in the write gate

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 14:40:04 +08:00
IronMurphy 5b83cb502c fix(middleware): publish externalized outputs atomically (#6109)
* fix(middleware): publish externalized outputs atomically

_externalize reported failure with None but had already created the file, so a
write that failed part-way (disk full, interrupted request) left a truncated
file under the final name in the outputs directory. Write to a sibling temp file
and os.replace it into place, removing the temp file when the write fails.

* style: drop the unused pytest import from the externalize test
2026-10-01 10:52:48 +08:00
huyengiang101086-specandhuyengiang101086-spec baae72d876 fix(channels): point the Discord missing-dependency hint at the discord extra (#6054)
* test(scripts): cover discord extra auto-detection

The discord mapping in scripts/detect_uv_extras.py worked but was neither
listed in the "currently maps" docstring nor covered by a test, unlike the
buzz mapping that grew up beside it.

* fix(channels): point the Discord missing-dependency hint at the discord extra

discord.py is the only IM transport that is not a core dependency — it ships in
the optional `discord` extra, which detect_uv_extras.py, backend/Dockerfile and
deploy.sh already expand. Its ImportError hint told users to run `uv add
discord.py` instead, which rewrites pyproject.toml/uv.lock in a tree every
restart syncs with `--locked`. Mirror the buzz hint wording and note the extra
in config.example.yaml next to the buzz block.

---------

Co-authored-by: huyengiang101086-spec <270684046+huyengiang101086-spec@users.noreply.github.com>
2026-10-01 10:51:37 +08:00
xiaodu55andxiaodu55 5b0ef628e2 fix(skills): drain custom-skill mutation tails across cancellation (#6078)
* fix(skills): drain custom-skill mutation tails across cancellation

* fix(skills): log drained skill mutation failures and drain state/install tails

Address review feedback on #6078: record cancelled-then-failed mutation
tails inside the drained unit (exception type only, per the managed-models
precedent), and extend the drained tail to the remaining write-to-cache
shapes in update_skill and the skill install path.

* test(skills): pin the lock contract under the drained cancellation timing

The cancelled update_skill caller now drains its mutation tail, so
CancelledError surfaces only after the tail settles. Start the MCP
writer without awaiting the cancelled task first, assert the section
stays owned while the drained tail is in flight, then await both and
assert the final order.

---------

Co-authored-by: xiaodu55 <xiaodu55@users.noreply.github.com>
2026-10-01 10:38:35 +08:00
Hyeonsang ChoandWillem Jiang 1f1192fd80 fix(agents): keep the todo completion reminder when a model call is retried (#6132)
* fix(agents): keep the todo completion reminder when a model call is retried

TodoMiddleware drained its queued completion reminder inside
wrap_model_call before calling the handler. LLMErrorHandlingMiddleware
sits outside it and retries a failed call by running that wrap again, so
the retry went out with an empty queue while the run had already spent
one of its two reminders on the lost one.

Drain the reminder once per wrap, and when the handler raises put it
back in front of the queue without counting it again, then re-raise. A
run whose reminder state was cleared while the call was in flight does
not get it back.

* docs(changelog): reference #6132 in the todo reminder retry entry

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 10:37:09 +08:00
Grapette.L 8df4286f41 fix(channels): split Telegram messages by UTF-16 code units, not code points (#6067)
* fix(channels): split Telegram messages by UTF-16 code units, not code points

* perf(telegram): clip stream previews without walking the whole reply

Every cumulative stream update measured the full answer and then split the full
answer to keep only its first chunk, so a long reply cost two complete passes
per update on the async channel path. _first_utf16_chunk stops at the first
character that does not fit, and the preview is built from that bounded prefix.

The tests now measure chunk length with the UTF-16 codec instead of a copy of
the production formula, cover the width helper and the bounded clip directly,
and fail if the preview stops being bounded by the limit.
2026-10-01 10:24:25 +08:00
SuperSgdkandWillem Jiang f74a290c42 fix(uploads): reject reserved staging filenames before batch ingestion (#6122)
* fix(uploads): reject filenames reserved for temporary staging

Guard thread and shelf upload names, validate SDK batches before copying, and preserve legacy shelf downloads. Add regression coverage and document the reserved staging namespace.

* fix(uploads): report reserved names before HTTP batch ingestion

Preflight reserved staging basenames before thread storage or sandbox setup. Return HTTP 400 with a rename hint; add HTTP and strict blocking-I/O regressions and update upload documentation.

* fix(uploads): reject Windows-style staging basenames on POSIX

Normalize backslashes before the HTTP reserved-name preflight. Add POSIX basename regressions for single and mixed batches and strict blocking-I/O coverage. Clarify the preflight scope and upload documentation.

Follow-up to the review of PR #6122; AI-assisted with OpenAI Codex.

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 10:22:38 +08:00
90d466c917 feat(mcp): propagate cache resets across workers (#6126)
* feat(mcp): propagate cache resets across workers

Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>

* docs: keep MCP reset guidance within instruction budgets

---------

Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
Co-authored-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 10:21:24 +08:00
Guo-Yixin 4d5607229e fix(subagents): preserve direct-return tool results (#6083)
* fix(subagents): preserve direct-return tool results

* fix(subagents): include middleware tools and surface direct failures
2026-10-01 10:05:39 +08:00
Hyeonsang Cho fe22e3899a fix(channels): keep blocking SDK teardown off the event loop in stop() (#6134)
* fix(channels): keep blocking SDK teardown off the event loop in stop()

SlackChannel.stop() called SocketModeClient.close() inline. close() joins
the SDK's message-processor thread (~0.7s every call, measured) and waits
for in-flight listeners, whose blocking Web API calls can run up to the
WebClient timeout, so every run, stream and channel on the Gateway stalled
during shutdown and POST /api/channels/{name}/restart.

Feishu and DingTalk had the same shape: they joined their SDK threads
inline with a 5s timeout, and since those threads only return on a fatal
error, the join normally froze the loop for the full 5s.

The teardown now runs through asyncio.to_thread. Slack detaches the client
first so a still-queued connect skips it, and shields and tracks the close
task so a cancelled stop() leaves it running and a retried stop() awaits
the same close instead of closing twice.

tests/blocking_io/test_channel_stop_offloop.py anchors all three stop()
paths to the strict Blockbuster gate, driving a real slack-sdk client.

* docs(changelog): reference #6134

* fix(channels): join the Discord client thread off the event loop in stop()

DiscordChannel.stop() still joined its client thread inline with a 10s
timeout. The thread usually exits right after the cross-loop client close,
but a timed-out close or a slow _run_client() drain keeps it alive for the
whole timeout, stalling the shared Gateway loop the same way the Slack,
Feishu and DingTalk teardown did. Run the join through asyncio.to_thread so
stop() matches the channels/AGENTS.md teardown rule, and add Discord to the
strict Blockbuster anchor.
2026-10-01 10:00:53 +08:00
lihongyuan99 241a7b3b9b fix(frontend): key the Julia entry of extensionMap by its extension (#6130)
extensionMap is keyed by extension and valued by language, but the Julia row
was written backwards as `julia: "jl"`. checkCodeFile tests membership with
`extension in extensionMap`, so ".jl" never matched and a Julia artifact was
not a code file at all, while ".julia" resolved to a language named "jl".

Add both `jl: "julia"` and `julia: "julia"`, the same pair form the map already
uses for haskell/hs and elixir/ex, so no extension that classified before stops
classifying.
2026-10-01 09:58:23 +08:00
Willem Jiang 5c6c074542 chore(doc): updated the CHANGLOG with lasted changes (#6121) 2026-10-01 09:38:04 +08:00
Li-john1021 90c5d526e2 fix(paths): reject Windows reserved names on host-visible paths (#6102)
* fix(paths): reject Windows reserved names on host-visible paths

Uploads, skill support files, and sandbox path checks now share one helper that rejects CON, PRN, AUX, NUL, COM1-COM9, and LPT1-LPT9 (including names like con.txt) and components that end in a dot or space. The check is not limited to Windows so a Linux host cannot store a name that breaks when the same tree is used on Windows.

* fix(paths): validate quoted local bash paths as whole words

The absolute-path scan stopped at whitespace, so a quoted upload name such as report. final.txt or CON notes.txt was rejected as a truncated segment. Local bash checks now use the shlex argument for traversal and Windows-incompatible segments.

* fix(paths): do not block reads of existing host names

* style(paths): format with ruff 0.15.12
2026-10-01 09:11:19 +08:00
creed 9c75236ea1 fix(frontend): strip closing hashes from web fetch titles (#6103)
* fix(frontend): strip closing hashes from web fetch titles

Signed-off-by: 97three <2212371308@qq.com>

* fix(frontend): scan web fetch heading suffixes in linear time

Signed-off-by: 97three <2212371308@qq.com>

---------

Signed-off-by: 97three <2212371308@qq.com>
2026-10-01 08:51:40 +08:00
ShxiaoandWillem Jiang 52b21d13ac fix(middleware): validate exact byte size on sandbox tool output externalization (#6112)
* fix(middleware): validate exact byte size on sandbox tool output externalization

* test(middleware): simulate partial write in sandbox externalization size test

* Fix PR number in CHANGELOG.md

Updated the changelog to reflect the correct issue number for a truncated file write.

* Update CHANGELOG_zh.md for issue #6112

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 08:38:08 +08:00
ee3159a01d fix(threads): re-anchor re-persisted history rows to their earliest seq (#6128)
* fix(threads): re-anchor re-persisted history rows to their earliest seq

reconcileThreadHistoryRows collapsed a repeated (run_id, identity) to
its latest-seq copy, deleting the earliest-seq row before
buildVisibleHistoryMessages could apply its earliest-seq-wins
anchoring. A message re-persisted under a newer seq therefore jumped
to the tail of the thread, contradicting the documented contract in
this function's docstring, in buildVisibleHistoryMessages, and in the
backend get_message_seqs earliest-seq-wins rule.

Re-anchor the surviving row to the earliest seq its identity held in
the merged window after deduping, and make the retained-history
stability check value-based so re-anchored rows do not force a state
update on every pass.

Fixes #6125

* fix(threads): address history reconciliation review

---------

Co-authored-by: liwenjie200543 <175601277+liwenjie200543@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 08:20:42 +08:00
4ced9dddf5 fix(vllm): fall back to legacy reasoning_content field (#6048)
* fix(vllm): fall back to legacy reasoning_content field

* fix(vllm): document the reasoning_content fallback and pin precedence

The previous change set left the docs describing the provider as
preserving only `reasoning`:

- the module docstring now mentions the legacy `reasoning_content`
  wire field, and the vLLM Provider bullet in models/AGENTS.md is
  extended to match the repo's documentation-update policy;

- _create_chat_result evaluated choice["message"] twice, once per key;
  hoist it into a single local and read both keys from it;

- two regression tests pin that a payload carrying both fields keeps
  `reasoning`, so a future refactor cannot silently invert the
  precedence (the existing tests only covered the absent-`reasoning`
  case).

Reported-by: willem-bd

* refactor(vllm): share reasoning fallback and cover follow-up payloads

---------

Co-authored-by: inchang-ing <197932532+inchang-ing@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 08:12:32 +08:00
NanPanandWillem Jiang 1414e9fae7 fix(memory): drain persistent mutations on cancellation (#6092)
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 07:59:28 +08:00
sjr666666andsjr666666 357bea8285 docs(zh): add the missing Startup Modes section to README_zh (#6129)
Mirror the English README's Startup Modes section (introduced and refined in #5995) into the Chinese version: the startup mode matrix, the stop/restart commands, and the SKIP_FRONTEND_BUILD opt-in note. README_zh previously had no documented way to run services in daemon or production mode or to stop/restart them.

Co-authored-by: sjr666666 <204130665+sjr666666@users.noreply.github.com>
2026-10-01 07:57:14 +08:00
yetugeandWillem Jiang 300abb36d2 fix(auth): reject a non-ASCII ID token nonce instead of raising (#6115)
* fix(auth): reject a non-ASCII ID token nonce instead of raising

The nonce comparison in validate_id_token still used the local
secrets.compare_digest on str operands, which raises TypeError when either
side holds non-ASCII characters. The nonce claim is provider-controlled
text, so such a provider response turned the callback's sso_failed redirect
(OIDCValidationError) into an unhandled 500.

Route it through the shared constant_time_equals helper #6076 introduced
for the same five compare sites, and drop the now-unused local helper.

Signed-off-by: yetuge <2219677952@qq.com>

* fix(auth): reject a non-string ID token nonce claim like a mismatch

A provider returning a truthy non-string nonce claim (e.g. an int)
slipped past the missing-claim guard and crashed the comparison helper
with AttributeError, which the callback turned into a 500 instead of
the sso_failed redirect. Guard the claim type before comparing so every
malformed nonce takes the same OIDCValidationError path.

Signed-off-by: yetuge <2219677952@qq.com>

---------

Signed-off-by: yetuge <2219677952@qq.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-10-01 07:43:16 +08:00
Storm 69004bd52c fix(sandbox): give fresh bash.exec runs immediate stdin EOF (#6117)
Why this change:
In Lark broker mode the shim forwards piped stdin to the broker only
after EOF. AIO's /v1/bash transport keeps a subprocess stdin pipe open
for writes, so a bare `lark-cli ...` never saw EOF and the shim blocked
forever on input that would never arrive. An earlier brace-group/eval
approach fixed the hang but changed top-level parsing (aliases/extglob)
or let a trailing backslash consume the appended delimiter.

What changed:
- _run_bash_exec prefixes the original command with `exec < /dev/null`
  on its own line, giving the fresh, per-invocation session immediate
  EOF on default stdin. Explicit pipes/heredocs/files still override
  fd0. The plain prefix keeps top-level parsing and appends no
  delimiter a trailing backslash can consume.
- The persistent /v1/shell PTY transport keeps commands unchanged; the
  shim already treats TTY stdin as no input.
- The shim drains non-TTY stdin under grace/tail idle budgets
  (DEERFLOW_LARK_BROKER_STDIN_{GRACE,TAIL}_SECONDS, default 2s, valid
  (0, 600]); an idle pipe exits 124 without contacting the broker, and
  read errors fail loudly instead of becoming empty input.
- Broker logs each exec as argc/exit/elapsed only (never argument
  values) and warns on concurrency-cap rejections.

Problem solved:
Lark commands without explicit stdin hung indefinitely in broker mode,
stalled producers could execute partial writes, and broker logs risked
capturing argument values. Docs updated: README, harness AGENTS.md,
docker/lark-cli-broker README.
2026-10-01 07:34:32 +08:00
Wenchao An 4c05e88c52 feat: add inline @ references with multiple skill activation (#6063)
* feat(frontend): unify composer references with an @ picker

* feat: edit references inline and activate multiple skills

* docs: describe inline references and multiple skill selection

* fix: prevent sent drafts from returning and repair CI checks

* fix: reconcile inline references and address review feedback

* fix: preserve references through discovery errors and polish
2026-09-30 23:09:05 +08:00
Daoyuan Li 5c67d61d65 fix(frontend): detect artifact types from file basenames (#6091) 2026-09-30 20:11:26 +08:00
ZJPex 8e282c4c24 fix: preserve existing task notes during parallel additions (#5954)
* fix: preserve task notes during parallel additions

* docs: scope task note guidance to continuity module

* fix: reserve task note slots using resolved execution keys

Signed-off-by: ZJPex <3258236335@qq.com>

* refactor: share task-note argument key and align documentation language

Signed-off-by: ZJPex <3258236335@qq.com>

* fix: ignore malformed sibling arguments in task-note reservations

Signed-off-by: ZJPex <3258236335@qq.com>

* fix: skip structurally invalid task-note reservations

Signed-off-by: ZJPex <3258236335@qq.com>

---------

Signed-off-by: ZJPex <3258236335@qq.com>
2026-09-30 20:03:35 +08:00
Fish 2000bffa8a fix(harness): anchor DeerFlowClient agent-name validation with fullmatch (#6104) 2026-09-30 19:51:47 +08:00
25671538af fix(cli): exit non-zero when a headless run ends in an LLM error fallback (#6056)
* fix(cli): exit non-zero when a headless run ends in an LLM error fallback

LLMErrorHandlingMiddleware turns provider failures (e.g. an expired
credential returning 401) into an AIMessage flagged
`deerflow_error_fallback` instead of raising. The Gateway run status and
the subagent executor already treat that flag as a failure. #5963 gave
`deerflow --print` / `--json` an error boundary for raised exceptions,
but a flagged fallback never raises, so both modes still printed the
fallback text and exited 0.

Both headless modes now follow #5963's failure contract when the run's
final AI message carries the flag: `--print` keeps the fallback text on
stdout and adds an `Error:` line on stderr; `--json` appends the terminal
`{"type": "error"}` record. Both exit 1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(cli): share final-answer accumulation with the embedded client

* docs: scope headless accumulator guidance to TUI

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-30 19:48:32 +08:00
hataaandWillem Jiang 1cd3eb742b feat(authz): enforce skill authorization at assembly and slash-activation (#4063 Phase 3) (#4541)
* feat(authz): enforce skill authorization at assembly and slash-activation (#4063 Phase 3)

Phase 3 / Skills — the second of three resource-type PRs (Models, Skills,
Sandbox). The RBAC provider already maps "skill" → config key "skills"
(rbac.py _RESOURCE_POLICY_KEYS), so no schema change is needed.

Layer 1 (assembly-time, mirrors Phase 1B):
- filter_available_skills_by_authorization() in skill_filter.py filters the
  skill-name allowlist (set[str] | None) by the provider's "skill" policy.
  In lead agent, runs after _available_skill_names() so the catalog
  (describe_skill) and SkillActivationMiddleware share one filtered set.
  In subagent executor _load_skills(), filters config.skills before disk load.
- filter_resources_by_authorization() in enforcement.py — generic version of
  filter_tools_by_authorization for any resource with a name attribute.

Layer 2 (slash-activation):
- No new middleware code needed. SkillActivationMiddleware._resolve_activation
  already checks "reference.name not in self._available_skills". Since Layer 1
  filters the allowlist, denied skills cannot be slash-activated.

authorization.enabled: false is a complete no-op. available_skills=None (no
agent allowlist) + enabled: resolves all skill names from config and filters.
12 new tests + 243 existing authz/skills tests pass.

* test(subagents): wait on the chained future instead of polling done() in caller-loop teardown test

* feat(authz): enforce action-scoped skill:activate at runtime and reuse the catalog loader for filter candidates

* feat(authz): arm skill:activate on the subagent chain, unify the client skill-surface user, and record the Layer 2 reversal in the design notes

* chore(authz): align the subagent filter's fallback user with the chain's bucket convention

* feat(authz): gate autonomous skill loading on skill:activate, close the bootstrap assembly leak, and reuse one provider per subagent assembly

Review round 4 (willem-bd, CHANGES_REQUESTED):

- [P1] Autonomous load path: describe_skill filters its match set on the
  action-scoped skill:activate decision (all-denied is indistinguishable
  from "no match"), and ToolErrorHandlingMiddleware stamps skill_context
  entries only when the action allows it — a denied SKILL.md read gets a
  skill_context_denied marker instead, so durable context, allowed-tools
  policy, and autonomous secret bindings never activate a denied skill.
  File visibility itself stays Layer 1 (sandbox-projection semantics).
- [P1] Bootstrap assembly: build_middlewares/apply_prompt_template now
  receive the authorization-filtered set instead of the raw
  _BOOTSTRAP_SKILL_NAMES, and an empty skill_names index is preserved
  when deferred discovery is on rather than round-tripping through None
  into the legacy full-metadata prompt that re-advertises a denied
  skill (deferred-off keeps the legacy rendering; the filtered empty
  allowlist blanks it on denial).
- [P2] SubagentExecutor resolves ResolvedSkillAuthorization once per
  assembly (_resolve_skill_authorization, memoized per executor) and
  shares it across the Layer 1 filter (authorization=), the activation
  middleware, and both new gates — uncached provider factories can no
  longer hand the layers different instances/policy snapshots.

All gates share one helper (skill_filter.skill_activation_allowed) so
slash, describe, and skill-file-load cannot drift; provider errors keep
the configured fail-closed/fail-open policy. skill_authorization is
threaded through _build_runtime_middlewares into both runtime builders.

Tests: describe gating (deny/fail-closed/fail-open/disabled), read
stamping (denied marker vs entry) + extract_skills skip without warning,
bootstrap three-way param (allow/deny x deferred on/off), lead + subagent
chain wiring for the stamp gate, and the distinct-instance provider
factory repro. Each fix is mutation-verified (reverting it fails its
test). Full A/B suite run: no new failures vs HEAD.

* feat(authz): re-authorize persisted skill_context before reuse and route async paths through aauthorize

Review round 5 (willem-bd):

- [P1] Persisted entries: skill_context survives across runs, so both
  consumers now re-check the action-scoped skill:activate decision before
  reuse. SkillToolPolicyMiddleware._active_skills_for_paths skips a
  denied skill (same skip as a disabled/unallowlisted one; all-denied
  keeps the existing fail-closed-to-builtins semantics) and
  _in_context_secret_sources binds no secrets for a denied entry.
  skill_authorization is threaded into SkillToolPolicyMiddleware on both
  chains.
- [P2] Async provider API: skill_activation_allowed_async mirrors the sync
  helper through await provider.aauthorize() (same pattern as
  authorize_sandbox_execution_async). Async hooks precompute an
  activation_decisions map on the event loop (slash target + persisted
  entry names, one batch of awaits) and hand it to the worker-thread
  handler, which consults the map instead of the sync API:
  SkillActivationMiddleware.awrap_model_call (activation + secret
  binding), ToolErrorHandlingMiddleware.awrap_tool_call (_amaybe_stamp
  awaits the decision before stamping), and SkillToolPolicyMiddleware's
  awrap_model_call/awrap_tool_call (entry names + slash source name).
  describe_skill becomes a StructuredTool with both a sync function
  (invoke -> authorize) and a coroutine (ainvoke -> aauthorize), so async
  agent execution never routes a loop-affine provider through the sync
  API while sync callers keep their semantics.

Tests: allow-read-then-deny-next-run regressions for secret binding and
allowed-tools (partial deny does not poison, all-denied fails closed to
builtins), and an _AsyncOnlyProvider (sync authorize always fails,
aauthorize works) driving four async-path assertions (activation, stamp
allow/deny, describe, tool policy). All six fixes are mutation-verified;
full A/B suite run against HEAD shows no new failures.

* fix(authz): authorize skill activation by declared Skill.name, not directory basename

ShenAC-SAC's P2: bundled skills may declare a name that differs from their
directory (skills/public/vercel-deploy-claimable declares name: vercel-deploy),
but the runtime gates derived authorization targets from the container path.

- New deerflow/skills/container_registry.py: build_container_path_registry()
  + canonical_skill_name() — the one path-to-Skill.name resolution shared by
  the activation, tool-policy, and read-stamping middlewares.
- ToolErrorHandlingMiddleware resolves the read path through the (user-scoped,
  via new user_id kwarg threaded from both runtime builders) registry before
  the skill:activate decision, sync and async (registry lookup in a worker
  thread, aauthorize() on the loop). RBAC allow by the declared name now
  activates; allow by the directory name does not.
- SkillToolPolicyMiddleware async hooks canonicalize the policy paths first
  (three-phase: thread registry -> loop aauthorize -> thread filter reusing
  the registry; the cached per-model-step decision still short-circuits),
  so the decision map is keyed by the names _active_skills_for_paths checks —
  no more map miss falling back to the sync provider API from the thread.
  _entry_names removed: persisted entry["name"] is path-derived by
  construction; entries are resolved by path instead.
- SkillActivationMiddleware collects slash names (already canonical) plus
  persisted entry *paths*, canonicalized off-loop via
  _canonical_names_for_paths().

Regressions (test_skills_authorization.py 45->50): sync stamp target, RBAC
allow declared-name/deny directory-name both directions, async stamp via
_AsyncOnlyProvider, composed async slash -> tool-policy keeps the skill's
tools with zero sync fallback, persisted-entry secret binding survives the
name mismatch. All three mutations (stamp target, policy map key, entry
canonicalization) fail their tests.

* fix(authz): close the sync-fallback class on skill activation paths

willem-bd's two P2s on the canonical-name fix, plus a fourth instance of
the same class found while sweeping the membership:

- SkillToolPolicyMiddleware: a transient prepass registry failure is now
  preserved (_REGISTRY_LOAD_FAILED marker) — the worker-side filter applies
  its own fail-closed treatment instead of silently retrying storage,
  which recovered skills absent from the decision map and fell back to
  the synchronous provider API from the thread (fail-open retaining a
  denied skill's allowed-tools).
- SkillActivationMiddleware: an empty entry-path list no longer pays a
  full skill-tree scan on ordinary authorization-enabled async model
  steps (previously a passive path with zero registry I/O).
- 4th instance (reproduced with the PR fixtures, then fixed): the entry
  prepass failure on the secret-binding path let the consumer-side
  _load_skill_registry_by_path silently recover and resolve entry names
  outside the decision map — the same sync fallback, fail-open binding
  the aauthorize-denied skill's declared secret. The fix goes one step
  further than failure propagation: the prepass hands down the registry
  *snapshot* the decision map was keyed by, and entry sources resolve
  against that snapshot, so a map miss is impossible by construction
  (this also closes the narrower window where both loads succeed but
  disagree because storage changed mid-step). The slash source is
  deliberately unaffected: it does not consult the decision map
  (validated once at activation) and keeps its fresh per-call load.
- Class membership swept: skill_activation_allowed() sync call sites are
  exactly four (stamping sync path and describe's sync implementation are
  legitimately sync); the only two async-reachable fallback points (the
  _activation_allowed helpers) are both construction-guarded now.
  Structural hardening (typed async-batch decision object where a map
  miss is fail-closed by construction) recorded in the plan notes as a
  follow-up, anchored at those two helpers.
- _RegistryArg type alias replaces the degenerate dict|object|None union
  in the policy middleware signatures.

Tests (test_skills_authorization.py 50->54): transient prepass failure ->
builtins-only with exactly one storage attempt and zero provider calls;
ordinary async step -> zero registry scans; two-call scenario -> entry
binding zeroed while the slash binding survives; entries resolve against
the prepass snapshot under a mid-step rename. Mutations verified red:
marker reverted to None, over-broad slash kill, snapshot bypass.

* fix(authz): correct secret-binding reference points and gate the rendered skill reminder

Review round 9 (willem-bd) plus construction-audit follow-ups on the same
class — every async divergence between two loads of the same data now has
an explicit, pinned reference point:

- Slash secret source resolves from a FRESH post-activation registry: the
  prepass snapshot is taken before the aauthorize awaits, so binding the
  slash source from it would activate NEW_KEY content while injecting
  OLD_KEY when a skill's declared secrets change in that window. The
  'one scan per step' bonus that motivated sharing the snapshot is
  withdrawn — consistency wins. The two secret sources resolve
  independently: a transient fresh-load failure zeroes only the slash
  binding; entries keep binding from their own snapshot.
- Same-skill dominance: when a skill is both the slash source (fresh) and
  a persisted entry (snapshot), only the activation-era declaration binds
  — unioning both eras injected a secret the skill no longer declares.
  Aligns with the tool-policy rule (slash activation dominates for the
  run); different-skill entries still stack additively.
- The rendered 'Active skills' reminder (DurableContextMiddleware) no
  longer advertises skills the provider denies — surface consistency with
  describe_skill. Publication/consumption split: the activation
  middleware publishes per-step path-keyed skill:activate decisions into
  the run context (async: reused from the prepass snapshot, zero extra
  I/O; sync: one scan + authorize() only when entries exist — passive
  steps stay scan-free), the renderer consumes them and filters its
  render copy (state untouched; unpublished = permissive, matching the
  other run-context carriers). New carrier key listed in
  REDACTED_CONTEXT_KEYS with a redaction regression.
- Wiring regressions on both chains for everything added above (user_id
  to the stamp gate, skill_authorization + token pairing + activation-
  before-durable ordering to the durable renderer): each mutation
  verified red (M4-M13 across this change set).
- Tests 54->65: slash post-activation lookup, source independence,
  dominance, sync-chain sync-API path, render filter (async/sync/
  disabled with load-count assertions), chain wiring x2, redaction
  listing. Plan notes record rounds 9-2..6 including the reference-point
  post-mortem and the audit axes exhausted.

* fix(authz): anchor same-skill dominance on the slash identity, not binding results

Review round 10 (willem-bd) plus a self-audit extension the new checklist
items caught in the first fix itself:

- R10 (P2): the same-skill exclusion was derived from successfully BOUND
  slash sources. When the prepass snapshots required-secrets [OLD_KEY]
  and the operator strips the declaration during the await window, the
  slash source binds nothing (required_secrets gate) — the exclusion set
  stayed empty and the stale entry view kept injecting OLD_KEY for a
  skill that now declares no secrets. The exclusion now derives from the
  AUTHENTICATED slash activation identity (run-context source path,
  owner-token checked): the identity's declared name resolves from the
  fresh registry first, then the entry snapshot, and joins the exclusion
  set regardless of binding success.
- Self-audit (value->gate->downstream on the fix itself): the declared
  name can also change mid-run (same path, foo->bar). A name-anchored
  exclusion then holds only the new name while the snapshot entry
  resolves the same path to the old one — both eras bound again. The
  PATH is the stable identity shared by the activation and every earlier
  read: entries at the slash identity's path are now excluded before
  registry resolution (exclude_paths), independent of what either
  version calls the skill; the name-based set remains for cross-path
  same-name shadowing.

Tests 65->67: no-secrets-era dominance (ACTIVE secrets None under the
OLD_KEY-to-empty interleaving), path-anchored dominance across a mid-run
rename (only the activation-era key binds). Mutations verified red:
identity-derived exclusion disabled (M14), path exclusion disabled
(M16).

* test(authz): pass explicit config to lead middleware wiring test

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-30 19:45:07 +08:00
yetuge 8432e4b71a fix(gateway): drain agent and user-profile writes across cancellation (#6087)
* fix(gateway): drain agent and user-profile writes across cancellation

A client that disconnects mid-request cancels the handler task. On the four
persistent-write endpoints of /api/agents - create, update, delete, and the
USER.md write - a bare asyncio.to_thread then either cancels a still-queued
worker (the write silently never happens) or detaches from a running one and
drops its failure, because the handler's except never runs.

Route the four writes through await_drained like the managed-subagent (#6023),
managed-model (#6024), and custom-skill (#6078) mutations, logging a lost
worker failure with the exception type only before the drain consumes it.
Expected domain errors (AgentExistsError on create) re-raise unlogged; the
caller still maps them to a 409 while connected. Reads stay bare: abandoning
them loses nothing.

Signed-off-by: yetuge <2219677952@qq.com>

* fix(gateway): address review nits on the agents write drain

Make expected_errors a positional-only parameter before *args so the
ParamSpec signature is checker-clean under PEP 612 (the one call site with
extra worker arguments passes () explicitly), acknowledge that
non-cancelled failures log twice like artifacts.py's drained commit does,
and drop the unused request stub helper from the tests.

Signed-off-by: yetuge <2219677952@qq.com>

---------

Signed-off-by: yetuge <2219677952@qq.com>
2026-09-30 19:43:01 +08:00
NanPanandWillem Jiang 4343ced685 fix(mcp): drain personal config mutations on cancellation (#6093)
* fix(mcp): drain personal config mutations on cancellation

* test(mcp): cover cancelled personal config mutation

* test(mcp): remove unused cancellation event

* test: remove unused event loop from personal MCP cancellation regression

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-30 19:39:11 +08:00
Grapette.L f61372764e fix(agents): anchor agent-name validation against a trailing newline (#5960)
* fix(agents): anchor agent-name validation against a trailing newline

Two copies of the shared ^[A-Za-z0-9-]+$ grammar still used re.match,
whose $ also matches before a final newline, while every other copy uses
fullmatch. POST /api/agents {"name": "reviewer\n"} therefore passed the
router guard and died inside the strict file store as an HTTP 500, and
DeerMem accepted the same name as an agents/{name}/facts directory.

* test(agents): state which trailing-newline params the regression actually covered

.match only accepted "reviewer\n" on main; "reviewer\n\n" and "reviewer \n"
were already rejected, so the docstrings should not read as if every param in
these parametrizations was broken before the fix.
2026-09-30 19:38:09 +08:00
a83aebe702 fix(docs): harness docs document an async client API that does not exist (#5843)
* docs: replace nonexistent async client API in harness docs

DeerFlowClient has no async methods: client.astream() and client.ainvoke()
do not exist, so the "Create Your First Harness" tutorial and the harness
integration guide fail with AttributeError on their first call. Overrides
were also shown as a nested config={"configurable": {...}} dict, which
stream()/chat() swallow via **kwargs and silently ignore.

Switch the examples to the shipping API (sync stream()/chat(), flat kwarg
overrides, agent_name as a constructor arg) in the EN and ZH copies of both
pages. backend/docs/STREAMING.md is left alone: its astream mentions are
design discussion and LangGraph internals, not user-facing examples.

Assisted-by: Claude Code / claude-opus-5
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>

* docs: serialize StreamEvent before StreamingResponse, use real event names

Review follow-up on the two points raised in #5843.

P2 — the streaming examples handed a generator of `StreamEvent` dataclasses
straight to Starlette. `StreamingResponse` calls `.encode()` on any chunk that
is not `str`/`bytes`, so the first event raised
`AttributeError: 'StreamEvent' object has no attribute 'encode'`. Both
occurrences (the `run_agent` snippet and the inline FastAPI route, which shipped
the dataclass `repr()` through an f-string) now serialize to SSE frames built
from `event.type` / `event.data`.

P3 — the tutorial showed mapping literals with `messages` and `thread_state`,
neither of which exists. The real contract is
`StreamEventType = Literal["values", "messages-tuple", "custom", "end"]`
(backend/packages/harness/deerflow/client.py:128) and the stream yields
dataclasses, so the section now shows `StreamEvent(...)` with the real names
plus a consuming loop.

The zh tutorial had no equivalent section at all; added for parity.

Assisted-by: Claude Code / claude-opus-5
Machine: A-Mac16-2019-PaloAlto
Account: tonydzi
Operator: robot:pr-reply-daily

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: drop the nonexistent load_config() from every harness snippet

Review follow-up on #5843 (@willem-bd): the rewritten tutorial still opened
with `from deerflow.config import load_config`, which raises ImportError
before `client.stream(...)` is ever reached — `deerflow/config/__init__.py`
exports `get_app_config` and friends, and the string `load_config` does not
appear anywhere under `backend/packages/harness/deerflow/` (the only near
match is `load_configured_extension_middlewares`).

The call was also unnecessary: `DeerFlowClient.__init__` already resolves
configuration itself (client.py:220-222 — `reload_app_config(config_path)`
when a path is given, otherwise `get_app_config()`, which auto-loads from
the default path or DEER_FLOW_CONFIG_PATH).

Fixed the class rather than the one snippet you flagged: the same dead
import appeared 10 times across 4 pages (EN + ZH tutorial and integration
guide). Every one is gone. The "configuration in embedded mode" section now
shows the API that exists — `DeerFlowClient(config_path="...")` — instead of
`load_config(config_path="...")`.

Also from the review:
- the `end` example showed only `total_tokens`; `stream()` always yields
  input/output/total (client.py:920), and the snippet right below prints the
  whole usage dict, so a reader saw keys the docs omitted. Both EN and ZH now
  show the full payload, and the `values` example visibly elides `artifacts`
  and `summary_text` instead of pretending they are absent.
- the two SSE comments on the ZH integration page stayed English while the
  example above them was translated; translated.

Verified statically, not executed: `grep` for the symbol across the harness
package (zero hits) and a read of `DeerFlowClient.__init__`. I did not stand
up a backend to run the tutorial end to end.

Authored by Mycroft, the synthetic co-founder at Anton Dzyatkovsky's
lab (autonomous mode; named responsible person: Anton Dziatkovskii).

Assisted-by: Claude Code / Opus 5 (review response, implementation)
Machine: MacBook-Anton
Account: a
Operator: mycroft (autonomous routine github-thread-watch)
Signed-off-by: tonydzi <194927794+tonydzi@users.noreply.github.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: fix native artifact SSE encoding and custom config advice

---------

Signed-off-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>
Signed-off-by: tonydzi <194927794+tonydzi@users.noreply.github.com>
Co-authored-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-30 19:36:54 +08:00
Grapette.L 0110845f66 fix(config): guard the recovered stream cleanup delay like its heartbeat sibling (#5850) 2026-09-30 19:26:50 +08:00
TotoroandTotoro-qaq 295437692a fix(agent,goal): stop coaching scripts when no bash tool is bound, and show the goal evaluator Human Input Card answers (#6082)
* fix(agent): drop script guidance from the lead prompt when no bash tool is bound

The workspace section told the agent how to write scripts and commands
("prefer relative paths", "avoid hardcoding /mnt/user-data/... inside
generated scripts") even when no tool could run them. With the default
LocalSandboxProvider host bash is off and the bash tool is removed, so
agents wrote helper scripts nothing could execute, then stalled or
computed by hand.

apply_prompt_template() takes bash_available (default True, so other
callers keep today's text). The lead agent (both assembly branches) and
the embedded client pass whether their authorized tools, bound or
deferred, include bash. Without it:
- the two script bullets become one line: no bash tool is bound, work
  out results directly and write them with write_file;
- the subagent section no longer shows direct bash examples, and marks
  the bash subagent unavailable, since it inherits the lead's tool groups;
- the ACP hint no longer suggests `bash cp`.

* fix(goal): show the goal evaluator the user's answer to a Human Input Card

The web UI sends a Human Input Card answer as a hidden human message
(hide_from_ui with a human_input_response) and shows it on the card.
format_visible_conversation() skipped every hidden message, so the
evaluator saw the question but never the answer. Following the user's
answer then read as the assistant guessing, and the goal stood down with
needs_user_input although the user had answered.

A hidden human message whose human_input_response passes
read_human_input_response() now adds a line
"User (Human Input Card answer): <value>" from the structured value; the
question is already on the card's own line. Other hidden messages, such
as goal continuations, stay out. Over the 12,000-char cap the latest
answer the cut drops is kept at the top next to the latest request, and
is never taken for the request itself. Without card answers the cut is
unchanged.

* fix(agent): keep the subagent examples and bash hint in line with the bound tools

Review follow-up for the lead prompt without bash.

- The last delegation example follows the tools. With a bash subagent on
  offer it is unchanged. A lead with bash and no bash subagent keeps
  "Run a routine test, build, or git command directly." without the Bash
  subagent sentence. A lead without bash gets a file-work example instead.
- The unavailable-bash description no longer points at the sandbox
  provider. Since #2253 it renders only when the bash subagent is
  registered and the lead binds no bash, where the sandbox is not the cause.

---------

Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
2026-09-30 17:54:42 +08:00
Hyeonsang ChoandWillem Jiang 14cd840bbf fix(gateway): reject non-ASCII tokens instead of raising in constant-time compares (#6076)
* fix(gateway): reject non-ASCII tokens instead of raising in constant-time compares

hmac.compare_digest raises TypeError when a str operand holds non-ASCII
characters. Starlette decodes header bytes as latin-1 and query values as
UTF-8, so a single 0xE9 byte in a CSRF token, GitHub webhook signature,
internal auth token, or OIDC state turned the expected 403/401 into a 500.

Route the five client-facing comparisons through a new
app.gateway.utils.constant_time_equals helper that compares the UTF-8
encodings ("surrogatepass" so even a lone surrogate cannot raise). ASCII
results are unchanged, and no bypass was possible before the fix.

* fix(provisioner): reject non-ASCII API keys instead of raising; drive OIDC test over HTTP

The standalone provisioner compared X-API-Key with secrets.compare_digest on
str operands, so a non-ASCII header raised TypeError (500) the same way the
Gateway sites did. It cannot import app.gateway, so encode inline.

The OIDC state regression test now mounts the auth router and sends the
percent-encoded state through a real signed state cookie instead of stubbing
get_state_cookie and calling the handler directly.

* docs(changelog): reference #6076

---------

Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
2026-09-30 17:41:11 +08:00
Wenchao An 1e462ae4a2 fix(frontend): consolidate run duration into reasoning headers (#6057)
* fix(frontend): show run duration in reasoning headers

* fix(frontend): cover all reasoning duration disclosures

* fix(frontend): collapse restored reasoning and cache duration targets

* chore: remove PR-only reasoning screenshot
2026-09-30 17:22:41 +08:00
kuseandliwenjie200543 8ecff3be95 test(config): make refresh_skew rejection test load-bearing (#6080)
test_refresh_skew_rejects_negative_and_boolean omitted the required
token_url, so McpOAuthConfig(refresh_skew_seconds=-1) failed validation
on the missing field alone -- deleting the ge=0 bound and the boolean
validator would leave the test green. Pass token_url (the same URL the
refresh_skew_seconds=0 boundary test uses) so only the guards can
satisfy pytest.raises; verified by stripping both guards from
extensions_config.py, which now fails exactly the two parametrized
cases, and restoring them, which returns the suite to 42 passed.

Also move SandboxOwnershipConfig up beside the other config imports
and drop the redundant in-test `import pytest` (review nits on #6026).

Generated-by: ZCode (GLM, coding agent)

Co-authored-by: liwenjie200543 <liwenjie200543@users.noreply.github.com>
2026-09-30 17:21:43 +08:00
PeaceMaker-bestandPeaceMaker-best b73a4e79bf feat(memory): scope management API by agent (#5565)
* feat(memory): scope management API by agent

Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>

* fix(memory): preserve scoped management isolation

Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>

---------

Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
Co-authored-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
2026-09-30 16:50:07 +08:00
Hyeonsang Cho af67738521 fix(agents): treat a null max_total_subagents override as unset (#6088)
* fix(agents): treat a null max_total_subagents override as unset

An API caller sending "max_total_subagents": null in the run context
failed that run with a TypeError: the key was present, so
dict.get(key, default) returned None and building SubagentLimitMiddleware
compared None with an int.

Resolve the per-run delegation cap through one helper,
effective_total_subagents_per_run, which treats None as "use
subagents.max_total_per_run" and clamps to 1-50. The lead agent's
middleware and assembly paths, DeerFlowClient, and the system prompt all
use it, so the extension host policy and release policy now report the
enforced cap instead of an out-of-range request.

* docs(changelog): reference #6088 in the null max_total_subagents entry
2026-09-30 16:10:29 +08:00
xihongshichaojidan8 d0e4d3525c fix(frontend): show media icons for APNG, AVIF and WebM (#6089) 2026-09-30 15:58:45 +08:00