* test(extensions): make the extension-manager suite runnable on Windows hosts and under host uv mirrors
- use the file's existing platform-aware .venv interpreter idiom in the two
rollback tests (Scripts/python.exe on nt);
- compare the uv add source argument as a Path so the assertion is
separator-neutral;
- pin UV_DEFAULT_INDEX in an autouse fixture so host mirror configuration
cannot break the source-build tests; tests installing their own local
index keep precedence.
Fixes#6141
* test(extensions): drop the extra-index env channels from the suite environment
UV_INDEX and UV_EXTRA_INDEX_URL entries take priority over the pinned
default index and could still break the source-build tests. UV_NO_CONFIG
is intentionally not used: it would also drop the fixture's own
[tool.uv] pyproject settings the command-shape assertions rely on.
---------
Co-authored-by: xiaodu55 <xiaodu55@users.noreply.github.com>
* refactor(gateway): hoist the persistent-write drain into one shared helper
Each persistent-write surface forked its own per-router drain wrapper:
_drained_write in routers/agents.py (action label, expected_errors
pass-through, lost-failure type-only logging), _run_store_mutation in
routers/subagents.py (bare drain), and the inline await_drained call in
routers/managed_models.py. Hoist the superset into
app.gateway.persistent_writes.run_drained_write and point all three
routers at it, per the review suggestion on #6087 (its family
predecessors #6078/#6092/#6093 have since landed).
The subagents and managed-models call sites now also emit the type-only
lost-failure log line on a drained-then-failed write (previously
silent), matching the tradeoff acknowledged in the #6087 review.
artifacts.py and personal_mcp.py keep their direct await_drained forms:
the former drains a coroutine with its own rollback flow, the latter
needs no label or expected-error handling.
Router tests move their monkeypatch targets to the shared name; the
lost-failure assertion pins the helper's logger. AGENTS.md points the
next persistent-write surface at the canonical implementation.
Signed-off-by: yetuge <2219677952@qq.com>
* fix(gateway): quiet expected domain errors and scope the AGENTS.md drain claim
Review follow-ups on #6151:
- The subagents create/update call sites pass (ManagedSubagentExistsError,)
and (FileNotFoundError,) so the shared helper stays silent on routine
4xx outcomes (duplicate-name 409, stale-store 404 race), matching
create_agent's (AgentExistsError,).
- The managed-models save passes (HTTPException,), so the helper stops
double-logging failures that _save_with_failure_logging already
reports and stays silent on its routine 4xx outcomes.
- The AGENTS.md sentence now names the three converted routers instead
of implying every persistent-write surface uses the helper; the
shortened form also keeps the file under the guidance byte budget
that agent-guidance enforces.
Signed-off-by: yetuge <2219677952@qq.com>
---------
Signed-off-by: yetuge <2219677952@qq.com>
* feat(authz): extension tool authorization PR2 — middleware-declared tool path
Layer 1 now covers tools declared by middlewares, which LangChain merges
into the bound tool set after the host's explicit-list filter:
- new agents/middlewares/tool_declarations.py: record the ordinary pass's
verdict, collect declarations from the built stack, decide only the
delta seeded by the first pass (a name denied for this build stays
denied), narrow the stack build-locally on independent state-preserving
copies (direct __dict__ write, refused up front for tools data
descriptors, never mutating caller-owned instances), and verify the
exact bound stack after the final normalization copy.
- IsolatedMiddleware ships a state-preserving __copy__ so narrowing
survives full/delta normalization and copies of copies.
- TodoMiddleware degrades when write_todos is denied: no system-prompt
injection, no completion/context-loss reminders, no jump_to.
- Layer 2 forwards host-resolved tool provenance into AuthzRequest.context
and binds the infrastructure exemption to the concrete host-created
tool object instead of a name set (tool=None goes to the provider).
- Wired into all four assembly paths (lead x2, native subagent with the
decision offloaded via asyncio.to_thread, DeerFlowClient).
Tests: new test_middleware_declared_tools.py (35) and
test_layer2_tool_provenance.py (7); blocking-IO anchor pinning the
subagent offload (teeth verified red->green); placement-guarantee and
enforcement suites extended; executor test doubles converted for the
async _create_agent. Behavior change called out in CHANGELOG: with
authorization enabled and a tools policy that excludes write_todos,
plan-mode builds no longer bind it (built-in RBAC default unchanged).
* fix(authz): close residual gaps from review — narrow non-collectable declarations, dedupe candidates, warn on skipped pass
- apply_declared_tool_view now removes non-BaseTool declarations from the
bound view when authorization is enabled (they would otherwise bind via
LangChain's create_tool auto-conversion with no provider decision);
verify_declared_tool_view fails loudly if one survives normalization
- decide_declared_tools dedupes never-submitted names so the provider sees
each candidate once
- subagent executor logs a warning when the provider is set but the seeded
Layer-1 state is missing, instead of skipping the declaration pass silently
- docs: middleware-declared tools must be BaseTool instances with usable names
* fix(authz): fail loudly on non-list/tuple middleware tools declarations
_iter_declared warned and returned () for a 'tools' attribute that is not
a list/tuple, but LangChain's factory iterates the attribute as-is — so a
set/generator/dict_values declaration bound with no Layer-1 decision.
Raise DeclaredToolViewError instead, matching the module's fail-loud
standard for every other un-narrowable shape. Disabled path remains a
strict no-op. Docs updated to match.
* feat(persistence): add notification delivery outbox for scheduled task IM push
New notification_deliveries table and repository (migration 0012) backing
the outbox pattern for issue #4254: idempotent enqueue keyed on task_run_id
plus event plus provider plus target, atomic claim of due rows, and
exponential-backoff retry state moving pending to sending to sent or failed.
Execution status stays in scheduled_task_runs; delivery status lives here.
* feat(scheduler): enqueue and deliver scheduled-task notifications over IM channels
Closes the loop for issue #4254. On run completion the scheduler writes
outbox rows for the task owner's connected channel identities (success and
failure events; interrupts stay silent), and a new NotificationDeliveryWorker
polls due rows and pushes them through the channel's proactive send path.
Channel gains send_notification; WeCom implements it via the WS client's
send_message with errcode checking, the base class raises for providers
without proactive push.
Delivery is best-effort and never shadows execution status: enqueue and send
failures are logged and recorded in the outbox, which retries with backoff.
Channel liveness is re-checked at delivery time, and the worker stops before
the channel service on shutdown. Feature activates only when channel
connections are enabled; no new dependencies.
Verified end to end against a live WeCom bot: scheduled run completion
produced an outbox row that was delivered in under one second.
* feat(scheduler): include scheduled run result summary in IM notifications
The v1 notification only carried the run outcome skeleton (status, task id,
run id). Completed runs now include the agent's final answer so the IM push
is actually useful headlessly.
- The delivery worker accepts an optional resolve_run_summary(run_id,
user_id) resolver; for run_completed events it enriches the rendered text
with a "Result:" line (bounded to 1000 chars). The lookup is best-effort:
resolver failures or missing summaries degrade to the skeleton text and
never block or fail the delivery.
- The summary is read at delivery time from runs.last_ai_message (the
convenience column written on run completion), not snapshotted into the
outbox payload; retried deliveries re-read the freshest value and the
outbox row payload stays skeleton-only. The claimed outbox row is never
mutated in place.
- Gateway wiring scopes the read to the outbox row's owner_user_id so a
stale row cannot pull another user's run content.
- Failed runs keep the error-only skeleton: a partial answer on a failed
run would be misleading.
Tests: 9 new worker cases (rendering incl. truncation, sync/async
resolvers, failure/missing fallbacks, no resolver call for failed runs).
Full scheduled + outbox + channels regression green (441 passed), ruff
check + format clean.
* fix(scheduler): address review feedback on notification delivery reliability
Review follow-ups for the scheduled-task IM notification outbox:
1. Crash recovery for "sending" rows: claim now stamps updated_at and the
worker reconciles rows stuck in "sending" past a 10-minute lease back to
"pending" before every claim. Per-row try/except isolates a crashing
delivery from the rest of the claimed batch.
2. Double-send race under multiple workers: claim returns exactly the rows
its own UPDATE...RETURNING flipped (guarded by status='pending'), so a
rival worker's concurrent claim can no longer be handed out twice.
3. Restore non-ASCII comments and drop the BOM accidentally introduced in
test_scheduled_task_service.py.
4. Manual "run now" completions no longer enqueue IM notifications; the
user is already watching the UI. Missing trigger metadata fails safe
towards notifying.
5. Channel-outage failures no longer consume the retry budget: they are
parked with count_attempt=False on a flat 15-minute backoff so an
hours-long outage cannot exhaust retries and drop the notification.
Tests: repository 16, worker 21, scheduled-task service 31; targeted
regression across notification/scheduler/channel suites: 511 passed.
* fix(scheduler): address round-2 review on IM notification delivery
Park WeCom transport/outage failures without burning retries, bound worker stop, include task titles, align enqueue with delivery, and document the IM push behavior.
Co-authored-by: Cursor <cursoragent@cursor.com>
* docs(scheduler): complete IM notification docs and drop unused list API
Add backend/channels/migrations AGENTS notes, CHANGELOG and config.example hints for the outbox push path, and remove the unused list_by_task_run repository method flagged in review.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(scheduler): address round-3 review on IM notification delivery
Cap channel-down parking by row age and parked attempts, redact IM egress
content, narrow WeCom transport exceptions, add lifespan wiring tests, and
run ruff format so lint-backend passes.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(scheduler): redact entire PEM blocks in IM egress text
The previous BEGIN-only pattern left key material and the END marker in notifications.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(persistence): align 0014 parked_attempts with bootstrap schema tests
Backfill existing rows with a temporary server default, then drop it so create_all matches Alembic, and move HEAD anchors to 0014.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(persistence): format notification tests and shorten alembic revision id
ruff format --check failed on test_notification_delivery_worker.py.
0015_notification_delivery_parked_attempts exceeds alembic_version.version_num
(VARCHAR 32) and truncates on Postgres CI; use 0015_parked_attempts instead.
* test: fix post-merge CI failures after migration renumber
Update migration 0015 test HEAD expectation to 0017_parked_attempts and add queue_timeout_seconds to the lifespan notification worker mock scheduler config.
* fix(persistence): chain the notification outbox migrations after 0026
Main's head moved from 0016 to 0026_mcp_task_lease_tokens, so the two
outbox revisions are renumbered and re-parented:
- 0017_notification_deliveries -> 0027_notification_deliveries
- 0018_parked_attempts -> 0028_parked_attempts
Only the revision ids, parents and the migration guide change.
* fix(scheduler): adapt the notification hook and its docs to main
- tests: the completion tests assert on the atomic complete_run record;
new cases cover a completion that was not recorded, a task deleted
meanwhile, a failed task read, and the outbox being off.
- AGENTS.md: drop the root note and shorten the backend note to a pointer;
the two notes pushed two effective guidance chains past the 96 KiB limit.
- CHANGELOG: move the entry from the released 2.1.0 section to Unreleased
and add its link references.
* test(persistence): move the chain-head pin to the outbox migrations
test_migration_0026 pinned 0026 as the head, so it failed once the outbox
revisions chained after it. It now checks chain membership like the older
migration tests, and a new test file owns the head pin, checks the parents
and the 32-character revision id limit, and runs the upgrade and downgrade.
* fix(gateway): wire scheduled-task notifications after the channel service starts
The outbox wiring sat in the scheduler block and asked for the running
channel service there. #5035 moved the channel service start to after the
scheduler, and the two changes merged without a conflict, so on every
startup the lookup returned None and the feature disabled itself with the
"write-only outbox" warning. The lifespan tests patched the lookup to a
constant and kept passing.
The wiring now runs in its own step after the channel service start. The
scheduler is built as on main, and ScheduledTaskService gains
attach_notification_outbox() so the enqueue side is enabled once the
delivery worker is running. The guard, the worker and the shutdown order
are unchanged.
Tests: the lifespan helper can model the real accessor (None until the
channel service has started); a new case fails without this fix.
* fix(scheduler): enqueue from scheduler start and keep the push when the title read fails
Follow-up to the startup-order fix, from review.
- The enqueue side is wired with the scheduler again (it only needs the
durable table), and only the delivery worker waits for the channel
service. Attaching it later left a window after the scheduler started in
which a finished run wrote no outbox row. If the channel service is
missing or the worker cannot start, the enqueue side is detached, so the
outbox still never becomes write-only.
- A failed task read after complete_run no longer drops the notification:
the title is optional and the message falls back to the task id. A task
deleted meanwhile still stays silent.
* fix(channels): park WeCom notifications while the SDK socket is closed
The WeCom SDK raises a plain RuntimeError when its socket is not open, and
send_notification only mapped ConnectionError, TimeoutError and OSError to
ChannelUnavailable. A disconnect therefore spent the counted retry budget
and settled the row as failed after about 15 minutes, although the docs
say transport outages park without consuming retries.
send_notification now asks the client's is_connected first. The new test
uses a real, never-connected SDK client: it raised RuntimeError after three
retries before, and parks now.
* docs(scheduler): describe which outcomes notify; mirror the changelog entry
- README (en/zh) and CONFIGURATION.md: only runs that finish as success or
failed notify; occurrences that end without a finished run (launch error,
queue timeout, restart recovery) do not. Parking lasts up to about a day.
- CHANGELOG: cite the issue as text like the rest of the file, and add the
entry and link reference to CHANGELOG_zh.md.
- AGENTS.md: drop the backend pointer so this branch adds nothing to the
guidance chains closest to their size limit; the scheduler-side rules
and the lifespan order now live in app/channels/AGENTS.md.
* docs(changelog): cite #6135 next to #4843 for the notification outbox entry
* fix(scheduler,channels): re-check the IM binding at delivery time; park WeCom transport errors, count its rejections
Review round 1 on #6135.
The delivery worker now takes resolve_connections (the Gateway wires
ChannelConnectionRepository.list_connections) and, before the channel
liveness check and the summary read, requires the row's owner to still
hold a connected (provider, external_account_id). A disconnected target
is finalized with mark_failed(terminal=True); a lookup failure or an
unexpected result shape parks the row. Terminal outbox rows ignore
later status writes.
WeCom: the SDK raises RuntimeError from send_message for a socket that
is not open and for a dropped connection, but also for a platform
rejection ("Reply ack error: errcode=..."). Only that prefix keeps
counting; other RuntimeErrors, ConnectionError/TimeoutError/OSError and
websockets' ConnectionClosed park the row as ChannelUnavailable.
* fix(channels): treat a WeCom ACK rejection as proof the transport is up; pin the rejection prefix through the retry branch and the real connection-row shape
---------
Co-authored-by: GaoTianyuanhr <75788039+GaoTianyuanhr@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: GaoTianyuanhr <GaoTianyuanhr@users.noreply.github.com>
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
* fix(middlewares): include file line count in read-before-write block message and clarify read_file range error
* fix(middlewares): address review feedback on line count pluralization and readability
* fix(middlewares): address round-2 review feedback for PR #6020
Stamping bug (_attach_read_mark, ggnnggez item 1):
Originally planned as a standalone issue; folded here since it is the exact mechanism behind the #6019 gate bypass. _attach_read_mark now returns early when message.content is in READ_FILE_NO_CONTENT_RESULTS - {READ_FILE_EMPTY}. In-band errors (invalid range, start_line exceeds file length, etc.) are returned with status='success' by LangChain so the existing message.status == 'error' guard did not catch them.
HTTP code collision [P3] (willem-bd):
Pre-stamp TOOL_META_KEY on the blocked ToolMessage before normalize_tool_result so files with 401, 403, 404, or 500 lines are not misclassified as auth/server errors by _classify_error_text. normalize_tool_message returns early when an existing TOOL_META_KEY entry is found.
Line-count parity (ggnnggez, inline L317):
Replace splitlines() with count_file_lines() in read_file_contract.py using count('\n'), matching LocalSandbox text-mode iteration and the _truncate_read_file_output formula. splitlines() over-counts for files with \f or \u2028 separators (e.g. pdftotext output).
Read hint (ggnnggez, inline _BLOCK_MESSAGE):
For empty files the ranged-read suggestion is omitted. For non-empty files the hint gives a concrete start_line=max(1, N-29), end_line=N.
Constant rename (ggnnggez, read_file_contract.py):
READ_FILE_INVALID_RANGE is the new canonical name; READ_FILE_EMPTY_RANGE kept as a backward-compat alias. _LEGACY_READ_FILE_EMPTY_RANGE removed (unreachable after the #6019 contract update).
Tests:
- TestBlockMessageLineCount parametrized with \f and \u2028 parity cases
- TestStampingGateBugFix: 4 no-content variants, gate-closed integration probe, empty-file-still-stamps assertion
- test_gate_block_on_numeric_line_count_is_recoverable: 401/403/404/500
- test_contract_constants_membership replaced with test_contract_constants_parity_with_localsandbox
---
NOT INCLUDED: ggnnggez item 2 (file length in read_file range error strings). READ_FILE_NO_CONTENT_RESULTS uses exact string membership and is also consumed by build_skill_usage for skill-read detection; making those strings dynamic requires a dedicated contract refactor. Deferred to a follow-up PR.
Note: test_web_fetch_relative_links failures in the full suite are pre-existing on upstream/main and unrelated to this change.
* test: clarify blank ranged reads in the write gate
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(middleware): publish externalized outputs atomically
_externalize reported failure with None but had already created the file, so a
write that failed part-way (disk full, interrupted request) left a truncated
file under the final name in the outputs directory. Write to a sibling temp file
and os.replace it into place, removing the temp file when the write fails.
* style: drop the unused pytest import from the externalize test
* test(scripts): cover discord extra auto-detection
The discord mapping in scripts/detect_uv_extras.py worked but was neither
listed in the "currently maps" docstring nor covered by a test, unlike the
buzz mapping that grew up beside it.
* fix(channels): point the Discord missing-dependency hint at the discord extra
discord.py is the only IM transport that is not a core dependency — it ships in
the optional `discord` extra, which detect_uv_extras.py, backend/Dockerfile and
deploy.sh already expand. Its ImportError hint told users to run `uv add
discord.py` instead, which rewrites pyproject.toml/uv.lock in a tree every
restart syncs with `--locked`. Mirror the buzz hint wording and note the extra
in config.example.yaml next to the buzz block.
---------
Co-authored-by: huyengiang101086-spec <270684046+huyengiang101086-spec@users.noreply.github.com>
* fix(skills): drain custom-skill mutation tails across cancellation
* fix(skills): log drained skill mutation failures and drain state/install tails
Address review feedback on #6078: record cancelled-then-failed mutation
tails inside the drained unit (exception type only, per the managed-models
precedent), and extend the drained tail to the remaining write-to-cache
shapes in update_skill and the skill install path.
* test(skills): pin the lock contract under the drained cancellation timing
The cancelled update_skill caller now drains its mutation tail, so
CancelledError surfaces only after the tail settles. Start the MCP
writer without awaiting the cancelled task first, assert the section
stays owned while the drained tail is in flight, then await both and
assert the final order.
---------
Co-authored-by: xiaodu55 <xiaodu55@users.noreply.github.com>
* fix(agents): keep the todo completion reminder when a model call is retried
TodoMiddleware drained its queued completion reminder inside
wrap_model_call before calling the handler. LLMErrorHandlingMiddleware
sits outside it and retries a failed call by running that wrap again, so
the retry went out with an empty queue while the run had already spent
one of its two reminders on the lost one.
Drain the reminder once per wrap, and when the handler raises put it
back in front of the queue without counting it again, then re-raise. A
run whose reminder state was cleared while the call was in flight does
not get it back.
* docs(changelog): reference #6132 in the todo reminder retry entry
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(channels): split Telegram messages by UTF-16 code units, not code points
* perf(telegram): clip stream previews without walking the whole reply
Every cumulative stream update measured the full answer and then split the full
answer to keep only its first chunk, so a long reply cost two complete passes
per update on the async channel path. _first_utf16_chunk stops at the first
character that does not fit, and the preview is built from that bounded prefix.
The tests now measure chunk length with the UTF-16 codec instead of a copy of
the production formula, cover the width helper and the bounded clip directly,
and fail if the preview stops being bounded by the limit.
* fix(uploads): reject filenames reserved for temporary staging
Guard thread and shelf upload names, validate SDK batches before copying, and preserve legacy shelf downloads. Add regression coverage and document the reserved staging namespace.
* fix(uploads): report reserved names before HTTP batch ingestion
Preflight reserved staging basenames before thread storage or sandbox setup. Return HTTP 400 with a rename hint; add HTTP and strict blocking-I/O regressions and update upload documentation.
* fix(uploads): reject Windows-style staging basenames on POSIX
Normalize backslashes before the HTTP reserved-name preflight. Add POSIX basename regressions for single and mixed batches and strict blocking-I/O coverage. Clarify the preflight scope and upload documentation.
Follow-up to the review of PR #6122; AI-assisted with OpenAI Codex.
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(channels): keep blocking SDK teardown off the event loop in stop()
SlackChannel.stop() called SocketModeClient.close() inline. close() joins
the SDK's message-processor thread (~0.7s every call, measured) and waits
for in-flight listeners, whose blocking Web API calls can run up to the
WebClient timeout, so every run, stream and channel on the Gateway stalled
during shutdown and POST /api/channels/{name}/restart.
Feishu and DingTalk had the same shape: they joined their SDK threads
inline with a 5s timeout, and since those threads only return on a fatal
error, the join normally froze the loop for the full 5s.
The teardown now runs through asyncio.to_thread. Slack detaches the client
first so a still-queued connect skips it, and shields and tracks the close
task so a cancelled stop() leaves it running and a retried stop() awaits
the same close instead of closing twice.
tests/blocking_io/test_channel_stop_offloop.py anchors all three stop()
paths to the strict Blockbuster gate, driving a real slack-sdk client.
* docs(changelog): reference #6134
* fix(channels): join the Discord client thread off the event loop in stop()
DiscordChannel.stop() still joined its client thread inline with a 10s
timeout. The thread usually exits right after the cross-loop client close,
but a timed-out close or a slow _run_client() drain keeps it alive for the
whole timeout, stalling the shared Gateway loop the same way the Slack,
Feishu and DingTalk teardown did. Run the join through asyncio.to_thread so
stop() matches the channels/AGENTS.md teardown rule, and add Discord to the
strict Blockbuster anchor.
extensionMap is keyed by extension and valued by language, but the Julia row
was written backwards as `julia: "jl"`. checkCodeFile tests membership with
`extension in extensionMap`, so ".jl" never matched and a Julia artifact was
not a code file at all, while ".julia" resolved to a language named "jl".
Add both `jl: "julia"` and `julia: "julia"`, the same pair form the map already
uses for haskell/hs and elixir/ex, so no extension that classified before stops
classifying.
* fix(paths): reject Windows reserved names on host-visible paths
Uploads, skill support files, and sandbox path checks now share one helper that rejects CON, PRN, AUX, NUL, COM1-COM9, and LPT1-LPT9 (including names like con.txt) and components that end in a dot or space. The check is not limited to Windows so a Linux host cannot store a name that breaks when the same tree is used on Windows.
* fix(paths): validate quoted local bash paths as whole words
The absolute-path scan stopped at whitespace, so a quoted upload name such as report. final.txt or CON notes.txt was rejected as a truncated segment. Local bash checks now use the shlex argument for traversal and Windows-incompatible segments.
* fix(paths): do not block reads of existing host names
* style(paths): format with ruff 0.15.12
* fix(frontend): strip closing hashes from web fetch titles
Signed-off-by: 97three <2212371308@qq.com>
* fix(frontend): scan web fetch heading suffixes in linear time
Signed-off-by: 97three <2212371308@qq.com>
---------
Signed-off-by: 97three <2212371308@qq.com>
* fix(middleware): validate exact byte size on sandbox tool output externalization
* test(middleware): simulate partial write in sandbox externalization size test
* Fix PR number in CHANGELOG.md
Updated the changelog to reflect the correct issue number for a truncated file write.
* Update CHANGELOG_zh.md for issue #6112
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(threads): re-anchor re-persisted history rows to their earliest seq
reconcileThreadHistoryRows collapsed a repeated (run_id, identity) to
its latest-seq copy, deleting the earliest-seq row before
buildVisibleHistoryMessages could apply its earliest-seq-wins
anchoring. A message re-persisted under a newer seq therefore jumped
to the tail of the thread, contradicting the documented contract in
this function's docstring, in buildVisibleHistoryMessages, and in the
backend get_message_seqs earliest-seq-wins rule.
Re-anchor the surviving row to the earliest seq its identity held in
the merged window after deduping, and make the retained-history
stability check value-based so re-anchored rows do not force a state
update on every pass.
Fixes#6125
* fix(threads): address history reconciliation review
---------
Co-authored-by: liwenjie200543 <175601277+liwenjie200543@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(vllm): fall back to legacy reasoning_content field
* fix(vllm): document the reasoning_content fallback and pin precedence
The previous change set left the docs describing the provider as
preserving only `reasoning`:
- the module docstring now mentions the legacy `reasoning_content`
wire field, and the vLLM Provider bullet in models/AGENTS.md is
extended to match the repo's documentation-update policy;
- _create_chat_result evaluated choice["message"] twice, once per key;
hoist it into a single local and read both keys from it;
- two regression tests pin that a payload carrying both fields keeps
`reasoning`, so a future refactor cannot silently invert the
precedence (the existing tests only covered the absent-`reasoning`
case).
Reported-by: willem-bd
* refactor(vllm): share reasoning fallback and cover follow-up payloads
---------
Co-authored-by: inchang-ing <197932532+inchang-ing@users.noreply.github.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
Mirror the English README's Startup Modes section (introduced and refined in #5995) into the Chinese version: the startup mode matrix, the stop/restart commands, and the SKIP_FRONTEND_BUILD opt-in note. README_zh previously had no documented way to run services in daemon or production mode or to stop/restart them.
Co-authored-by: sjr666666 <204130665+sjr666666@users.noreply.github.com>
* fix(auth): reject a non-ASCII ID token nonce instead of raising
The nonce comparison in validate_id_token still used the local
secrets.compare_digest on str operands, which raises TypeError when either
side holds non-ASCII characters. The nonce claim is provider-controlled
text, so such a provider response turned the callback's sso_failed redirect
(OIDCValidationError) into an unhandled 500.
Route it through the shared constant_time_equals helper #6076 introduced
for the same five compare sites, and drop the now-unused local helper.
Signed-off-by: yetuge <2219677952@qq.com>
* fix(auth): reject a non-string ID token nonce claim like a mismatch
A provider returning a truthy non-string nonce claim (e.g. an int)
slipped past the missing-claim guard and crashed the comparison helper
with AttributeError, which the callback turned into a 500 instead of
the sso_failed redirect. Guard the claim type before comparing so every
malformed nonce takes the same OIDCValidationError path.
Signed-off-by: yetuge <2219677952@qq.com>
---------
Signed-off-by: yetuge <2219677952@qq.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
Why this change:
In Lark broker mode the shim forwards piped stdin to the broker only
after EOF. AIO's /v1/bash transport keeps a subprocess stdin pipe open
for writes, so a bare `lark-cli ...` never saw EOF and the shim blocked
forever on input that would never arrive. An earlier brace-group/eval
approach fixed the hang but changed top-level parsing (aliases/extglob)
or let a trailing backslash consume the appended delimiter.
What changed:
- _run_bash_exec prefixes the original command with `exec < /dev/null`
on its own line, giving the fresh, per-invocation session immediate
EOF on default stdin. Explicit pipes/heredocs/files still override
fd0. The plain prefix keeps top-level parsing and appends no
delimiter a trailing backslash can consume.
- The persistent /v1/shell PTY transport keeps commands unchanged; the
shim already treats TTY stdin as no input.
- The shim drains non-TTY stdin under grace/tail idle budgets
(DEERFLOW_LARK_BROKER_STDIN_{GRACE,TAIL}_SECONDS, default 2s, valid
(0, 600]); an idle pipe exits 124 without contacting the broker, and
read errors fail loudly instead of becoming empty input.
- Broker logs each exec as argc/exit/elapsed only (never argument
values) and warns on concurrency-cap rejections.
Problem solved:
Lark commands without explicit stdin hung indefinitely in broker mode,
stalled producers could execute partial writes, and broker logs risked
capturing argument values. Docs updated: README, harness AGENTS.md,
docker/lark-cli-broker README.
* fix(cli): exit non-zero when a headless run ends in an LLM error fallback
LLMErrorHandlingMiddleware turns provider failures (e.g. an expired
credential returning 401) into an AIMessage flagged
`deerflow_error_fallback` instead of raising. The Gateway run status and
the subagent executor already treat that flag as a failure. #5963 gave
`deerflow --print` / `--json` an error boundary for raised exceptions,
but a flagged fallback never raises, so both modes still printed the
fallback text and exited 0.
Both headless modes now follow #5963's failure contract when the run's
final AI message carries the flag: `--print` keeps the fallback text on
stdout and adds an `Error:` line on stderr; `--json` appends the terminal
`{"type": "error"}` record. Both exit 1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(cli): share final-answer accumulation with the embedded client
* docs: scope headless accumulator guidance to TUI
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* feat(authz): enforce skill authorization at assembly and slash-activation (#4063 Phase 3)
Phase 3 / Skills — the second of three resource-type PRs (Models, Skills,
Sandbox). The RBAC provider already maps "skill" → config key "skills"
(rbac.py _RESOURCE_POLICY_KEYS), so no schema change is needed.
Layer 1 (assembly-time, mirrors Phase 1B):
- filter_available_skills_by_authorization() in skill_filter.py filters the
skill-name allowlist (set[str] | None) by the provider's "skill" policy.
In lead agent, runs after _available_skill_names() so the catalog
(describe_skill) and SkillActivationMiddleware share one filtered set.
In subagent executor _load_skills(), filters config.skills before disk load.
- filter_resources_by_authorization() in enforcement.py — generic version of
filter_tools_by_authorization for any resource with a name attribute.
Layer 2 (slash-activation):
- No new middleware code needed. SkillActivationMiddleware._resolve_activation
already checks "reference.name not in self._available_skills". Since Layer 1
filters the allowlist, denied skills cannot be slash-activated.
authorization.enabled: false is a complete no-op. available_skills=None (no
agent allowlist) + enabled: resolves all skill names from config and filters.
12 new tests + 243 existing authz/skills tests pass.
* test(subagents): wait on the chained future instead of polling done() in caller-loop teardown test
* feat(authz): enforce action-scoped skill:activate at runtime and reuse the catalog loader for filter candidates
* feat(authz): arm skill:activate on the subagent chain, unify the client skill-surface user, and record the Layer 2 reversal in the design notes
* chore(authz): align the subagent filter's fallback user with the chain's bucket convention
* feat(authz): gate autonomous skill loading on skill:activate, close the bootstrap assembly leak, and reuse one provider per subagent assembly
Review round 4 (willem-bd, CHANGES_REQUESTED):
- [P1] Autonomous load path: describe_skill filters its match set on the
action-scoped skill:activate decision (all-denied is indistinguishable
from "no match"), and ToolErrorHandlingMiddleware stamps skill_context
entries only when the action allows it — a denied SKILL.md read gets a
skill_context_denied marker instead, so durable context, allowed-tools
policy, and autonomous secret bindings never activate a denied skill.
File visibility itself stays Layer 1 (sandbox-projection semantics).
- [P1] Bootstrap assembly: build_middlewares/apply_prompt_template now
receive the authorization-filtered set instead of the raw
_BOOTSTRAP_SKILL_NAMES, and an empty skill_names index is preserved
when deferred discovery is on rather than round-tripping through None
into the legacy full-metadata prompt that re-advertises a denied
skill (deferred-off keeps the legacy rendering; the filtered empty
allowlist blanks it on denial).
- [P2] SubagentExecutor resolves ResolvedSkillAuthorization once per
assembly (_resolve_skill_authorization, memoized per executor) and
shares it across the Layer 1 filter (authorization=), the activation
middleware, and both new gates — uncached provider factories can no
longer hand the layers different instances/policy snapshots.
All gates share one helper (skill_filter.skill_activation_allowed) so
slash, describe, and skill-file-load cannot drift; provider errors keep
the configured fail-closed/fail-open policy. skill_authorization is
threaded through _build_runtime_middlewares into both runtime builders.
Tests: describe gating (deny/fail-closed/fail-open/disabled), read
stamping (denied marker vs entry) + extract_skills skip without warning,
bootstrap three-way param (allow/deny x deferred on/off), lead + subagent
chain wiring for the stamp gate, and the distinct-instance provider
factory repro. Each fix is mutation-verified (reverting it fails its
test). Full A/B suite run: no new failures vs HEAD.
* feat(authz): re-authorize persisted skill_context before reuse and route async paths through aauthorize
Review round 5 (willem-bd):
- [P1] Persisted entries: skill_context survives across runs, so both
consumers now re-check the action-scoped skill:activate decision before
reuse. SkillToolPolicyMiddleware._active_skills_for_paths skips a
denied skill (same skip as a disabled/unallowlisted one; all-denied
keeps the existing fail-closed-to-builtins semantics) and
_in_context_secret_sources binds no secrets for a denied entry.
skill_authorization is threaded into SkillToolPolicyMiddleware on both
chains.
- [P2] Async provider API: skill_activation_allowed_async mirrors the sync
helper through await provider.aauthorize() (same pattern as
authorize_sandbox_execution_async). Async hooks precompute an
activation_decisions map on the event loop (slash target + persisted
entry names, one batch of awaits) and hand it to the worker-thread
handler, which consults the map instead of the sync API:
SkillActivationMiddleware.awrap_model_call (activation + secret
binding), ToolErrorHandlingMiddleware.awrap_tool_call (_amaybe_stamp
awaits the decision before stamping), and SkillToolPolicyMiddleware's
awrap_model_call/awrap_tool_call (entry names + slash source name).
describe_skill becomes a StructuredTool with both a sync function
(invoke -> authorize) and a coroutine (ainvoke -> aauthorize), so async
agent execution never routes a loop-affine provider through the sync
API while sync callers keep their semantics.
Tests: allow-read-then-deny-next-run regressions for secret binding and
allowed-tools (partial deny does not poison, all-denied fails closed to
builtins), and an _AsyncOnlyProvider (sync authorize always fails,
aauthorize works) driving four async-path assertions (activation, stamp
allow/deny, describe, tool policy). All six fixes are mutation-verified;
full A/B suite run against HEAD shows no new failures.
* fix(authz): authorize skill activation by declared Skill.name, not directory basename
ShenAC-SAC's P2: bundled skills may declare a name that differs from their
directory (skills/public/vercel-deploy-claimable declares name: vercel-deploy),
but the runtime gates derived authorization targets from the container path.
- New deerflow/skills/container_registry.py: build_container_path_registry()
+ canonical_skill_name() — the one path-to-Skill.name resolution shared by
the activation, tool-policy, and read-stamping middlewares.
- ToolErrorHandlingMiddleware resolves the read path through the (user-scoped,
via new user_id kwarg threaded from both runtime builders) registry before
the skill:activate decision, sync and async (registry lookup in a worker
thread, aauthorize() on the loop). RBAC allow by the declared name now
activates; allow by the directory name does not.
- SkillToolPolicyMiddleware async hooks canonicalize the policy paths first
(three-phase: thread registry -> loop aauthorize -> thread filter reusing
the registry; the cached per-model-step decision still short-circuits),
so the decision map is keyed by the names _active_skills_for_paths checks —
no more map miss falling back to the sync provider API from the thread.
_entry_names removed: persisted entry["name"] is path-derived by
construction; entries are resolved by path instead.
- SkillActivationMiddleware collects slash names (already canonical) plus
persisted entry *paths*, canonicalized off-loop via
_canonical_names_for_paths().
Regressions (test_skills_authorization.py 45->50): sync stamp target, RBAC
allow declared-name/deny directory-name both directions, async stamp via
_AsyncOnlyProvider, composed async slash -> tool-policy keeps the skill's
tools with zero sync fallback, persisted-entry secret binding survives the
name mismatch. All three mutations (stamp target, policy map key, entry
canonicalization) fail their tests.
* fix(authz): close the sync-fallback class on skill activation paths
willem-bd's two P2s on the canonical-name fix, plus a fourth instance of
the same class found while sweeping the membership:
- SkillToolPolicyMiddleware: a transient prepass registry failure is now
preserved (_REGISTRY_LOAD_FAILED marker) — the worker-side filter applies
its own fail-closed treatment instead of silently retrying storage,
which recovered skills absent from the decision map and fell back to
the synchronous provider API from the thread (fail-open retaining a
denied skill's allowed-tools).
- SkillActivationMiddleware: an empty entry-path list no longer pays a
full skill-tree scan on ordinary authorization-enabled async model
steps (previously a passive path with zero registry I/O).
- 4th instance (reproduced with the PR fixtures, then fixed): the entry
prepass failure on the secret-binding path let the consumer-side
_load_skill_registry_by_path silently recover and resolve entry names
outside the decision map — the same sync fallback, fail-open binding
the aauthorize-denied skill's declared secret. The fix goes one step
further than failure propagation: the prepass hands down the registry
*snapshot* the decision map was keyed by, and entry sources resolve
against that snapshot, so a map miss is impossible by construction
(this also closes the narrower window where both loads succeed but
disagree because storage changed mid-step). The slash source is
deliberately unaffected: it does not consult the decision map
(validated once at activation) and keeps its fresh per-call load.
- Class membership swept: skill_activation_allowed() sync call sites are
exactly four (stamping sync path and describe's sync implementation are
legitimately sync); the only two async-reachable fallback points (the
_activation_allowed helpers) are both construction-guarded now.
Structural hardening (typed async-batch decision object where a map
miss is fail-closed by construction) recorded in the plan notes as a
follow-up, anchored at those two helpers.
- _RegistryArg type alias replaces the degenerate dict|object|None union
in the policy middleware signatures.
Tests (test_skills_authorization.py 50->54): transient prepass failure ->
builtins-only with exactly one storage attempt and zero provider calls;
ordinary async step -> zero registry scans; two-call scenario -> entry
binding zeroed while the slash binding survives; entries resolve against
the prepass snapshot under a mid-step rename. Mutations verified red:
marker reverted to None, over-broad slash kill, snapshot bypass.
* fix(authz): correct secret-binding reference points and gate the rendered skill reminder
Review round 9 (willem-bd) plus construction-audit follow-ups on the same
class — every async divergence between two loads of the same data now has
an explicit, pinned reference point:
- Slash secret source resolves from a FRESH post-activation registry: the
prepass snapshot is taken before the aauthorize awaits, so binding the
slash source from it would activate NEW_KEY content while injecting
OLD_KEY when a skill's declared secrets change in that window. The
'one scan per step' bonus that motivated sharing the snapshot is
withdrawn — consistency wins. The two secret sources resolve
independently: a transient fresh-load failure zeroes only the slash
binding; entries keep binding from their own snapshot.
- Same-skill dominance: when a skill is both the slash source (fresh) and
a persisted entry (snapshot), only the activation-era declaration binds
— unioning both eras injected a secret the skill no longer declares.
Aligns with the tool-policy rule (slash activation dominates for the
run); different-skill entries still stack additively.
- The rendered 'Active skills' reminder (DurableContextMiddleware) no
longer advertises skills the provider denies — surface consistency with
describe_skill. Publication/consumption split: the activation
middleware publishes per-step path-keyed skill:activate decisions into
the run context (async: reused from the prepass snapshot, zero extra
I/O; sync: one scan + authorize() only when entries exist — passive
steps stay scan-free), the renderer consumes them and filters its
render copy (state untouched; unpublished = permissive, matching the
other run-context carriers). New carrier key listed in
REDACTED_CONTEXT_KEYS with a redaction regression.
- Wiring regressions on both chains for everything added above (user_id
to the stamp gate, skill_authorization + token pairing + activation-
before-durable ordering to the durable renderer): each mutation
verified red (M4-M13 across this change set).
- Tests 54->65: slash post-activation lookup, source independence,
dominance, sync-chain sync-API path, render filter (async/sync/
disabled with load-count assertions), chain wiring x2, redaction
listing. Plan notes record rounds 9-2..6 including the reference-point
post-mortem and the audit axes exhausted.
* fix(authz): anchor same-skill dominance on the slash identity, not binding results
Review round 10 (willem-bd) plus a self-audit extension the new checklist
items caught in the first fix itself:
- R10 (P2): the same-skill exclusion was derived from successfully BOUND
slash sources. When the prepass snapshots required-secrets [OLD_KEY]
and the operator strips the declaration during the await window, the
slash source binds nothing (required_secrets gate) — the exclusion set
stayed empty and the stale entry view kept injecting OLD_KEY for a
skill that now declares no secrets. The exclusion now derives from the
AUTHENTICATED slash activation identity (run-context source path,
owner-token checked): the identity's declared name resolves from the
fresh registry first, then the entry snapshot, and joins the exclusion
set regardless of binding success.
- Self-audit (value->gate->downstream on the fix itself): the declared
name can also change mid-run (same path, foo->bar). A name-anchored
exclusion then holds only the new name while the snapshot entry
resolves the same path to the old one — both eras bound again. The
PATH is the stable identity shared by the activation and every earlier
read: entries at the slash identity's path are now excluded before
registry resolution (exclude_paths), independent of what either
version calls the skill; the name-based set remains for cross-path
same-name shadowing.
Tests 65->67: no-secrets-era dominance (ACTIVE secrets None under the
OLD_KEY-to-empty interleaving), path-anchored dominance across a mid-run
rename (only the activation-era key binds). Mutations verified red:
identity-derived exclusion disabled (M14), path exclusion disabled
(M16).
* test(authz): pass explicit config to lead middleware wiring test
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(gateway): drain agent and user-profile writes across cancellation
A client that disconnects mid-request cancels the handler task. On the four
persistent-write endpoints of /api/agents - create, update, delete, and the
USER.md write - a bare asyncio.to_thread then either cancels a still-queued
worker (the write silently never happens) or detaches from a running one and
drops its failure, because the handler's except never runs.
Route the four writes through await_drained like the managed-subagent (#6023),
managed-model (#6024), and custom-skill (#6078) mutations, logging a lost
worker failure with the exception type only before the drain consumes it.
Expected domain errors (AgentExistsError on create) re-raise unlogged; the
caller still maps them to a 409 while connected. Reads stay bare: abandoning
them loses nothing.
Signed-off-by: yetuge <2219677952@qq.com>
* fix(gateway): address review nits on the agents write drain
Make expected_errors a positional-only parameter before *args so the
ParamSpec signature is checker-clean under PEP 612 (the one call site with
extra worker arguments passes () explicitly), acknowledge that
non-cancelled failures log twice like artifacts.py's drained commit does,
and drop the unused request stub helper from the tests.
Signed-off-by: yetuge <2219677952@qq.com>
---------
Signed-off-by: yetuge <2219677952@qq.com>
* fix(agents): anchor agent-name validation against a trailing newline
Two copies of the shared ^[A-Za-z0-9-]+$ grammar still used re.match,
whose $ also matches before a final newline, while every other copy uses
fullmatch. POST /api/agents {"name": "reviewer\n"} therefore passed the
router guard and died inside the strict file store as an HTTP 500, and
DeerMem accepted the same name as an agents/{name}/facts directory.
* test(agents): state which trailing-newline params the regression actually covered
.match only accepted "reviewer\n" on main; "reviewer\n\n" and "reviewer \n"
were already rejected, so the docstrings should not read as if every param in
these parametrizations was broken before the fix.
* docs: replace nonexistent async client API in harness docs
DeerFlowClient has no async methods: client.astream() and client.ainvoke()
do not exist, so the "Create Your First Harness" tutorial and the harness
integration guide fail with AttributeError on their first call. Overrides
were also shown as a nested config={"configurable": {...}} dict, which
stream()/chat() swallow via **kwargs and silently ignore.
Switch the examples to the shipping API (sync stream()/chat(), flat kwarg
overrides, agent_name as a constructor arg) in the EN and ZH copies of both
pages. backend/docs/STREAMING.md is left alone: its astream mentions are
design discussion and LangGraph internals, not user-facing examples.
Assisted-by: Claude Code / claude-opus-5
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>
* docs: serialize StreamEvent before StreamingResponse, use real event names
Review follow-up on the two points raised in #5843.
P2 — the streaming examples handed a generator of `StreamEvent` dataclasses
straight to Starlette. `StreamingResponse` calls `.encode()` on any chunk that
is not `str`/`bytes`, so the first event raised
`AttributeError: 'StreamEvent' object has no attribute 'encode'`. Both
occurrences (the `run_agent` snippet and the inline FastAPI route, which shipped
the dataclass `repr()` through an f-string) now serialize to SSE frames built
from `event.type` / `event.data`.
P3 — the tutorial showed mapping literals with `messages` and `thread_state`,
neither of which exists. The real contract is
`StreamEventType = Literal["values", "messages-tuple", "custom", "end"]`
(backend/packages/harness/deerflow/client.py:128) and the stream yields
dataclasses, so the section now shows `StreamEvent(...)` with the real names
plus a consuming loop.
The zh tutorial had no equivalent section at all; added for parity.
Assisted-by: Claude Code / claude-opus-5
Machine: A-Mac16-2019-PaloAlto
Account: tonydzi
Operator: robot:pr-reply-daily
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: drop the nonexistent load_config() from every harness snippet
Review follow-up on #5843 (@willem-bd): the rewritten tutorial still opened
with `from deerflow.config import load_config`, which raises ImportError
before `client.stream(...)` is ever reached — `deerflow/config/__init__.py`
exports `get_app_config` and friends, and the string `load_config` does not
appear anywhere under `backend/packages/harness/deerflow/` (the only near
match is `load_configured_extension_middlewares`).
The call was also unnecessary: `DeerFlowClient.__init__` already resolves
configuration itself (client.py:220-222 — `reload_app_config(config_path)`
when a path is given, otherwise `get_app_config()`, which auto-loads from
the default path or DEER_FLOW_CONFIG_PATH).
Fixed the class rather than the one snippet you flagged: the same dead
import appeared 10 times across 4 pages (EN + ZH tutorial and integration
guide). Every one is gone. The "configuration in embedded mode" section now
shows the API that exists — `DeerFlowClient(config_path="...")` — instead of
`load_config(config_path="...")`.
Also from the review:
- the `end` example showed only `total_tokens`; `stream()` always yields
input/output/total (client.py:920), and the snippet right below prints the
whole usage dict, so a reader saw keys the docs omitted. Both EN and ZH now
show the full payload, and the `values` example visibly elides `artifacts`
and `summary_text` instead of pretending they are absent.
- the two SSE comments on the ZH integration page stayed English while the
example above them was translated; translated.
Verified statically, not executed: `grep` for the symbol across the harness
package (zero hits) and a read of `DeerFlowClient.__init__`. I did not stand
up a backend to run the tutorial end to end.
Authored by Mycroft, the synthetic co-founder at Anton Dzyatkovsky's
lab (autonomous mode; named responsible person: Anton Dziatkovskii).
Assisted-by: Claude Code / Opus 5 (review response, implementation)
Machine: MacBook-Anton
Account: a
Operator: mycroft (autonomous routine github-thread-watch)
Signed-off-by: tonydzi <194927794+tonydzi@users.noreply.github.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: fix native artifact SSE encoding and custom config advice
---------
Signed-off-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>
Signed-off-by: tonydzi <194927794+tonydzi@users.noreply.github.com>
Co-authored-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
* fix(agent): drop script guidance from the lead prompt when no bash tool is bound
The workspace section told the agent how to write scripts and commands
("prefer relative paths", "avoid hardcoding /mnt/user-data/... inside
generated scripts") even when no tool could run them. With the default
LocalSandboxProvider host bash is off and the bash tool is removed, so
agents wrote helper scripts nothing could execute, then stalled or
computed by hand.
apply_prompt_template() takes bash_available (default True, so other
callers keep today's text). The lead agent (both assembly branches) and
the embedded client pass whether their authorized tools, bound or
deferred, include bash. Without it:
- the two script bullets become one line: no bash tool is bound, work
out results directly and write them with write_file;
- the subagent section no longer shows direct bash examples, and marks
the bash subagent unavailable, since it inherits the lead's tool groups;
- the ACP hint no longer suggests `bash cp`.
* fix(goal): show the goal evaluator the user's answer to a Human Input Card
The web UI sends a Human Input Card answer as a hidden human message
(hide_from_ui with a human_input_response) and shows it on the card.
format_visible_conversation() skipped every hidden message, so the
evaluator saw the question but never the answer. Following the user's
answer then read as the assistant guessing, and the goal stood down with
needs_user_input although the user had answered.
A hidden human message whose human_input_response passes
read_human_input_response() now adds a line
"User (Human Input Card answer): <value>" from the structured value; the
question is already on the card's own line. Other hidden messages, such
as goal continuations, stay out. Over the 12,000-char cap the latest
answer the cut drops is kept at the top next to the latest request, and
is never taken for the request itself. Without card answers the cut is
unchanged.
* fix(agent): keep the subagent examples and bash hint in line with the bound tools
Review follow-up for the lead prompt without bash.
- The last delegation example follows the tools. With a bash subagent on
offer it is unchanged. A lead with bash and no bash subagent keeps
"Run a routine test, build, or git command directly." without the Bash
subagent sentence. A lead without bash gets a file-work example instead.
- The unavailable-bash description no longer points at the sandbox
provider. Since #2253 it renders only when the bash subagent is
registered and the lead binds no bash, where the sandbox is not the cause.
---------
Co-authored-by: Totoro-qaq <279883115+Totoro-qaq@users.noreply.github.com>
* fix(gateway): reject non-ASCII tokens instead of raising in constant-time compares
hmac.compare_digest raises TypeError when a str operand holds non-ASCII
characters. Starlette decodes header bytes as latin-1 and query values as
UTF-8, so a single 0xE9 byte in a CSRF token, GitHub webhook signature,
internal auth token, or OIDC state turned the expected 403/401 into a 500.
Route the five client-facing comparisons through a new
app.gateway.utils.constant_time_equals helper that compares the UTF-8
encodings ("surrogatepass" so even a lone surrogate cannot raise). ASCII
results are unchanged, and no bypass was possible before the fix.
* fix(provisioner): reject non-ASCII API keys instead of raising; drive OIDC test over HTTP
The standalone provisioner compared X-API-Key with secrets.compare_digest on
str operands, so a non-ASCII header raised TypeError (500) the same way the
Gateway sites did. It cannot import app.gateway, so encode inline.
The OIDC state regression test now mounts the auth router and sends the
percent-encoded state through a real signed state cookie instead of stubbing
get_state_cookie and calling the handler directly.
* docs(changelog): reference #6076
---------
Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
test_refresh_skew_rejects_negative_and_boolean omitted the required
token_url, so McpOAuthConfig(refresh_skew_seconds=-1) failed validation
on the missing field alone -- deleting the ge=0 bound and the boolean
validator would leave the test green. Pass token_url (the same URL the
refresh_skew_seconds=0 boundary test uses) so only the guards can
satisfy pytest.raises; verified by stripping both guards from
extensions_config.py, which now fails exactly the two parametrized
cases, and restoring them, which returns the suite to 42 passed.
Also move SandboxOwnershipConfig up beside the other config imports
and drop the redundant in-test `import pytest` (review nits on #6026).
Generated-by: ZCode (GLM, coding agent)
Co-authored-by: liwenjie200543 <liwenjie200543@users.noreply.github.com>
* fix(agents): treat a null max_total_subagents override as unset
An API caller sending "max_total_subagents": null in the run context
failed that run with a TypeError: the key was present, so
dict.get(key, default) returned None and building SubagentLimitMiddleware
compared None with an int.
Resolve the per-run delegation cap through one helper,
effective_total_subagents_per_run, which treats None as "use
subagents.max_total_per_run" and clamps to 1-50. The lead agent's
middleware and assembly paths, DeerFlowClient, and the system prompt all
use it, so the extension host policy and release policy now report the
enforced cap instead of an out-of-range request.
* docs(changelog): reference #6088 in the null max_total_subagents entry