Compare commits

...
Author SHA1 Message Date
nesquena-hermesandn 27b7064c4c Release: decode state.db content sentinel in the message projection (#7492) (#7590)
Co-authored-by: n <a@n>
2026-09-15 17:16:53 -07:00
totalitarianandtotalitarian 963e1bae8a Decode state.db content sentinel in the WebUI message projection (#7492)
Decode the state.db content sentinel in the WebUI message projection so structured/multimodal content renders instead of a placeholder, with an out-of-band tuple identity that no scalar can forge and a per-call ContextVar memo for the canonical serialisation.

Co-authored-by: totalitarian <totalitarian@users.noreply.github.com>
2026-09-15 17:12:45 -07:00
nesquena-hermesandn 294577a1d1 Release: native composer field-sizing with fallback parity (#6760) (#7589)
Co-authored-by: n <a@n>
2026-09-15 16:56:11 -07:00
starship-sandstarship-s 2afcdf2676 perf: use native composer sizing where supported (#6760)
Use native CSS field-sizing for composer auto-grow where supported, keeping the JS path as fallback. Includes a fallback-parity fix so an empty composer holds its resting height instead of measuring the placeholder.

Co-authored-by: starship-s <starship-s@users.noreply.github.com>
2026-09-15 16:51:24 -07:00
nesquena-hermesandn 8fa1244279 Release: nest delegated subagent sessions under their messaging parent (#7263) (#7588)
Co-authored-by: n <a@n>
2026-09-15 16:36:51 -07:00
Kenanandlowlandsheperd b56dd423e7 fix(sidebar): nest delegated subagents under messaging parents (#7263)
Nest delegated subagent sessions under the parent conversation that spawned them instead of forcing them to the sidebar top level.

Co-authored-by: lowlandsheperd <lowlandsheperd@users.noreply.github.com>
2026-09-15 16:31:54 -07:00
webtecnicaandwebtecnica 68008e824b fix: classify Signal gateway sessions as messaging (#7572)
Classify Signal gateway sessions as messaging so they group, label, and list like other gateway platforms.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-15 16:16:57 -07:00
nesquena-hermesandn 572c84005f Release: preserve remote POSIX workspace paths across session, streaming and upload resolvers (#7168) (#7586)
Co-authored-by: n <a@n>
2026-09-15 15:44:39 -07:00
alfred-rootson b71c88f292 fix(workspace): preserve remote POSIX paths across Session & streaming callsites (#7131) (#7168)
Closes the remote POSIX workspace path preservation gap across Session and streaming. Profile-aware runtime path resolution.

Co-authored-by: alfred-rootson <alfred.wojtasordzik@gmail.com>
2026-09-15 15:37:11 -07:00
94fd2da830 Release: semantic fingerprint for the Codex catalog dependency (#7556) (#7558)
* fix(config): semantic fingerprint for the Codex catalog dependency (#7540)

* fix(#7556): guard the recursive codex-cache strip against RecursionError

Codex found a valid-but-deeply-nested models_cache.json crashes
_codex_models_cache_fingerprint with RecursionError (the recursive strip ran
outside the fallback try), 500ing /api/models. Move _strip_volatile_codex_cache_fields
inside the existing encode try/except so a transform failure degrades to the
stat-based fingerprint. +1 mutation-proven regression (forces the strip to raise;
fails with RecursionError if the guard is reverted).

* docs(changelog): stamp #7556 (semantic fingerprint for Codex catalog dependency)

---------

Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-09-13 17:39:39 +00:00
nesquena-hermesandnesquena-hermes e36f773891 test(#7510): clean up ruff F401/F841 in the raw-pending-approval test (#7552)
Follow-up to #7551: the merged test file carried an unused `route_approvals as ra`
import and two unused `entry =` assignments (advisory lint gate was red, though not
a required check). Drop them; the two `entry` bindings that are asserted on are kept.

Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-09-13 03:16:48 +00:00
e2ce040f1f Release: actionable ids for raw id-less pending approvals (#7510) (#7551)
* fix(approvals): mint actionable ids for raw id-less local pending entries

A raw id-less approval entry in _pending (agent-side _pending_result drops
{command, pattern_key, pattern_keys, description} verbatim when no gateway
notifier is registered for the session) is unactionable by contract: the
frontend owner-capture requires an approval_id, so the approval card can
neither be approved nor dismissed - the 1.5s poll re-serves it forever.

reconcile_gateway_pending_mirror_locked already mints gwlocal:<token> ids
for id-less _gateway_queues producers; extend the same contract to raw
non-mirror _pending entries with a stable gwlocal-mirrorless:<uuid> id
minted once onto the stored entry dict, so the id persists across polls
and frontend dismiss markers stick.

The exact-id respond path needs no further change: once the entry carries
an id, _resolve_approval_legacy pops it (also clearing the stale
stream-pointer 409 guard), stale explicit ids still fail closed (#527),
and the yolo Skip-all release drains it.

* docs(changelog): stamp #7510 (actionable ids for raw id-less pending approvals)

---------

Co-authored-by: CharlesMcquade <6466275+CharlesMcquade@users.noreply.github.com>
Co-authored-by: n <a@n>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-09-13 03:07:49 +00:00
e4d9b799e4 Release: manual title regeneration for queued-user transcripts (#7545) (#7547)
* fix(api): scan to first complete exchange for title regeneration (#7543)

_first_exchange_snippets() aborted at a second consecutive user message
before collecting assistant text, so transcripts opening with queued user
rows (channel-backed imports) hit the missing_exchange rejection and the
deterministic fallback got persisted with HTTP 200 — invisible in the UI
and unrecoverable by retrying.

Keep scanning until the first complete user+assistant pair instead, and
fall back to the last complete exchange when the opening user message
yields no user text (privacy scrubbing) instead of failing with
empty_user_message.

Tests: tests/test_issue7543_title_first_exchange.py (8 cases, incl. the
issue-pinned transcript shape and walker-mocked discriminator tests for
both fixes).

* test(api): drop local-machine notes and sys.path hack from 7543 test module

- Docstring described the change as a local-only uncommitted hotfix with
  machine-specific paths — wrong for an upstream PR and unsafe to publish.
- Drop the defensive REPO_ROOT sys.path insert; the repo conftest already
  puts the package root on sys.path (matches the neighboring test modules).

* fix(api): drop unreachable latest-exchange fallback from title regeneration

Greptile flagged the mocked-walker fallback test (P2) as fabricating
internal state, and an exhaustive enumeration of realistic transcript
shapes (5460 combinations, lengths 1-6) confirms it: both walkers share
the same text extraction and the forward walker scans every message, so
'first walker empty while latest walker yields a pair' is unreachable
from real transcripts. The fallback branch was dead code.

Remove it together with its two walker-mocked tests; the fixed first-
exchange walker alone resolves the reported issue shape. Replace the
mocked tests with an observable E2E error-path test through the real
walkers, and make the aux mock guard-faithful (empty assistant_text ->
missing_exchange per the generate_title_raw_via_aux contract) so the
fail-before proof actually exercises the endpoint contract.

* fix(#7543): scope first-exchange scan-past to the manual-regen path only

The full suite caught that broadening the SHARED _first_exchange_snippets also
activated the automatic in-stream background-title path for consecutive-user-row
transcripts, moving stream teardown from the synchronous stream_end branch to the
background-title-thread branch (test_issue3929 emits_done regressed: stream_end
was no longer in the synchronous queue). #7543 is specifically about the MANUAL
'Regenerate title' endpoint, so gate the scan-past behavior behind an opt-in
scan_past_consecutive_users flag (default False = master behavior) and pass True
only from generate_session_title_for_session. Auto in-stream path is now
byte-identical to master; added a default-behavior regression guard test.

* docs(changelog): stamp #7545 (manual title regen for queued-user transcripts)

---------

Co-authored-by: Raylands <Raylands@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-09-12 21:12:52 +00:00
b1286878a4 Release: models_discovered/discover_models live-catalog fix (#7409) (#7538)
* Apply reviewed change

* fix(models): honor discover_models opt-out + preserve discovered catalog on live-probe failure

Codex gate found two CORE edge cases in the #7404 fix:
1. discover_models:false must override models_discovered:true (explicit opt-out
   pins the configured allowlist; honor Agent string forms false/no/0).
2. a transient empty live-catalog probe dropped a discovered provider's models
   (cached empty ~24h); now falls back to configured discovered IDs merged with
   any static catalog (deduped, ordered first).

Mirrors the Hermes Agent canonical semantics (model_switch_providers._discover_flag,
model_switch._models_config_is_allowlist). Adds 3 mutation-proven regressions.

* docs(changelog): stamp #7409 (models_discovered/discover_models live catalog)

---------

Co-authored-by: evgenyponomarev <712802+evgenyponomarev@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-09-12 05:32:16 +00:00
nesquena-hermes 52bcfa3cb9 Merge pull request #7537 from nesquena/stage-relmgr-0912c
exp-v0.52.299 — graceful SIGINT shutdown (#7315)
2026-09-11 20:28:03 -07:00
nesquena-hermes 529488057e docs(changelog): stamp #7315 (SIGINT graceful shutdown) 2026-09-12 03:22:07 +00:00
nesquena-hermes c10bbd8e33 stage-fix(#7315): keep explicit signal.SIGTERM literal + add SIGINT within 749-line budget
Two tests pin opposite constraints on server.py: test_sprint10 requires <750 lines,
and test_issue5543 AST-asserts a signal.signal() call whose first arg literal is
signal.SIGTERM. The loop refactor satisfied the size budget but the loop var _sig
failed the AST literal check. Register both signals explicitly in one shared
try/except: keeps the SIGTERM literal, adds SIGINT (#7078), stays at 749 lines.
2026-09-12 03:03:10 +00:00
n c8a443454c Stage: PR #7315 — SIGTERM+SIGINT graceful shutdown by @aniruddhaadak80 (closes #7078) 2026-09-12 02:44:44 +00:00
nesquena-hermes a8f3933fd0 fix(server): register SIGTERM+SIGINT via one loop to stay under the 749-line budget
The contributor's SIGINT addition (#7315) pushed server.py to 754 lines, tripping
test_sprint10::test_server_py_under_750_lines (master was 749). Fold the two
signal.signal() try/except blocks into a single loop over (SIGTERM, SIGINT) so
SIGINT graceful shutdown lands (#7078) without growing the file past the enforced
split budget. Same _request_shutdown handler, same fail-safe on non-main-thread.
2026-09-12 02:29:28 +00:00
nesquena-hermes 00afa1eb31 Merge pull request #7536 from nesquena/stage-relmgr-0912a
exp-v0.52.298 — static MIME types (#7372) + upload TOCTOU (#7337)
2026-09-11 18:49:42 -07:00
nesquena-hermes 8e9d41380a docs(changelog): stamp #7372 (static MIME types) + #7337 (upload TOCTOU) 2026-09-12 01:44:47 +00:00
Aniruddha Adak 21f134056d fix(server): register SIGINT handler for graceful shutdown (#7078)
ctl.sh start spawns the server from a non-interactive bash background job,
which causes the child to inherit SIGINT=SIG_IGN. The Stop server buttonsends SIGINT (os.kill(pid, signal.SIGINT)), which is silently ignored.

Register _request_shutdown for SIGINT in addition to SIGTERM so the
shutdown button works for ctl.sh-managed daemons.
2026-09-12 01:29:23 +00:00
n c196306fcb Stage: PR #7337 — chat-attachment upload anchored O_EXCL (TOCTOU) by @zicochaos 2026-09-12 01:26:51 +00:00
n cb5e1974d1 Stage: PR #7372 — serve static binaries with correct MIME types by @Oidaladn0 2026-09-12 01:26:51 +00:00
Sebastian B f7523f63a2 fix(api): create chat-attachment uploads with anchored O_EXCL (TOCTOU)
handle_upload deduped filenames with an exists() check and then wrote with
plain write_bytes — two concurrent uploads of the same name both passed the
check and the last writer silently won. Mirror the #3398 workspace-upload
hardening: O_CREAT|O_EXCL|O_NOFOLLOW anchored open from the session
attachment dir; a raced duplicate now 409s instead of overwriting and a
raced symlink cannot redirect the write.
2026-09-12 00:55:45 +00:00
Hermes c076fe9e09 fix: detect MIME types for static binaries 2026-09-12 00:55:23 +00:00
nesquena-hermesandnesquena-hermes 2ad70a2815 docs(changelog): stamp #7525 (restore cron delivery platforms) (#7527)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-09-11 05:32:02 +00:00
nesquena-hermesandnesquena-hermes f79325498e fix(cron): restore messaging platforms in the delivery picker (#7525)
* fix(cron): restore messaging platforms in the delivery picker

The Agent moved _KNOWN_DELIVERY_PLATFORMS from cron.scheduler to
cron.scheduler_delivery. _handle_cron_delivery_options imported the old
path inside a bare except and fell back to an empty frozenset, so the
import error was swallowed and GET /api/crons/delivery-options silently
returned only local and origin — telegram, discord, slack, feishu and
every other messaging platform disappeared from the cron delivery picker
with no error surfaced anywhere.

Resolve the authority across both module paths (current first, legacy
fallback) so the picker keeps working across Agent versions, and add a
regression test that pins the resolution order and cross-checks the
endpoint output against the resolved authority.

* test(cron): skip the authority-resolution test where the Agent is absent

CI runs agent-free (tests/browser_smoke.py), so neither cron.scheduler_delivery
nor cron.scheduler is importable there and the authority resolves empty. The
new regression test asserted non-empty and therefore failed shard 3 on both
Python versions. Skip instead: with no cron package there is no authority to
resolve and nothing meaningful to assert. The test still fails loudly wherever
the Agent IS installed, which is where a future relocation would bite.

* test(cron): register the new test in _AGENT_DEPENDENT_TESTS

Use the project's existing skip mechanism (tests/conftest.py) rather than an
ad-hoc inline pytest.skip, matching the five sibling delivery-options tests.
CI runs agent-free, so there is no cron package and no authority to resolve
there; with the Agent installed the test still asserts loudly, which is the
environment where a future relocation would actually bite.

Verified both directions: 6/6 pass with the Agent present, and forcing
AGENT_MODULES_AVAILABLE=False skips all 6 (was 5) as CI will.

---------

Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-09-10 22:30:57 -07:00
nesquena-hermesandnesquena-hermes 88b8e7450a docs(changelog): stamp #7519 + #7522 (auxiliary model picker fixes, thanks @webtecnica) (#7526)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-09-11 05:17:11 +00:00
nesquena-hermes b7d81ce89e Merge pull request #7522 from webtecnica/fix/7521-aux-picker-endpoint-error-groups
fix(settings): keep endpoint-error provider groups in the auxiliary model pickers (#7521)
2026-09-10 22:13:23 -07:00
nesquena-hermes c5315bda6a Merge branch 'master' into fix/7521-aux-picker-endpoint-error-groups 2026-09-10 22:07:56 -07:00
nesquena-hermes f6f344063d Merge pull request #7519 from webtecnica/fix/7486-aux-provider-preserve
fix(settings): preserve configured aux provider when absent from catalog (#7486)
2026-09-10 22:07:26 -07:00
nesquena-hermes 80eb26301c Merge branch 'master' into fix/7486-aux-provider-preserve 2026-09-10 22:03:09 -07:00
nesquena-hermesandnesquena-hermes ce483fc70f docs(changelog): stamp #7518 (OpenRouter z-ai/ namespace, thanks @webtecnica) (#7524)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-09-11 04:52:39 +00:00
nesquena-hermes 8cd1e544ed Merge pull request #7518 from webtecnica/fix/7514-openrouter-zai-namespace
fix(onboarding): use OpenRouter namespace z-ai/ for Z.AI models (#7514)
2026-09-10 21:50:59 -07:00
webtecnica 617881feab test: drive the auxiliary loader behaviorally instead of grepping panels.js (#7521) 2026-09-10 23:18:44 -03:00
webtecnica dc1d41db04 fix: keep endpoint-error provider groups in auxiliary model pickers (#7521) 2026-09-10 23:11:12 -03:00
webtecnica e34980d969 test(7514): pass strict=True to zip() in the namespace projection test
Fixes the B905 finding flagged by the forward ruff gate on #7518.
2026-09-10 22:39:18 -03:00
webtecnica 4f109aebbe fix(settings): preserve configured aux provider when absent from catalog (#7486) 2026-09-10 22:37:09 -03:00
webtecnica 4359ed4fee fix(onboarding): use OpenRouter namespace z-ai/ for Z.AI models (#7514) 2026-09-10 22:30:44 -03:00
nesquena-hermesandn b0f1536f87 Release exp-v0.52.294: add GLM-5.3 Flash to Z.AI catalog (#7323) (#7517)
Co-authored-by: n <a@n>
2026-09-10 20:10:08 +00:00
rh-idandnesquena-hermes 40a6a7696a feat: add GLM-5.3 Flash to Z.AI model list (#7323)
* feat: add GLM-5.3 Flash to Z.AI model list

Add glm-5.3-flash next to glm-5.3 in _PROVIDER_MODELS["zai"] and mirror
it in _FALLBACK_MODELS, following the existing glm-4.5 / glm-4.5-flash
sibling pattern. Reasoning-effort gating needs no change: the >= 5.2
version gate already classifies the id at the full effort tier.

Adds a regression suite asserting catalog presence, label, ordering,
fallback entry, and end-to-end propagation through
get_available_models().

* test: pin glm-5.3-flash sibling adjacency in zai provider catalog

Per PR review feedback (#7323): the provider-catalog ordering test now
asserts glm-5.3-flash immediately after glm-5.3 and glm-5.2 immediately
after the flash sibling, matching the adjacency the fallback test
already pins. The base<flash relative check is subsumed; flash<glm-5.1
is kept to pin 5.1 after the whole 5.3 generation.

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-09-10 13:08:49 -07:00
nesquena-hermesandn f4569ade88 Release exp-v0.52.293: extract _handle_session_get refactor (#7340) (#7516)
Co-authored-by: n <a@n>
2026-09-10 19:50:42 +00:00
zicochaosandnesquena-hermes a6b187fc84 refactor(routes): extract the 585-line GET /api/session handler to _handle_session_get (#7340)
* refactor(routes): extract the 585-line GET /api/session handler to _handle_session_get

Pure move out of handle_get's if/elif chain into a named module-level handler, matching the pattern the other ~200 routes already use. Body is verbatim (one-level dedent only); AST identity of the moved statements verified; no behavior change — the GET /api/session test slice passes identically before and after.

* test: repoint four source-window tests to _handle_session_get

The moved block's source-window extractors (chat_start synthesiser/read-only
pins, regressions runtime-streaming-state window, issue1436 load-path
markers, security-redaction inspect) now anchor on the extracted handler;
assertions unchanged.

* test: repoint messaging-session loader pin to _handle_session_get

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-09-10 12:49:14 -07:00
nesquena-hermesandn a71f6c6feb Release exp-v0.52.292: recovered-reasoning-card duplication fix (#7443) (#7513)
Co-authored-by: n <a@n>
2026-09-10 19:30:53 +00:00
656575a400 Fix geometric duplication of recovered reasoning-only rows at turn settle (refs #7388, independent of #7416) (#7443)
* Fix geometric duplication of recovered reasoning-only rows at turn settle

_restore_display_reasoning_metadata() re-inserted every display-only
reasoning row at its API-safe position in the new result. After context
compaction that position is past the end of result_messages, so the row
was appended after the new reply; _message_identity() is None for an
empty assistant row, so the display merge could not dedup it. Each
settled turn therefore re-added every existing empty
_recovered_from_run_journal anchor once more (1,1,2,4,8,... blocks).

Only restore a reasoning-only row in front of its own API-safe
successor; a compacted result with no aligned slot is skipped.

Refs #7388, #7416 (independent: that PR fixes linear journal replay).

* tests: keep docstrings within the 5-line block budget

* Prefer stable ids when aligning a reasoning row's successor at restore

The successor-alignment guard compared the historical successor with the
aligned current row via _message_identity(), which ignores stable ids and
returns equal keys for repeated prompts ("continue"). A distinct current
turn with the same text false-matched the historical successor, every
recovered anchor before it was re-inserted, and 64 anchors became 128
after one restore->merge settle (nesquena-hermes gate on ed2076eda).

Alignment is now stable-id-first: when either row carries an id, both
must carry the same id; _message_identity() is the fallback only when
neither row has one. Adds a production-order regression through
_settle_result_messages() with 64 anchors, stable ids and repeated
"continue" prompts (RED on ed2076eda and origin/master: 128 != 64), plus
a unit test pinning all four id-presence branches.

* Fail closed on stable-id identity: positive-int only, unique, no content fallback

_restore_display_reasoning_metadata() compared ids with raw equality, so a
persisted historical successor with id=True (or 1.0) matched a distinct
current row minted as id=1 and restored the reasoning card before the
wrong turn. Ids are now accepted only when type(value) is int and > 0,
both rows must carry a valid equal id whenever either has an id field,
and the id must be the sole owner in both API-safe projections.
_assign_stable_message_ids() re-mints invalid ids so they cannot survive
into a collision with minted integers.

* tests: explicit zip strict= for the ruff B905 gate

* Resolve the active-turn boundary before restoration; never carry historical ids/rows into it

A historical assistant successor with the same content as the distinct
current assistant ("Done." / "Done.") let _restore_reasoning_metadata copy
the historical stable id onto the current row; the uniqueness guard then
accepted that forged identity and 64 recovered cards settled to 128.

_active_turn_boundary() is decided once in _prepare_marker_clean_writeback
before any restore (token > strict context prefix > last matching user row
> 0) and threaded through _settle_result_messages into both restores.
Rows at/after it receive no historical id/metadata and no historical
reasoning-only row. Assistant-only results are boundary zero.

* Keep the pinned _restore_reasoning_metadata signature; boundary goes through _restore_reasoning_metadata_before_boundary

tests/test_sprint49.py pins the two-arg signature as source text, so the
boundary-aware core lives in _restore_reasoning_metadata_before_boundary
and the original name stays a thin no-boundary wrapper.

* Content-only prefix is not turn ownership: fail closed to boundary 0; sync route passes result/Agent turn authority

- _active_turn_boundary: a previous-context prefix matched by _message_identity
  (ids/timestamps ignored) only narrows where the current user may be; the
  boundary is the prompt-matching user AT OR AFTER that prefix, else 0.
- _handle_chat_sync: resolve authority via _resolve_active_turn_authority
  (result, agent) instead of hard-coding None.
- Regressions: three-row current-only collision (boundary 0, one 64 block),
  exact-authority + append-only controls, ID-less-successor settle/save/load
  pinned at the persisted-display boundary, sync-route authority pass-through.

---------

Co-authored-by: carlotestor <carlotestor@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-09-10 12:29:33 -07:00
nesquena-hermes 71689c6ce7 Merge pull request #7505 from nesquena/release/pr7398-keycmd-cache
Release: preserve agent cache for key_cmd credentials (#7398)
2026-09-09 22:49:15 -07:00
n a284d478d4 docs: credit key_cmd cache fix (#7398) 2026-09-10 05:43:14 +00:00
n f08c8426e1 test: keep key_cmd laziness assertion at construction boundary 2026-09-10 04:48:14 +00:00
n f74b8a7c86 Merge remote-tracking branch 'origin/pr/7398-head' into release/exp289-pr7398 2026-09-10 04:41:05 +00:00
Arush Khasru b39bf1a113 Fix key_cmd credentials in agent cache signatures 2026-09-10 04:40:26 +00:00
394ab31eea release: #7351 bind turn workspace as agent cwd (#7504)
* rebase #7351 onto master (bind turn workspace as agent session cwd)

Co-authored-by: samfoy <samfoy@users.noreply.github.com>

* fix test: use deployed agent.runtime_cwd API (_SESSION_CWD.get) not nonexistent _session_cwd_override

The two failing assertions referenced agent.runtime_cwd._session_cwd_override(), a
helper not present in the installed Agent (which exposes _SESSION_CWD, resolve_agent_cwd,
set/clear_session_cwd). The production binder is correct and verified: binding workspace=/tmp
sets _SESSION_CWD.get()=='/tmp' and resolve_agent_cwd()=='/tmp', reset restores _UNSET.
Line 343 was redundant with the surrounding assertions; line 386 now reads _SESSION_CWD.get().

Co-authored-by: samfoy <samfoy@users.noreply.github.com>

* docs(changelog): stamp #7351 turn workspace cwd binding

---------

Co-authored-by: n <a@n>
Co-authored-by: samfoy <samfoy@users.noreply.github.com>
2026-09-10 03:40:45 +00:00
47e394895d release: #7160 stale-runtime report + hardened Agent-update marker read (#7502)
* rebase #7160 onto master (report unverified Agent updates, no auto-restart)

Co-authored-by: hrdwdmrbl <hrdwdmrbl@users.noreply.github.com>

* harden agent-update marker read: no-follow, fstat regular-file, bounded (fixes gate CORE)

The Agent update marker is attacker-adjacent shared state. _read_live_agent_update()
used Path.read_text(), which follows symlinks, blocks forever on a FIFO, and reads an
unbounded regular file — a rejected/stale-runtime request could hang or OOM. Now opens
O_RDONLY|O_NONBLOCK|O_NOFOLLOW, fstat-verifies a small regular file, and reads a bounded
max, classifying anything else as 'unknown' (fail-closed). Adds FIFO/oversized/symlink/
happy-path regression tests.

Co-authored-by: hrdwdmrbl <hrdwdmrbl@users.noreply.github.com>

* fix marker-read portability: gate os.open fast-path on O_NONBLOCK+O_NOFOLLOW

os.O_NONBLOCK is Unix-only; accessing it unconditionally raised AttributeError
on native Windows, breaking every stale-revision barrier (chat/compression/etc).
Now both flags resolve via getattr and the os.open() fast path is taken only
when BOTH are available (_MARKER_SAFE_OPEN_AVAILABLE); otherwise fall back to an
lstat-only classification (absent vs unknown) so a fallback platform can never
symlink-traverse or crash. Adds a Windows-fallback regression test.

Co-authored-by: hrdwdmrbl <hrdwdmrbl@users.noreply.github.com>

* docs(changelog): stamp #7160 stale-runtime report + hardened marker read

---------

Co-authored-by: n <a@n>
Co-authored-by: hrdwdmrbl <hrdwdmrbl@users.noreply.github.com>
2026-09-10 02:21:04 +00:00
nesquena-hermes ae311e4377 Merge pull request #7500 from nesquena/release/exp-v0.52.288-bound-sidebar-cache
exp-v0.52.288: bound sidebar session-list cache
2026-09-09 17:23:57 -07:00
n a8f46e899e docs: credit bounded sidebar cache fix (#7489) 2026-09-10 00:09:11 +00:00
nandMartin Dell 3e616ad455 fix: harden bounded sidebar cache projection (#7489)
Preserve nested request isolation, make the response allowlist single-source,
and strip heavy fields from existing index rows after reconciliation.

Co-authored-by: Martin Dell <martindell@users.noreply.github.com>
2026-09-09 23:36:41 +00:00
n 989d7f3668 Merge remote-tracking branch 'origin/pr-7489' 2026-09-09 22:49:54 +00:00
Martin Dell 7150b8fb2b fix: bound sidebar session-list cache 2026-09-09 21:22:38 +00:00
cfdfe4db8f release: #7261 provider-native auxiliary model IDs (approved) (#7491)
* fix(settings): persist provider-native auxiliary model IDs

Co-authored-by: lowlandsheperd <lowlandsheperd@users.noreply.github.com>

* docs(changelog): stamp #7261 provider-native aux model IDs

---------

Co-authored-by: n <a@n>
Co-authored-by: lowlandsheperd <lowlandsheperd@users.noreply.github.com>
2026-09-09 20:49:47 +00:00
351f664021 release: #7317 workspace switcher stored order (approved) (#7490)
* fix: workspace selector follows Workspaces view order (drop alpha re-sort)

Co-authored-by: CharlesMcquade <CharlesMcquade@users.noreply.github.com>

* docs(changelog): stamp #7317 workspace switcher stored-order

---------

Co-authored-by: n <a@n>
Co-authored-by: CharlesMcquade <CharlesMcquade@users.noreply.github.com>
2026-09-09 20:38:57 +00:00
2477d992f4 release: lock-disciplined STREAMS reads (#7339) (#7488)
* fix(streaming): lock-disciplined STREAMS reads via peek_stream()

Writers pop STREAMS under STREAMS_LOCK, but four hot reads used bare
STREAMS.get(): the streaming worker start, the gateway run start, and the two
SSE attach paths in routes.py. Each is individually saved today by GIL-atomic
dict.get plus a None-guard, but the check-then-act windows (get → register,
get → subscribe) race teardown pops, and the bare pattern invites unprotected
copies. peek_stream() centralizes the locked read; all four sites migrate with
no behavioral change.

* test: dual-patch config.STREAMS alongside routes.STREAMS in SSE seam tests

peek_stream() reads the canonical api.config.STREAMS registry; tests that
inject a fake registry by patching only the routes-level re-export left the
helper reading the real (empty) registry and hung the SSE loop to the
pytest-timeout ceiling. Mirror the test_issue6623 pattern: patch both
bindings.

* docs(changelog): stamp #7339 lock-disciplined STREAMS reads

---------

Co-authored-by: Sebastian B <sb@wmc.sh>
Co-authored-by: n <a@n>
2026-09-09 20:22:36 +00:00
b1a5ade137 release: multimodal title-generation fix (#7306) (#7484)
* fix(title): preserve multimodal title generation

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna
Assisted-by: Claude Code:claude-opus-5

* docs(title): explain multimodal title generation

Assisted-by: Hermes Agent:gpt-6-astra

* docs(changelog): stamp #7306 multimodal title-generation fix

---------

Co-authored-by: starship-s <45587122+starship-s@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-09-09 18:50:13 +00:00
842b2ac471 release: composer footer-fit jitter fix (#7275) (#7483)
* fix(composer): stop footer fit from jittering the transcript during SSE

_fitComposerFooter() measured overflow by REMOVING the .cf-icons/.cf-burger
stage classes, reading scrollWidth, then adding them back. Between those two
steps the footer is laid out at full width: the composer grows a few px and
#messages loses the same amount of clientHeight, then gets it back on the
next frame.

For a pinned reader that is a visible up/down jitter of the whole transcript,
because the fit pass runs on every context-indicator update — which fires
continuously while an SSE turn is streaming.

Measured with a real Chromium probe instrumenting the scrollTop setter,
scroll events, ResizeObserver and MutationObserver:

  before: 4 non-JS backward jumps of exactly -8px per streamed turn,
          each one ~30ms after a `ctx-indicator-wrap` -> `composer-footer`
          class mutation pair
  after:  0 backward jumps over 5 streamed turns on 2 different sessions

Fix: freeze the footer's layout box (explicit height + visibility:hidden)
for the duration of the measurement, so the intermediate expanded geometry
is never committed to the screen, and restore the resolved stage in the same
task. The adaptive behaviour is unchanged — verified across 8 viewport
widths (1600 -> 420px), with the correct full/icons/burger stage at each and
no residual inline style left behind.

Note: the guard added in the tail-jitter work is what keeps the reader pinned
here; before it, the same footer jitter unpinned the reader instead, which
masked the oscillation behind a worse bug.

* test(composer): cover footer fit freeze across full/icons/burger stages

Add a behavioural node driver test for _fitComposerFooter() that runs the
actual function from static/ui.js against a small layout model where the
stage classes dictate the footer's natural height and the left cluster's
content width. Every class mutation, inline-style write and overflow
measurement commits a layout sample, so the recorded footer heights and
messages clientHeights are what a browser would have painted.

The matrix starts from each stage (full, icons, burger), forces each
overflow outcome, and runs with empty and caller-owned prior inline
styles. It asserts the resolved stage classes, that the footer border-box
height and messages clientHeight stay pinned while the ladder is probed
(and do not move at all in the steady state), that no intermediate stage
is ever painted, and that prior inline height/visibility are restored
verbatim. It also covers the zero-height skip and an exception during
measurement, and includes a control run with the freeze made ineffective
to prove the harness reports the original jitter. Against the pre-fix
ui.js the height, jitter, painted-stage and style-restoration assertions
fail.

* docs(changelog): stamp #7275 composer footer-fit jitter fix

* ci: re-trigger (playwright browser install hung on GH runners)

* ci: re-trigger (google chrome apt mirror hash-sum mismatch, external infra)

---------

Co-authored-by: Anthony Ruiz <ruizanthony@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-09-09 18:11:21 +00:00
nesquena-hermes b1792e64e7 Merge pull request #7480 from nesquena/stage/7270
Release exp-v0.52.282: canonical server entrypoint for Windows restart (#7270, @rodboev)
2026-09-08 23:38:46 -07:00
n 1607767408 changelog: #7270 canonical Windows restart entrypoint (@rodboev) 2026-09-09 06:32:16 +00:00
n f9443e17b2 stage: merge #7270 (canonical Windows restart entrypoint) 2026-09-09 06:14:20 +00:00
nesquena-hermes ec63149e81 Merge pull request #7479 from nesquena/stage/7336
Release exp-v0.52.281: reject malformed/non-object JSON bodies with clean 400s (#7336, @zicochaos)
2026-09-08 21:44:39 -07:00
n 58df481bc5 changelog: #7336 reject malformed/non-object JSON bodies (@zicochaos) 2026-09-09 04:40:08 +00:00
nandzicochaos 1e17127f49 fix: treat whitespace-only body as empty ({}) not a 400
Gate finding on #7336: read_body() rejected whitespace-only bodies that
worked on master (a body-optional DELETE/PATCH/PUT can send b' \n'), returning
400 Invalid JSON body. Normalize a fully-whitespace body to {} before json.loads,
matching the pre-#7336 lenient behavior for that shape. Adds a parametrized
regression test.

Co-authored-by: zicochaos <zicochaos@users.noreply.github.com>
2026-09-09 04:21:46 +00:00
n 1f57afd84c stage: merge #7336 (reject malformed JSON bodies) 2026-09-09 04:02:32 +00:00
nesquena-hermes f4fa4bf77d Merge pull request #7477 from nesquena/stage/7476
Release exp-v0.52.280: warn once per state.db missing source column (#7476, @rodrigogs)
2026-09-08 19:33:49 -07:00
n 547ca6bb5e changelog: #7476 warn once per state.db missing source column (@rodrigogs) 2026-09-09 02:28:40 +00:00
n 927c4fa1b6 stage: merge #7476 (warn once per state.db) 2026-09-09 02:10:36 +00:00
nesquena-hermes 0c6b07c52c Merge pull request #7475 from nesquena/stage/7474
Release exp-v0.52.279: publish conversation context for aux title client (#7474, @webtecnica)
2026-09-08 19:07:44 -07:00
n f129de62e1 changelog: #7474 publish conversation context for aux title client (@webtecnica) 2026-09-09 02:03:26 +00:00
n f990201f48 stage: merge #7474 (aux title context) 2026-09-09 01:45:11 +00:00
Rodrigo Gomes 12df57498a fix(sessions): warn once per state.db when 'source' column is missing
Observed: on a multi-profile install where one profile's state.db predates
the agent's `source` column, errors.log gained one identical line every
~15 s (276 lines in ~70 min):

  api.models WARNING agent session listing skipped: state.db at
  /data/hermes/profiles/planner/state.db has no 'source' column (older
  hermes-agent?). Agent sessions unavailable. Upgrade hermes-agent to
  fix this.

Mechanism: api/agent_sessions.py read_importable_agent_session_rows()
runs `PRAGMA table_info(sessions)` and warns unconditionally when the
column is absent (api/agent_sessions.py:564 before this change).
api/models.py get_cli_sessions(all_profiles=True) calls it for every
profile on every sidebar poll, behind only the 5 s
_CLI_SESSIONS_CACHE_TTL_SECONDS cache, so the same DB re-emits the same
warning on every cache miss. The condition is a property of the DB
file, not of the poll.

Fix: a module-level `_SOURCE_COLUMN_WARNED_DB_PATHS: set[str]` keyed by
the resolved db_path. The warning fires the first time a given path is
seen in the process and is suppressed after that; a different path
warns on its own. Message text is unchanged. Process-lifetime only: a
restart warns again, which is intended (the line remains the operator's
cue that the agent still needs upgrading). Plain set mutation under the
GIL is sufficient; a duplicate line from two racing first calls is
harmless.

Tests (tests/test_issue634.py, real pre-`source` sqlite DBs under
tmp_path, asserting on captured log records):
- same db_path called twice -> exactly one WARNING (failed before:
  `assert 2 == 1`)
- two db_paths, then the first again -> exactly two WARNINGs, one per
  path (failed before: `assert 3 == 2`)
The seven existing #634 source-shape tests still pass; neighbours
test_issue3762/3930/5455/1494 pass.

Deliberately not changed: the exception-path warning in
get_cli_sessions() (api/models.py), the read-only-open fallback warning,
and the 5 s cache TTL. The operator-side root cause (the agent
reconciler creating a pre-`source` state.db) is a separate upstream fix;
this only stops the repetition.
2026-09-08 22:38:07 -03:00
webtecnica ca653eca1d test: tolerate conversation_id kwarg in title mock 2026-09-08 22:15:46 -03:00
webtecnica 61f3cf7ed9 fix(title): publish conversation context for aux title calls
Title generation via the aux client ran without the Agent's ambient
conversation context, so relay targets that key on it (OpenCode Go)
rejected the completion call with HTTP 400.

Publish the webui session id as the Agent conversation context around
the aux completion (restoring it in a finally block) and pass it from
all three title entry paths: background update, background refresh and
the on-demand regenerate helper.

Refs #7470
2026-09-08 22:06:32 -03:00
nesquena-hermes 7d56234ea4 Merge pull request #7473 from nesquena/stage/7082
Release exp-v0.52.278: stop tearing live progress words (#7082, @ruizanthony)
2026-09-08 14:53:01 -07:00
n b52f5b1f0e changelog: #7082 stop tearing live progress words (@ruizanthony) 2026-09-08 21:48:18 +00:00
n 049a1ef3de stage: merge #7082 (stream word-tearing) 2026-09-08 21:16:39 +00:00
nesquena-hermes 1f2c7abef0 Merge pull request #7469 from nesquena/stage/7356
Release exp-v0.52.277: keep provider hint for slash-id picks (#7356, @Manny7717)
2026-09-08 11:45:02 -07:00
n 4717dc97f2 changelog: #7356 keep provider hint for slash-id picks (@Manny7717) 2026-09-08 18:38:41 +00:00
n ad620c3211 stage: merge #7356 (provider-hint slash-id) 2026-09-08 18:20:45 +00:00
nesquena-hermes 5efc635488 Merge pull request #7468 from nesquena/stage/7277
Release exp-v0.52.276: guard submitEdit against re-entry (#7277, @alanjds)
2026-09-08 10:46:47 -07:00
n 38462d7cbd changelog: #7277 guard submitEdit re-entry (data-loss prevention, @alanjds) 2026-09-08 17:40:31 +00:00
n 649083eb7f stage: merge #7277 (submitEdit re-entry guard) 2026-09-08 17:22:39 +00:00
nesquena-hermes 3fae64a9e1 Merge pull request #7466 from nesquena/stage/5888
Release exp-v0.52.275: read-only child session ordering (#5888, @igorroncevic)
2026-09-08 09:59:51 -07:00
n f765e8178d ci: re-trigger (test_warm_models_catalog_provenance flake — network-timing, unrelated to this sessions.js-only diff) 2026-09-08 16:53:34 +00:00
nandigorroncevic 6cac3244d7 changelog: #5888 read-only child session ordering (@igorroncevic)
Co-authored-by: igorroncevic <igorroncevic@users.noreply.github.com>
2026-09-08 16:44:00 +00:00
n a1f8c51e6b stage: #5888 read-only child session ordering (@igorroncevic) 2026-09-08 16:43:05 +00:00
nesquena-hermes 10104dd82b Merge pull request #7465 from nesquena/stage/7335
Release exp-v0.52.274: inline handler arg escaping (#7335, @zicochaos)
2026-09-08 09:32:01 -07:00
nandzicochaos 2d391310f8 changelog: #7335 jsArg inline-handler escaping (@zicochaos)
Co-authored-by: zicochaos <zicochaos@users.noreply.github.com>
2026-09-08 16:24:06 +00:00
n 990a92a18e stage: #7335 jsArg inline-handler escaping (@zicochaos) 2026-09-08 16:05:36 +00:00
nesquena-hermes 3599d1f896 Release: nested mixed ul/ol list hierarchy, iteratively serialized (#6716, @webtecnica) (#7463)
Release: nested mixed ul/ol list hierarchy, iteratively serialized (#6716, @webtecnica)
2026-09-08 08:22:50 -07:00
n de742f64ef docs(changelog): stamp #6716 nested mixed-list renderer + iterative serializer 2026-09-08 15:14:52 +00:00
n cedc2ffc22 Merge PR #6716 re-push for re-gate 2026-09-08 14:54:50 +00:00
nesquena-hermes be5c071750 Release: skill-not-found reply discloses truncation (#7428, @elight) (#7460)
Release: skill-not-found reply discloses truncation (#7428, @elight)
2026-09-07 23:56:12 -07:00
n 5df1470619 docs(changelog): stamp #7428 skill-not-found truncation disclosure 2026-09-08 06:48:29 +00:00
n 4311143c12 Merge PR #7428 for gate 2026-09-08 06:29:12 +00:00
nesquena-hermes 1ab000069c Release: port out-of-band normalizers to the transport API (#7402, @zicochaos) (#7459)
Release: port out-of-band normalizers to the transport API (#7402, @zicochaos)
2026-09-07 22:49:24 -07:00
n a946b616c8 docs(changelog): stamp #7402 transport-normalizer port 2026-09-08 05:41:51 +00:00
n 2a804ae38d Merge PR #7402 for gate 2026-09-08 05:23:11 +00:00
nesquena-hermes 4b47cfa12d Release: streaming echo-suffix fold O(n^2)->O(n) (#7288, @ruizanthony) (#7458)
Release: streaming echo-suffix fold O(n^2)->O(n) (#7288, @ruizanthony)
2026-09-07 22:05:28 -07:00
n 07f1a1405a docs(changelog): stamp #7288 streaming echo-fold perf fix 2026-09-08 04:51:10 +00:00
n edf6c6a686 Merge PR #7288 for re-gate on current master 2026-09-08 04:37:42 +00:00
webtecnica 2c0b4902d0 fix(renderer): emit nested lists iteratively to avoid stack overflow (#6700) 2026-09-08 01:16:01 -03:00
webtecnica 660d54af63 test(renderer): mirror nests same-marker child lists inside parent li (#6700)
greptile P2: the Python mirror's handle_ul/handle_ol closed the parent
<li> before opening the child <ul>/<ol> for same-marker nested lists
('- parent\n  - child'), so structural assertions passed without
validating renderMd()'s parent-child hierarchy. Child lists now render
inside the parent <li>, byte-matching renderMd() for the two-level case;
adds a mirror regression test asserting the exact nested structure.
2026-09-08 01:16:00 -03:00
webtecnica b810e5466f fix(renderer): preserve ordered flag on list items (#6700)
openItem() received ordered but dropped it from the item constructor, so
it.ordered was always undefined and every item looked unordered at the
task-list check. An ordered item whose text begins with [x] or [ ] (e.g.
'1. [x] shipped') was silently converted to a task icon and lost its
literal prefix. Task-list rendering now applies only to unordered items.
2026-09-08 01:16:00 -03:00
webtecnica 13dd97d60d fix(renderer): preserve nested mixed ul/ol list hierarchy (#6700) 2026-09-08 01:16:00 -03:00
nesquena-hermes 641878bd88 Release: extension archive dotfile allowlist (#6620, @asorourx) (#7457)
Release: extension archive dotfile allowlist (#6620, @asorourx)
2026-09-07 20:27:37 -07:00
n d53587d77a docs(changelog): stamp #6620 extension archive dotfile allowlist 2026-09-08 03:14:02 +00:00
n e69d681e82 Merge PR #6620 for exact-head re-gate 2026-09-08 03:01:00 +00:00
nesquena-hermes 6e6893f3ae Phase 1 Batch 2: passkey-behind-proxy + opt-in pipe trusted-groups (#7456)
Phase 1 Batch 2: passkey-behind-proxy (#6772) + opt-in pipe trusted-groups (#7331)
2026-09-07 19:21:27 -07:00
n 286c5e3336 docs(changelog): #7331 pipe separator is opt-in 2026-09-08 01:53:18 +00:00
n e7841a5988 gate fixes: opt-in pipe separator (#7331) + IPv6-safe Origin (#6772)
Two Codex adversarial-security-gate findings on the batch:

1. #7331 CORE (backward-compat break): making '|' an unconditional trusted-groups
   separator would re-split an existing deployment whose group NAME legitimately
   contains a literal '|' (e.g. 'unprivileged|admins' read as one group), silently
   changing its profile binding. Pipe splitting is now OPT-IN via
   HERMES_WEBUI_TRUSTED_GROUPS_PIPE_SEPARATOR; comma/newline stay the default.
   Tests: literal-pipe preserved by default, pipe splits only with the flag set.

2. #6772 follow-up (my own hardening introduced it): rebuilding the origin from
   parsed parts dropped the brackets around an IPv6 host, so 'http://[::1]:8787'
   became 'http://::1:8787' and never matched the browser's bracketed
   clientDataJSON.origin -> passkey login over IPv6 would fail. Re-bracket IPv6
   literals. Added rp_context regression tests (Origin-first, port preserved,
   IPv6 bracketed, malformed->Host fallback, userinfo stripped).
2026-09-08 01:53:01 +00:00
n 2e62c4f896 docs(changelog): stamp Phase-1 batch 2 security fixes (#6772, #7331) 2026-09-08 01:31:45 +00:00
n 289b539bc4 harden(passkeys): accept only a well-formed http(s) Origin for RPID derivation
Follow-up on #6772. Preferring Origin over Host for the WebAuthn RPID is correct
(the RPID must match the page origin, not the backend's), and it is not a trust
decision: the authenticator only releases a credential whose RPID is a registrable
suffix of the real page origin, clientDataJSON.origin must equal the stored origin,
rpIdHash must match the stored RPID, and the assertion signature is verified against
the stored credential public key.

Still, only accept a syntactically valid http(s) Origin carrying a hostname, and
rebuild the stored origin from the parsed scheme/host/port instead of echoing the raw
header — so a malformed, non-http (javascript:/file:), or userinfo-bearing Origin
falls through to the Host-derived path rather than being stored verbatim.
2026-09-08 01:31:26 +00:00
n 0f76602790 Merge PR #7331 2026-09-08 01:30:36 +00:00
n 1e00ef131b Merge PR #6772 2026-09-08 01:30:35 +00:00
nesquena-hermes 49c6f98f82 Phase 1 Batch 1: seven low-risk contributor fixes (#7455)
Phase 1 Batch 1: seven low-risk contributor fixes (+2 gate follow-ups)
2026-09-07 18:28:12 -07:00
n 7ee8e43392 docs(changelog): stamp Phase-1 batch 1 (#7427 #7207 #7447 #7385 #6768 #7430 #6330) 2026-09-08 00:59:17 +00:00
n a8d10887df fix(picker): don't let a separator-only search match every model
Codex gate finding on #7385: the new fold-matching (#7228) normalizes the search
term by stripping [\s._-], so a separator-only query ('.', '-', '_') folds to '' and
String.includes('') is always true — every model matched. Verified: searching '.'
returned all models, including ids with no dot.

Guard the folded clauses with a non-empty foldTerm, keeping the literal
name/id includes() checks so '-' still matches ids that really contain one.
Real searches ('ox alpha' / 'ox-alpha' / 'sonnet') are unchanged.
2026-09-08 00:58:05 +00:00
n 150c766875 fix(transcript): mount the render window when jumping to session start mid-stream
Codex gate finding on #7427: moving _scheduleMessageVirtualizedRender() below the
_freshProgrammaticScrollActive() guard (the idle re-render-loop fix) removed the ONLY
thing that mounted the new render window on the jumpToSessionStart() streaming path,
which deliberately skips renderMessages(). Reproduced in Chromium: 160 rows produced a
14,827px spacer and zero visible rows (blank transcript) where master mounted correctly.

Schedule the virtualized render explicitly in the streaming branch, after invalidating
_messageVirtualWindowKey, preserving #7427's listener guard and its idle-loop fix.
2026-09-08 00:26:54 +00:00
n 94b70ce8b8 Merge PR #7385 2026-09-08 00:16:17 +00:00
n d52596a269 Merge PR #7427 2026-09-08 00:16:17 +00:00
n f8ade25db9 Merge PR #7207 2026-09-08 00:16:17 +00:00
n e025f0aca6 Merge PR #7447 2026-09-08 00:16:17 +00:00
n d510c2f508 Merge PR #6768 2026-09-08 00:16:17 +00:00
n 6d248ecd0b Merge PR #7430 2026-09-08 00:16:17 +00:00
n df9187e5f8 Merge PR #6330 2026-09-08 00:16:16 +00:00
nesquena-hermes 51d4685b5f Fix WebUI ↔ agent v2026.9.7 API skew (session-title regression + agent-coupled tests) (#7454)
Fix WebUI ↔ agent v2026.9.7 API skew (silent session-title regression + agent-coupled tests)
2026-09-07 16:47:16 -07:00
n bee7b8a9e4 fix(state-sync): repoint to renamed agent SessionDB API + fix agent-coupled tests
The Hermes Agent's state-module refactor (agent commit 53db597201, released
v2026.9.7) renamed SessionDB.set_auto_title_if_empty -> set_auto_title(
session_id, title, *, source) and (agent commit 14791b4d4e) dropped the
tools.approval._ApprovalEntry back-compat facade (class moved, unchanged, to
tools.approval_gateway_wait).

Production impact (real, silent): api.state_sync.sync_session_title called the
removed set_auto_title_if_empty, whose AttributeError was swallowed by the
surrounding `except Exception: logger.debug(...)`, so WebUI auto-generated
session titles stopped persisting to state.db on agent >= v2026.9.7 — hermes
sessions list went blank again (the exact #6892 regression). Now calls
set_auto_title with LLM provenance, which preserves the identical semantics
(only fills NULL/lower-authority title, returns False untouched on a manual
rename, ValueError -> get_next_title_in_lineage #N de-dup).

Test fixes (agent-coupled; skip on GitHub CI which installs no agent, so they
only surface on an agent-equipped box like our gate):
 1. tools.approval._ApprovalEntry -> idempotent autouse conftest compat shim
    backfilling the class from its post-split location only when the installed
    agent lacks the dropped facade. No production code imports it.
 2. test_goal_command_webui::test_profile_goal_falls_back_when_session_db_path_is_frozen
    was missing the monkeypatch.syspath_prepend(...) that every sibling test in
    the file has, so its bare `import hermes_state` failed under full-suite
    ordering after a predecessor stripped the agent dir from sys.path. Added it.

Verified: authoritative 5-shard CI-equivalent split all green (0 failed);
each failure proven pre-existing on clean master (not this diff).
2026-09-07 23:34:45 +00:00
nesquena-hermes bd62772ca5 Release: Docker env-reload secret-truncation fix (#7433, @webtecnica)
Release: Docker env-reload secret-truncation fix (#7433, @webtecnica)
2026-09-07 15:56:36 -07:00
n 37bb44173e docs(changelog): stamp #7433 Docker env-reload secret fix 2026-09-07 22:19:51 +00:00
nandwebtecnica 95107ceb94 fix(docker): preserve trailing = in env values on su reload (#7432)
docker_init.bash load_env() reloaded the save_env() snapshot after the
su privilege drop with `while IFS='=' read -r key value`, and Bash
silently drops a single trailing `=` from the last field. Any secret
ending in one `=` (openssl rand -base64 32 output is 44 chars ending in
`=`) reloaded one character short, breaking shared-secret auth and
HERMES_WEBUI_PASSWORD matching in the container.

Read each line whole (`IFS= read -r line`) then split at the FIRST `=`
only (key=${line%%=*}, value=${line#*=}), preserving embedded and
trailing `=`. Blank and comment lines are skipped; lines without `=`
keep prior empty-value behavior.

Closes #7432.

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-07 22:19:51 +00:00
SuperAgent BEAM3 a452665dd5 test: add parameterized separator table for trusted-groups header parsing
Per review feedback on #7331: cover comma, pipe, newline, mixed
separators, repeated/adjacent separators, and surrounding whitespace
directly against _trusted_groups_header_value(), in addition to the
existing end-to-end bound_profile regression test.
2026-09-07 07:28:02 +00:00
MuhammadUsamaMX 77ee0b07f9 worklog: stop reading a placeholder lifecycle row as a compression start
routes.py appends a 'live anchor shell' row whenever a stream has events
but no visible rows projected yet:

    row_id:            lifecycle:<stream>:running
    role:              lifecycle
    status:            running
    text:              Working
    source_event_type: runtime_journal_snapshot

_sourceEventTypeForSnapshotAnchorRow skips the source_event_type branch
for exactly that value, falls through to the role check, and the bare
phase==='running' arm then classified the shell as 'compressing'. Every
session that hit the path painted a 'Compressing context' divider while
nothing was compressing - including brand-new sessions on their first
message.

It also persists. The classification is written back into the session's
anchor_activity_scenes as source_event_type: 'compressing', which the
source branch above then returns directly on the next load, so the
phantom divider survives a reload. Measured on this host: 46 such rows
across 24 sessions, every one carrying the text 'Working' and not one a
real compression notice.

Require an explicit cue instead - phase==='compressing' or one of the
compression text markers already listed below it. The function's own
comment asked for this ('do not default every lifecycle row to a running
compress divider'); the bare running arm did the opposite.

tests/test_compression_phantom_barrier.py: 8 passed.
2026-09-06 23:16:18 +05:00
evgenyponomarev aa99f6f7da Apply reviewed change 2026-09-05 01:01:47 +03:00
LeonardoLGDS e60323d5cd Fix #6799: perf(transcript): virtualization ON + long session = endless scroll to re-render loop (~5-6 re-renders/sec idle), CPU 100%+, text selection broken 2026-09-04 15:56:47 +00:00
Evan Light d7ed506d21 fix(skills): report the total when skill-not-found truncates available_skills (#7426)
The not-found reply capped available_skills at 20 with no total and no
truncation marker, so a caller with more skills than the cap was handed a
partial list that reads as the complete set. A skill sorting past the cut
looked uninstalled.

Keep the bound; stop it lying. The payload now carries total_skills and
available_skills_truncated, and the hint names how many of how many were
shown, matching the (list, truncated) shape already used for capped commit
subject lists in api/updates.py.
2026-09-04 15:40:17 +00:00
Anthony Ruiz b28559f62c test(streaming): exercise an echo that really falls outside the search window
The case labelled as an echo further back than the window still ended the
buffer with the suffix, so it matched regardless of the window. Replace it
with a buffer whose whitespace-padded echo spans more than the default
window, and add a dedicated test proving the same echo is rejected with the
default window and found once the window covers it, for both the optimised
scan and the reference oracle.
2026-09-03 09:41:27 +00:00
Sebastian B 6a0defbc2f test: agent-less CI compat for transport-port tests
The new normalize test now stubs agent.anthropic_adapter in sys.modules
(CI has no agent package — real import raised ModuleNotFoundError), and the
issue-1013 codex fake ports to the v0.21 _get_transport contract instead of
the deleted _normalize_codex_response method.
2026-09-02 15:29:04 +01:00
Sebastian B da5703d501 fix(agent-v0.21): port out-of-band normalizers to the transport API
hermes-agent v0.21 removed agent.anthropic_adapter.normalize_anthropic_response and the AIAgent._normalize_codex_response method; the sanctioned out-of-band contract is now agent._get_transport(<mode>).normalize_response(...) returning NormalizedResponse (.content). Port the anthropic_messages and codex_responses branches of generate_title_raw_via_agent (api/streaming.py) and _agent_text_completion (api/routes.py); on agent v0.21 those branches currently die on import/attribute errors and silently fall back, degrading title generation and handoff summaries. build_anthropic_kwargs, _anthropic_messages_create, and _run_codex_stream are unchanged in v0.21 and stay as-is.
2026-09-02 15:07:17 +01:00
Sebastian B ff32048ac5 fix: restore uv.lock to master's deliberate 3-line stub
Review ask on #7335: the +232-line lockfile regeneration was a local uv
side effect swept in by a broad 'git add -A' — this is a JS-only escaping
fix and the lockfile stub is intentional on master.
2026-09-02 14:02:20 +01:00
webtecnica ec4bac9fc9 fix: model picker search matches OpenRouter display names (#7228) 2026-08-31 15:29:28 -03:00
Alan Justino da Silva 10e0a5f214 Merge branch 'master' into claude/webui-edit-double-submit 2026-08-31 11:17:46 -03:00
Manny7717 1a31003e9e fix(config): gate custom provider hints on unique configured entries
Named custom providers are only routable when the slug resolves to a real,
unique custom_providers[] entry. A stale session provider such as
'custom:missing' must not be minted into an @custom:missing:... route -
resolve_model_provider() would take the named-provider lane and find no
matching endpoint. _unique_custom_provider_entry() returns None for unknown
slugs and raises AmbiguousCustomProviderError for collisions, matching the
point-of-return guard used by resolve_model_provider.

Addresses maintainer review on #7356. Adds two regression tests: unknown
named custom slug stays bare, ambiguous normalized slug fails closed.
2026-08-29 09:12:17 -05:00
Manny7717 e305d7154d style: ruff-format new #7333 regression test file 2026-08-28 20:30:57 -05:00
Manny7717 da367cc692 fix(config): keep provider hint for slash ids when session provider differs
model_with_provider_context() passed any slash-bearing model ID through
bare when the session provider was not OpenRouter/ACP/plugin and not
listed under config.yaml -> providers:. That dropped the hint for portal
providers (nous, nvidia, opencode-zen, opencode-go) and named custom
providers (custom:<slug>), so resolve_model_provider() fell back to the
profile default provider + base_url and sent a vendor-scoped id to the
wrong endpoint (HTTP 404, e.g. upstage/solar-pro4:free to api.x.ai).

Emit an explicit @provider:model hint for any known routable provider
that differs from the configured default; keep the bare passthrough for
unknown/ambiguous provider slugs so custom/proxy base_url routing stays
in charge. Closes #7333.
2026-08-28 20:29:05 -05:00
Anthony Ruiz c6a2c1b7d2 perf(streaming): fold the echo search window once instead of per cut index
`_strip_compact_echo_suffix` probed every candidate cut index and re-folded
the whole remaining tail for each probe, so the cost grew with
`window x suffix` rather than with the data actually inspected. It runs up to
four times per interim assistant message, on the reasoning buffer, the shared
stream mirror and up to two per-message segments.

The practical effect is worst on the longest message of a turn: 297ms for a
200-character message but 5.6s for a 6000-character one, measured on the live
process. That time is spent holding the GIL, so a single long final message
also stalls every other stream served by the same worker. Profiling a live
turn attributed ~5% of total CPU to this one helper.

Fold the window once, then walk backwards across exactly as many non-space
characters as the folded suffix holds. `str.isspace` and the `\s` pattern used
by `_compact_for_echo_compare` agree on all 1114112 Unicode code points, so
the folded view and the walk stay consistent.

Behaviour is unchanged, including the leftmost-cut tie-break on repeated text
and the bounded search window. The new test pins equivalence against a naive
reference oracle over ~2020 cases (exact, partial and absent echoes, mixed
whitespace, unicode, empty and None inputs, repetitions, out-of-window
echoes), asserts the fold no longer runs once per cut index, and caps the
runtime of a long-conclusion call.

Measured: 200 chars 297ms -> 0.149ms, 6000 chars 5645ms -> 0.720ms.
2026-08-27 23:58:18 +00:00
Anthony Ruiz 69e4fa3b98 fix(stream): adopt parser-owned fade row instead of recloning it
After rewind/cache-miss the keyed visible row stayed distinct from the
persistent parser candidate, so every later growth frame deep-cloned the
parser DOM into the old row. That cut off the previous tail word's
~620ms fade and broke word-node identity.

Promote the parser-owned candidate into the existing row's DOM position
once, mute the already-rendered prefix on the real parser body, and
return the adopted node so later keyed renders hit existing === node.

Regression: two post-transition growth frames keep the first new tail
span as the same .is-new node until an explicit animationend.
2026-08-27 23:58:14 +00:00
Anthony Ruiz 9388c6495b fix(stream): keep parser DOM as fade-row authority after rebuild
Cursorless markdown rebuilds cloned the live parser tree, then later
source growth appended raw bytes through _appendTransparentFadeText.
Incomplete emphasis (Hello ** / Hello **world**) became literal
**world** while the candidate already had <strong>world</strong>.

Parser-owned candidates now keep reconciling from the parsed DOM.
Pending and MEDIA tails stay on the parser. Rendered-prefix mute is
unchanged.
2026-08-27 23:58:14 +00:00
Anthony Ruiz 259631e46b fix(stream): resume cursor excludes parser pending tail on cursorless markdown rebuild
Greptile P1 follow-up on the fade-row rebuild: when a cursorless live
markdown row adopts the incremental node's cloned DOM, that DOM lags the
source text by the streaming parser's held-back tail (pending/text
buffers). Keeping the full source-length resume cursor made the next
reconciliation append nothing, so the pending characters stayed missing
from the live row until settlement.

Read the held-back length from the parser bound on the candidate body
(_smdBindParserIdentity) and trim the source-space cursor by exactly
that length, so the cursor matches the adopted content and the next
reconciliation re-appends the pending tail.

Regression: rebuild while the parser holds an unrendered tail, then
reconcile again — the pending text must appear in the live row before
settlement.
2026-08-27 23:58:14 +00:00
Anthony Ruiz 34f10d2545 fix(stream): mute rendered prefix and keep parsed markdown on fade-row rebuild
Review follow-ups for #7082:

- Maintainer should-fix: the no-cursor rebuild branch of
  _refreshTransparentFadeProseRow cleared the body and re-wrapped every
  word as .is-new, so the whole visible row dipped to opacity 0 and
  faded back (~620ms), and wholesale node replacement invited the
  one-time scroll-anchor bounce (#6257). Now the rendered text is
  snapshotted before the clear and the messages.js
  _streamFadeMuteRenderedPrefix idiom is re-applied after the rebuild,
  so only genuinely-new tail words animate. The helper is exported on
  window (__streamFadeMuteRenderedPrefix) the same way
  __anchorProseIncrementalNode is.

- Greptile P1: the same branch passed raw source text to
  _appendTransparentFadeText, flattening live markdown rows to literal
  syntax (links/emphasis/headings lost until settlement). The rebuild
  now adopts a deep clone of the candidate's parsed .msg-body (the
  incremental streaming-markdown output) when one exists, so the parsed
  DOM survives; the plain-text rebuild remains only as fallback for
  candidates without a parsed body. The resume cursor stays in source
  space when rendered text diverges from source (markdown), and shrinks
  to the rendered prefix for plain prose so the parser's held-back
  character is re-appended on the next delta instead of dropped.

Regression tests: rebuild must not re-animate the already-rendered
prefix (tail-only .is-new), and rebuild must keep parsed <a>/<strong>
DOM with no literal markdown in the visible text. The two existing
tests of this PR are unchanged and still pass.
2026-08-27 23:58:14 +00:00
Anthony Ruiz 59d2f295c8 fix(stream): stop tearing live progress words across block boundaries
Transparent-stream progress rows are built incrementally by
`_anchorProseIncrementalNode`, which streams source text into the vendored
streaming-markdown parser. That parser always holds the most recently written
character in its pending buffer, so a live row's rendered `.msg-body` text
trails its source text by one character.

`_refreshTransparentFadeProseRow` resumed appending from the
`data-stream-fade-text` cursor but fell back to `body.textContent` when the
attribute was absent — exactly the case for incrementally built rows, which
never set it. That fallback mixed two coordinate spaces (rendered text vs
source text), so the computed delta started one character early, in the middle
of a word. The delta was then appended as a sibling of the parser's open `<p>`,
and since `<p>` is a block box the tail rendered on its own line: a single word
split across two lines ("Ces deu" / "x fichiers").

Two independent fixes:

1. Resume strictly from the source-space cursor. When no cursor exists we
   cannot know how much source is already rendered, so rebuild the body
   instead of guessing a delta from rendered text.
2. Append into the trailing block element rather than the `.msg-body` root, so
   a continuation of the current sentence stays on the same line even if a
   cursor is ever stale.

Regression tests drive the real vendored parser to build a live row exactly as
production does, assert the parser really does lag its source, and fail if any
word is torn across a block boundary. The existing reconcile harness gains the
new helper dependency.
2026-08-27 23:58:14 +00:00
Sebastian B edaa255fa3 lint: drop unused json import in read_body validation test 2026-08-28 00:57:30 +01:00
Sebastian B 743b757a77 fix(api): reject malformed and non-object JSON bodies with clean 400s
read_body() silently turned malformed JSON into {} (clients then got a
misleading 'Missing field' 400) and passed non-dict JSON through to handlers
where body.get() raised AttributeError and surfaced as a generic 500.
Raise ValueError instead — handle_post already maps it to a clean 400 — and
give the PATCH/DELETE/PUT dispatchers the same mapping they lacked. Empty
bodies keep the documented {} default.
2026-08-28 00:49:50 +01:00
Sebastian B d4e102fd38 fix(ui): encode inline on* handler args with shared jsArg() (#3797 class)
esc() is HTML-only escaping; the browser decodes attribute entities BEFORE
parsing an inline handler, so onclick="fn('${esc(id)}')" let any id
containing a quote break out of the JS string literal. _kanbanJsArg solved
this for kanban dependencies only; promote it to jsArg() in ui.js and route
every inline-handler argument through it (panels.js cron/kanban/notes/
passkeys/gateway/checkpoints, sessions.js handoff hint, onboarding.js OAuth
device code). Adds a Playwright round-trip proof and a source guard against
reintroducing the esc()-only pattern.
2026-08-28 00:45:45 +01:00
SuperAgent BEAM3 bb5347e1cd fix: accept pipe-separated trusted-groups header values
Authentik's proxy provider on this deployment joins multiple group
names with "|" instead of "," (e.g. "admins|developpeur"). The
comma-only split in _trusted_groups_header_value() treated that whole
string as a single unmatched group name, silently dropping the
session to the unbound "default" profile even for identities that
are legitimate members of a mapped group. Single-group identities were
never affected since there was no separator to parse.

Accept comma, pipe, and newline as equivalent separators so the parser
works regardless of which the outpost happens to send.
2026-08-26 08:35:10 +00:00
nesquena-hermes e168b67e42 Merge pull request #7307 from nesquena/release/exp-v0.52.264
Release exp-v0.52.264: fast regenerate via bounded sidecar-anchored tail read (#7204, @webtecnica)
2026-08-25 18:03:26 -07:00
n ebc631f53a Release exp-v0.52.264: fast regenerate via bounded sidecar-anchored tail read (#7204, @webtecnica)
Adds the CHANGELOG entry for #7204 (closes #6826). Release-metadata only;
the code change ships via the merge of #7204 (contributor webtecnica).
2026-08-26 01:01:19 +00:00
nesquena-hermes 134327aa41 Merge pull request #7204 from webtecnica/fix/6826-issue
Bounded sidecar-anchored tail read for regenerate (closes #6826). Thanks @webtecnica — six rounds of adversarial re-gate to convergence.
2026-08-25 18:00:52 -07:00
nandwebtecnica 3fac79bb59 lint: drop redundant unused import in regeneration test (#6826)
Removes the local re-import of api.streaming._session_payload_with_full_messages
inside test_in_tail_duplicate_guard_refuses_bounded_and_full_revision_accepted;
it is already imported at module level (line 17) and unused in the function body,
tripping F401 on CI's hosted lint. Mechanical maintainer fix per auto-fix policy.

Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
2026-08-26 00:24:07 +00:00
webtecnica b36fbdd633 fix(regeneration): refuse bounded tail on in-tail duplicate keys + real WAL interleaving test (#6826 r5) 2026-08-26 00:24:07 +00:00
webtecnica 7abd4773ca fix(regeneration): snapshot reuses canonical projection+ordering; WAL data_version guard (#6826 r4) 2026-08-26 00:24:07 +00:00
webtecnica b33aff9d09 fix(regeneration): occurrence-count guard + single-snapshot (no TOCTOU) for bounded tail read (#6826 r3 round-3) 2026-08-26 00:24:07 +00:00
webtecnica 40ff74b6fc fix(regeneration): bounded tail read only when skipped state.db prefix is provably identical; compression-anchor coverage (#6826 r3) 2026-08-26 00:24:07 +00:00
webtecnica 3672e0038b fix(regeneration): bounded sidecar-anchored tail read instead of full transcript refetch (#6826) 2026-08-26 00:24:07 +00:00
webtecnica 8c9f505fab fix(ui): avoid full transcript fetch on regenerate (#6826) 2026-08-26 00:24:07 +00:00
Claude 5c21e2db90 Merge remote-tracking branch 'origin/master' into claude/webui-edit-double-submit 2026-08-24 19:31:10 +00:00
Claude fc528113f8 fix(webui): guard submitEdit against re-entry (destructive double-submit)
Clicking "Send edit" a second time while the first is still working starts a
second full truncate. Observed in the wild on a laggy instance: seven concurrent
POST /api/session/truncate, 5.2s / 8.3s / 6.7s / 7.4s / 8.1s / 8.5s, from one
user clicking repeatedly because the UI had not yet acknowledged the first click.

submitEdit's only guard was S.busy — but send() does not set that until the LAST
line of the function, after two multi-second awaits:

    if(!S.session || S.busy) return;              // S.busy still false
    const absoluteKeepCount = _oldestIdx + msgIdx;
    await _ensureAllMessagesLoaded();             // ~6.4s on a long session
    await api('/api/session/truncate', ...);      // 5-8s under load
    await send();                                 // S.busy finally set

On a 2000-message session that is a 10s+ window in which every further click
runs the whole destructive path again.

The load is the visible symptom; the hazard is worse. absoluteKeepCount is
captured BEFORE the awaits deliberately — at click time _oldestIdx is the loaded
window's offset and msgIdx is window-relative, so the sum is the true absolute
index (pinned by test_submit_edit_captures_absolute_before_await, the #2184
pattern). But the first call's _ensureAllMessagesLoaded() sets _oldestIdx = 0.
A second call entering after that computes `0 + msgIdx` from a msgIdx that is
still window-relative — a far smaller keep_count. Truncating a long session to
that would delete most of its history.

Fix: a module-scope _submitEditInFlight flag, raised before the first await and
cleared in a finally. Guarded at the function rather than at the click handler
because the edit is also submittable with Enter (the keydown handler clicks the
button), so a handler-only guard would leave that path open, and because a future
caller would silently reopen the hole. The finally means an early return or a
throw cannot wedge editing off for the rest of the page's life.

Deliberately NOT changed:
- The pre-await capture stays where it is. Recomputing keep_count after
  _ensureAllMessagesLoaded() looks like the obvious fix and is wrong: it would
  pair _oldestIdx=0 with a window-relative msgIdx and break #2184 outright.
- regenerateResponse has the same S.busy-only shape, but no longer client-side
  truncates (it delegates to the atomic startRegeneration) and validates the
  clicked index against latestAssistantIndex, so a double click is refused
  server-side rather than being destructive. Left alone to keep this diff to the
  fault.
- No UI feedback added. Extra clicks are now harmless, but the button still does
  not visibly acknowledge the first one, which is what provoked the clicking.
  Worth a follow-up; it is a UX change, not this correctness fix.

Not reproduced end-to-end: this environment cannot run a real agent turn, so the
seven-truncate storm comes from the reporter's network log, not from a local
repro. The guard is verified by source-level test only.

test_submit_edit_is_guarded_against_reentry pins the flag, that it is raised
before the first await and cleared in a finally after send(). Confirmed to FAIL
when the guard is reverted. The three pre-existing tests in that file still pass.
Broad sweep: 2457 passed, 13 failed — all pre-existing test_update_banner_fixes
ordering pollution, identical on a clean tree. One earlier run of the same sweep
showed a 14th failure that did not reproduce across two further runs and whose
name I did not capture; flagging it rather than claiming a clean baseline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018y9pr7ppZgDuCZCT6x96kY
2026-08-24 16:07:22 +00:00
Rod Boev df82b46060 ci: retry unrelated test shard (#7195) 2026-08-24 11:38:27 -04:00
Rod Boev 34c60bdf09 test(updates): avoid conftest module identity assumptions (#7195) 2026-08-24 11:28:56 -04:00
Rod Boev 2e1d2917dd test(updates): match Windows restart evidence and ordering (#7195) 2026-08-24 10:42:24 -04:00
Rod Boev fd9f06d02d fix(updates): harden Windows restart isolation (#7195) 2026-08-24 10:19:53 -04:00
Rod Boev 7fafeb7993 fix(updates): use canonical server entrypoint on Windows restart (#7195) 2026-08-24 10:01:21 -04:00
3b9c632a1f test: make model cache regressions hermetic (#7235)
* test: make model cache regressions hermetic

* test: wait for model cache worker cleanup

* test: fully invalidate model cache state

* ci: re-trigger checks (first-time contributor workflow approval)

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
Co-authored-by: n <a@n>
2026-08-24 02:53:47 -07:00
nesquena-hermesandn fcb28248a1 Release exp-v0.52.263: batch sidebar subagent source lookup (#7231, @starship-s) (#7266)
Co-authored-by: n <a@n>
2026-08-24 02:21:27 -07:00
starship-sandnesquena-hermes 037a35abfd perf(sessions): batch sidebar subagent source lookup (#7231)
Reuse the authoritative batched state.db sidebar metadata to classify delegated subagents before route filtering. Keep direct lookups for single-session recovery and mutation guards.

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-24 02:18:48 -07:00
nesquena-hermesandn df65251d58 release: exp-v0.52.262 — native-image mirror reconciliation (#7230) (#7236)
Co-authored-by: n <a@n>
2026-08-22 15:48:23 +00:00
starship-s 78b0425902 fix(sessions): reconcile native image message mirrors (#7230)
* fix(sessions): reconcile native image message mirrors

Keep the rich WebUI sidecar row when Hermes Agent stores the same
exact-timestamp native image turn as scalar screenshot-marker text.

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna
Assisted-by: OpenCode:opencode/x-preview-f-free
Assisted-by: Claude Code:claude-opus-5

* docs(replay): document native image mirror contract

Record the exact cross-store identity requirements and fail-closed
behavior for native image session reconciliation.

Assisted-by: Hermes Agent:gpt-5.6-sol

* fix(sessions): fail closed on state mirror identity gaps

Treat present-versus-absent candidate identity fields as ambiguous so
session restore preserves distinct state-only turns and their metadata.

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna
2026-08-22 08:47:08 -07:00
nesquena-hermesandn a00b02fccb Release: transcript-reopen bounded state.db cache (#7212, @ruizanthony) (#7226)
Co-authored-by: n <a@n>
2026-08-22 04:51:22 +00:00
ruizanthonyandAnthony Ruiz 00e4c5af6f perf(api): skip state load on cached transcript reopen (#7212)
* perf(session-load): memoize the compression-lineage display stitch

_webui_sidecar_lineage_messages_for_display re-merged multi-thousand-row
parent+child transcripts on EVERY GET /api/session (measured 5.1s per
request on a real 5.2k-message lineage, py-spy: merge_session_messages_
append_only + _matching_visible_duplicate dominate). The merge result is
now cached per session id, validity keyed by the exact stat signature of
the child sidecar AND every snapshot parent involved, bounded LRU(16).
Any write to any involved sidecar invalidates naturally; cache hits and
the populating call hand out shallow-copied rows so caller-side display
metadata cannot corrupt cached state.

Measured: 5.17s cold -> 0.002s warm on the real session; invalidation
after touch re-merges. Tests: tests/test_lineage_display_cache.py
(correctness, copy isolation, child/parent invalidation, no pollution
for lineage-less sessions).

* perf(api): memoize inactive-session display merges

* perf(api): skip state load on cached transcript reopen

* test(api): remove unused lineage-cache import

* fix(api): reject streaming key without state database

* test(api): isolate display cache fixture

* fix(api): bind cache publication to read signature

* fix(api): harden display cache invalidation

* fix(api): bind cache hits to lineage provenance

* fix(api): close remaining display cache scope races

* fix(api): scope display cache revision to target

* fix(api): evict unverifiable lineage provenance

* fix(api): keep incomplete lineage provenance uncached

---------

Co-authored-by: Anthony Ruiz <ruizanthony@users.noreply.github.com>
2026-08-21 21:49:59 -07:00
bsgdigital 15523b382d fix: preserve last char before tool call in streaming text
Replace .trim() with replace(/^\s+/, '') in _stripXmlToolCalls and
_stripXmlToolCallsDisplay to preserve trailing whitespace/characters
that appear before <function_calls> blocks in the streaming text.

Fixes: streaming text losing the last character before a tool call card.
2026-08-21 18:29:24 +00:00
nesquena-hermesandnesquena-hermes cfcc39194a Release exp-v0.52.260: opt-in bindable ports for the three-container Docker stack (#7145, @mercael91) (#7202)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-21 03:33:33 -07:00
43eb46c17f Fix Hermes agent port configuration (#7145)
* Fix Hermes agent port configuration

* fix: bind dashboard to 0.0.0.0 so port-forwarded hosts can reach it

Restore the full three-container compose file (previous commit replaced
it wholesale) and change only the dashboard port binding: 127.0.0.1:9119
prevents access via the server IP, which is exactly the issue reported
in #6876.

* compose: restore loopback-only dashboard binding, document opt-in

The dashboard runs --insecure; publishing it on 0.0.0.0 exposes an
unauthenticated service on every host interface. Restore the
127.0.0.1:9119:9119 binding that matches the file convention (all
services are loopback by default) and document the opt-in exactly
like the webui service does (#6876).

* fix: bindable ports for remote access (issue #6876)

---------

Co-authored-by: mercael91 <mercael91@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-21 10:28:45 +00:00
nesquena-hermesandnesquena-hermes 139b8d6703 Release exp-v0.52.259: resolve @provider:model through the shared parser (#7183, @darthzen) (#7200)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-20 23:48:29 -07:00
05ab4173cf fix(profiles): resolve @provider:model through the shared parser (#7183)
_split_webui_provider_model_value() and _strip_webui_provider_prefix()
split provider-qualified model values at the last colon. That amputates
any model name carrying its own tag: @ollama:qwen3.8:27b-mtp-q8_0
resolved to model 27b-mtp-q8_0 with provider ollama:qwen3.8, and the
truncated name was sent upstream, which 404'd. The splitter also cannot
distinguish a tagged model from a multi-segment custom provider ID, so
@custom:mykey:model-a:free was affected the same way.

Both sites now delegate to config._parse_provider_qualified_model_id(),
the grammar routes.py and the gateway adopted in #6723. The import is
function-local, matching the existing api.config imports in this module.

_split_webui_provider_model_value() persists its result into profile
config, so the mangled name outlived the session that produced it.

Fixes #7182.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-21 06:45:52 +00:00
nesquena-hermesandnesquena-hermes c54d3faca5 Release exp-v0.52.258: resolve session exports against document base URI (#7189, @lowlandsheperd) (#7199)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-20 23:38:26 -07:00
Kenanandnesquena-hermes 6c0f2c765c fix(export): resolve session exports against document base URI (#7189)
* fix(export): resolve session exports against document base URI

* test(export): cover subpath deep links

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-21 06:32:50 +00:00
nesquena-hermesandnesquena-hermes 8066462fcc test(#1699): make static-catalog invalidation assertions overflow-aware (#7197)
The two static-catalog-change invalidation tests asserted the newly-appended
opencode-go model appeared in the group's visible `models` list. That held when
the test was written (opencode-go had 14 static entries), but the curated list
has since grown to 27 — past the model-picker overflow threshold — so the
rebuilt group now splits into a visible `models` head (capped) plus an
`extra_models` overflow tail, and the appended entry (added at the end) lands in
`extra_models`. Cache invalidation itself works correctly; the assertion was
stale.

Assert the appended model is surfaced across the combined picker catalog
(`models` + `extra_models`) so the test stays a pure cache-invalidation
regression check, independent of the overflow cap. Verified non-vacuous (a
genuinely-absent id still fails). No production change.

Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-20 22:35:33 -07:00
nesquena-hermesandnesquena-hermes 23df7bd8f2 test(mcp): skip mcp_server tests cleanly on an unsupported mcp SDK (#7196)
The test module only did importorskip('mcp'), which checks the package is
present but not that it exposes the 1.x Server.list_tools/call_tool decorator
API that mcp_server.py is built on. On an environment carrying mcp>=2.0.0
(which dropped that API), importing mcp_server hard-exits (SystemExit) at its
own SDK version guard, so every test errored at setup instead of skipping.

Reuse the exact predicate mcp_server.py uses (hasattr(Server, 'list_tools'))
to module-level-skip when the installed SDK is out of the supported range.
CI, which pins mcp>=1.28,<2 via requirements-dev.txt, still runs the full
module; a box with a newer mcp now skips cleanly instead of producing
spurious errors/failures.

Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-20 22:24:00 -07:00
nesquena-hermesandnesquena-hermes 5f81ed2b1c Release exp-v0.52.257: fence compression with durable transcript revisions (#6554, @ruizanthony) (#7194)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-20 22:06:04 -07:00
c27f3d67a8 fix(compression): fence durable transcript revisions (#6554)
Co-authored-by: Anthony Ruiz <ruizanthony@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-20 22:01:52 -07:00
nesquena-hermesandnesquena-hermes 46c734a721 Release exp-v0.52.256: suppress false dashboard loopback warning (#6573, @webtecnica) (#7192)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-20 21:35:11 -07:00
webtecnicaandnesquena-hermes 65c185478c fix(ui): suppress false loopback-only warning when public URL is configured (#6573)
* fix(ui): suppress false loopback-only warning when public URL is configured

When a user configures status.browser_url (e.g., via a reverse proxy
at https://dashboard.example.com), the Dashboard link correctly opens
that public URL, but _applyDashboardStatus() still displayed the
'loopback-only' warning because it only checked the browser's hostname.

Fix: suppress the warning when status.browser_url is present — the
user has explicitly configured a public URL, so the loopback bind
is intentional and the warning is a false positive.

Closes #6459

* fix(ui): base loopback warning on resolved dashboard target hostname (#6573)

* fix(ui): complete loopback classifier for dashboard warning (#6573)

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-20 21:32:08 -07:00
nesquena-hermesandn 069a55db58 Release exp-v0.52.255: new chat inherits current conversation workspace (#7180, @ruizanthony) (#7186)
Co-authored-by: n <a@n>
2026-08-20 23:38:22 +00:00
ruizanthonyandAnthony Ruiz eee72f942e fix(workspace): new chat inherits the current conversation's workspace (#7180)
* fix(workspace): new chat inherits the current conversation's workspace

`newSession()` resolved the workspace of a new chat as:

    switchWs || S._profileDefaultWorkspace || S.session.workspace

`S._profileDefaultWorkspace` is global to the profile, hydrated from
`/api/profile/active`. Preferring it over the current conversation meant a new
chat opened from a MES conversation landed on the profile-global value instead
of MES. With several conversations running in parallel on different workspaces,
the context the user was actually in was systematically lost.

The precedence is worse than it looks, because that "profile default" is not a
stable configured value: `get_profile_default_workspace()` reads
`last_workspace.txt` BEFORE the configured key (api/workspace.py ~419-427), and
for the `default` profile that file is the GLOBAL one. `set_last_workspace()`
rewrites it on every `/api/chat/start` and `/api/session/update` of ANY
conversation. So "profile default wins over the current session" degrades in
practice to "whatever workspace another conversation last used wins over the
conversation I am in" -- observed live: the pointer flipped mid-operation
because a concurrent turn wrote to it.

New precedence:

    switchWs || S.session.workspace || S._profileDefaultWorkspace

Preserved invariants:
- the one-shot profile-switch flag still wins first and is still consumed (#823);
- a blank page (no loaded session) still falls back to the profile default,
  which remains required for the empty-composer workspace chip (#804/#5169);
- `_profileDefaultWorkspace` is never nulled by `newSession()` (#823).

This deliberately diverges from upstream #4755, which made the profile default
win. That decision assumed a stable configured default; on a shared/global
`last_workspace.txt` it inverts its own intent. The upstream test is rewritten
with the rationale inline so a future rebase does not silently restore the bug,
plus a new test that evaluates the real expression under node and proves the
blank-page fallback still holds.

Pre-existing failures in tests/test_workspace_stale_recovery.py (3) were
verified identical with and without this change.

* fix(workspace): recover stale inherited workspace

---------

Co-authored-by: Anthony Ruiz <ruizanthony@users.noreply.github.com>
2026-08-20 16:36:39 -07:00
nesquena-hermesandn 1f5e45527a Release exp-v0.52.254: skip redundant sidecar rewrites during a turn (#7172, @ruizanthony) (#7175)
Co-authored-by: n <a@n>
2026-08-20 11:49:13 +00:00
ruizanthonyandAnthony Ruiz 0e5a0fadae perf(streaming): skip redundant sidecar rewrites during a turn (#7172)
The periodic checkpoint thread fires whenever a tool call completes and
unconditionally rewrites the whole session sidecar. A completed tool call
is not proof that anything the checkpoint persists actually changed:
`run_conversation()` mutates an internal copy of the transcript, and the
turn bookkeeping fields (`active_stream_id`, `pending_user_message`,
`pending_started_at`) are written once before the run starts. So the
recurring rewrite normally re-persists byte-identical state.

Each rewrite serializes the entire session while holding the GIL, so its
cost scales with transcript size. Because the HTTP server and the agent
workers share one interpreter, that cost is added latency for every
concurrent request — not just for the session being written.

Measured on this deployment (1,622 sidecars, 6.65 GB, median 2.31 MB,
28.8% above 5 MB, max 35.9 MB):

  sidecar    write    per 10-min turn (40 wakeups)
  1.8 MB    ~100 ms    4.1 s  ->  0.2 s
  24.7 MB   ~1.4 s    41.6 s  ->  2.1 s

A checkpoint now writes only when the state it persists actually moved,
compared through a cheap fingerprint that allocates nothing. The
fingerprint covers every field the checkpoint can legitimately advance
mid-run: session identity (compression rotates the SID), turn
bookkeeping, and the length of the persisted message projections.

Fail-closed on both sides: an unreadable fingerprint returns None and
forces the write, and a 5-minute idle refresh keeps `updated_at` from
going stale on a long quiet turn (the sidebar sorts on it). Crash
recovery is unchanged — any real mutation is still persisted within one
15 s interval.

Co-authored-by: Anthony Ruiz <ruizanthony@users.noreply.github.com>
2026-08-20 04:47:38 -07:00
nesquena-hermesandn 45659362e7 Release exp-v0.52.253: honor agent.clarify_timeout incl 0=unlimited (#7163, @BBaFSE) (#7171)
Co-authored-by: n <a@n>
2026-08-20 09:30:17 +00:00
BBaFSEandnesquena-hermes 0b5d2f3698 fix(clarify): honor agent.clarify_timeout and support unlimited wait (#7163)
The WebUI clarify bridge resolved the prompt timeout from a local, divergent
copy (legacy clarify.timeout, 120s fallback) and treated 0 as invalid, so
agent.clarify_timeout: 0 (unlimited) could not be honored and there was no
way to make a clarification card wait for the user.

Resolve the timeout with the agent core's canonical resolve_clarify_timeout()
(clarify.timeout -> agent.clarify_timeout -> 3600, <= 0 = unlimited) against
the session's own profile config (the streaming worker thread does not inherit
the request's per-profile context), wait with a null deadline when unlimited,
and stop the frontend rendering an expired countdown for non-positive timeouts.

- api/streaming.py: per-session profile config resolution, canonical resolver
  with a compatibility fallback, and an extracted _await_clarify_response()
  helper that polls in 1s slices and honours cancellation.
- api/clarify.py: preserve timeout_seconds <= 0 and emit no expires_at.
- static/messages.js: _clarifyExpiryMs() returns 0 for non-positive timeouts
  so the card stays up without a countdown.
- tests/test_sprint42.py: resolution matrix, per-session profile divergence,
  unlimited resolve-later / cancel-exit, and no-expiry payload.

Closes #7159

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-20 09:28:23 +00:00
nesquena-hermesandn c6842e29ee Release exp-v0.52.252: route overlapping custom_providers by active endpoint (#7107, @rouVling) (#7169)
Co-authored-by: n <a@n>
2026-08-20 07:11:57 +00:00
4528279cb2 fix: route overlapping custom_providers models by active endpoint (#7107)
* fix: route overlapping custom_providers models by active endpoint

When two custom_providers[] entries list the same bare model id (e.g.
dogapi and packyapi both advertising 'claude-sonnet-5'),
resolve_model_provider scanned them in config write order and returned
the first match, silently ignoring the active provider's base_url. A
user whose model.base_url pointed at packyapi got routed to dogapi
because dogapi appeared earlier in config.

Now, when the active provider resolves to a named custom provider
(custom:<slug>, including via model.base_url -> named-slug matching),
that provider claims models it owns before the ordered scan runs. Falls
through to the legacy first-match scan when the active provider is bare
'custom' with no base_url to disambiguate, so existing single-endpoint
setups are unaffected.

Adds 3 regression tests covering: shared-model active-endpoint wins,
provider-unique model still routes by ownership, and bare-custom
no-base_url preserves legacy order.

* fix: extend active-endpoint routing to aux, handoff, and slug-collision paths

Review follow-up: the overlapping-id fix must reach three sibling
resolution paths that still routed on the wrong endpoint.

1. Auxiliary slot base_url (config.py set_auxiliary_model): a bare
   resolve_model_provider(model) ignored the SELECTED aux provider, so a
   custom:A slot was persisted with the active custom:B endpoint when both
   listed the model. Look up the selected provider's own custom entry via
   resolve_custom_provider_connection. (NOT model_with_provider_context
   here: the @custom:<slug>:model form resolves base_url to None, which
   would drop the base_url entirely.)

2. Handoff summary (routes.py): read s_obj.model_provider and pass
   model_with_provider_context(model, model_provider) so a session pinned
   to custom:A is not rerouted to the active custom:B. base_url is
   backfilled from the provider's own custom entry downstream.

3. Normalized-name slug collision (config.py resolve_model_provider): two
   distinct provider names can normalize to the same slug (e.g. "Foo Bar"
   and "foo-bar"). The active-slug guard now matches the EXACT entry by
   base_url instead of re-deriving it from the non-unique slug, and fails
   closed (falls through to the ordered scan) on genuine ambiguity.

Adds 4 regression tests: normalized-slug base_url disambiguation,
slug-collision fail-closed, aux-slot selected-provider base_url, and
session-provider-context handoff routing.

* fix: fail closed on custom-provider slug collisions and break aux deadlock

resolve_model_provider() returns a provider slug (custom:<slug>) and the
credential lookup (resolve_custom_provider_connection) later resolves the
API key from the FIRST same-slug entry regardless of base_url. So when two
distinct provider names normalize to the same slug (e.g. "Foo Bar" and
"foo-bar") and both list the requested model, returning the slug silently
pairs one entry's endpoint with another entry's credential. Now raise a
clear AmbiguousCustomProviderError instead of guessing, on both the
active-slug guard and the ordered-scan fall-through paths.

Also: an explicit, unambiguous named custom provider now stays authoritative
even when model.base_url is stale/absent and points at neither entry — a
stale URL is no longer evidence to demote it to config order.

And fix a self-deadlock in set_auxiliary_model(): it held the non-reentrant
_cfg_lock while calling resolve_custom_provider_connection() -> get_config()
-> reload_config_if_stale(), which re-acquires the same lock. The selected
custom provider's base_url is now resolved inline from the config_data
already loaded under the lock.

Tests: collision cases now assert the ambiguity error; added a stale-URL
explicit-provider regression and a watchdog-guarded aux no-deadlock
regression (both fail before the fix).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: fail closed on custom-provider slug collisions across all four paths

The previous cut only guarded the symmetric bare-model case, and it built
collision membership from entries that OWN the requested model. But credential
resolution scans ALL same-slug entries and first-matches regardless of
ownership, so several slug-only boundaries could still pair one entry's
endpoint with another entry's credential — including the asymmetric case where
only one colliding entry lists the model.

Introduce one shared, pure/lock-safe helper (_custom_provider_slug_key +
_unique_custom_provider_entry) that groups ALL named custom_providers entries by
a single canonical slug normalization and raises AmbiguousCustomProviderError
whenever a consumed slug maps to >=2 entries. Apply it at every slug-only
boundary so endpoint and credential can never come from different entries:

- bare resolution (resolve_model_provider): all-entry membership for both the
  active-slug guard and the ordered ownership scan (was owner-only).
- provider-qualified @custom:<slug>:model return (the session/send/handoff
  model_with_provider_context shape), which previously bypassed the check.
- credential lookup (resolve_custom_provider_connection), replacing its private
  first-match-by-slug scan and its duplicate _slugify.
- auxiliary persistence (set_auxiliary_model), replacing its own inline
  first-match-by-slug; the ambiguity now propagates (fails the save) instead of
  being swallowed, while other best-effort errors stay non-fatal. Still resolves
  from the in-scope config_data so it remains lock-safe (no deadlock regression).

Tests: add fail-closed regressions for the asymmetric A-first-non-owner/B-owner
bare case (active-endpoint and bare-custom paths), the model_with_provider_context
qualified handoff shape, the credential boundary called directly, and the
auxiliary-save ambiguity (asserts nothing is persisted). All five fail before
this change and pass after.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: one canonical custom-provider slug identity, point-of-return ambiguity, profile-scoped credentials

Deep re-gate follow-up (5 findings):

1. Collision detection now derives from the single slug PRODUCER
   (_custom_provider_slug_from_name) instead of a looser private key, so
   producer-level collisions on punctuated/parenthesized names (e.g.
   'Foo (Bar)' vs 'foo-bar', both -> custom:foo-bar) are detected. Previously
   the looser key normalized 'Foo (Bar)' to 'foo-(bar)' and missed the
   collision, reintroducing the endpoint-A/credential-B bug for those names.

2. The all-entry ambiguity check now runs ONLY at the point a custom:<slug> is
   about to be returned (via a _finalize() guard covering every custom-return in
   resolve_model_provider, including the final active-provider fallback), not up
   front. An unrelated collision on the active slug no longer blocks lanes that
   resolve to a different provider (@openrouter, a non-colliding @custom:other).

3. streaming.py now resolves endpoint AND credential inside ONE profile scope:
   the main path extends the scope over runtime-key + custom-provider resolution,
   and both self-heal retry paths bind the session profile. This stops a detached
   worker from pairing a named profile's endpoint with the default profile's key.

4. Handoff summary: AmbiguousCustomProviderError now returns HTTP 400 with the
   actionable rename message instead of degrading to a 200 local fallback the UI
   treats as success.

5. UI: aux + default model saves surface the server's message and abort (retain
   dirty state) instead of swallowing it into a generic failure / "settings
   saved"; the handoff catch shows the 400 rename message verbatim.

AmbiguousCustomProviderError gains a .message attribute for verbatim forwarding.

Tests: parenthesized-name collision fails closed on all direct paths; unrelated
lanes are not blocked by an active collision; custom overrides resolve inside the
passed profile scope; handoff returns 400 (not a persisted fallback); JS
error-surfacing source guards. Verified the finding #1/#2 cases fail before and
pass after (old key misses 'Foo (Bar)'/'foo-bar'; old up-front check over-blocks).

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-20 07:09:25 +00:00
nesquena-hermesandn 6bc16a68bc Release exp-v0.52.251: snapshot OS drop entries before traversal (#7142, @nightcityblade) (#7164)
Co-authored-by: n <a@n>
2026-08-20 04:24:10 +00:00
93814bbb76 fix(workspace): snapshot OS drop entries before traversal (#7142)
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-20 04:18:50 +00:00
nesquena-hermesandn dc9464f62c Release exp-v0.52.250: recognize Agent _row_id in replay merge (#7133, @ruizanthony) (#7161)
Co-authored-by: n <a@n>
2026-08-20 01:06:11 +00:00
43bcdcf97f fix(transcript): recognize Agent _row_id during replay merge (#7133)
* fix(transcript): recognize agent row ids in replay merge

* Strip imported _row_id so replay merge cannot trust caller provenance.

Session import already dropped _state_db_row_id aliases. After treating
Agent _row_id as identity, the same public/import scrubber must drop that
field too, or a crafted import can collide with a later state.db turn.

---------

Co-authored-by: Anthony Ruiz <ruizanthony@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-20 01:01:10 +00:00
fc1dc3a52f Release: preserve remote POSIX workspace paths on macOS host (#7132, @alfred-rootson) (#7150)
Co-authored-by: n <a@n>
Co-authored-by: alfred-rootson <alfred-rootson@users.noreply.github.com>
2026-08-19 13:44:11 +00:00
alfred-rootsonandn 7b38e9aceb fix(workspace): preserve remote POSIX paths without macOS synthetic resolution (#7132)
* fix(workspace): preserve remote POSIX paths without macOS synthetic resolution (#7131)

When hermes-webui runs on macOS with a remote terminal profile pointing to
a Linux host (e.g. terminal.cwd='/home/<user>'), path resolution previously
invoked Path.resolve(), expanding '/home' via macOS synthetic firmlinks to
'/System/Volumes/Data/home/<user>' and failing validation.

- Return Path(normalized_raw) directly from _remote_terminal_workspace_candidate
- Check _remote_terminal_workspace_candidate in _clean_workspace_list before local _safe_resolve
- Add unit test covering Linux /home/<user> remote paths without macOS expansion

* ci: trigger required checks (workflow did not run on fork head)

---------

Co-authored-by: n <a@n>
2026-08-19 13:42:05 +00:00
8883a9994d Release: regeneration atomicity — concurrent-send safe + imported-turn aware (#6677, @rodboev) (#7148)
Co-authored-by: n <a@n>
Co-authored-by: rodboev <rodboev@users.noreply.github.com>
2026-08-19 11:23:46 +00:00
4d03578192 fix(#6611): make regeneration start atomic (#6677)
* fix(#6611): make regeneration transaction canonical

* fix(#6611): close cross-provider regeneration blockers

* fix(#6611): harden regeneration transaction boundaries

* fix(#6611): bind canonical context and fork lineage

* fix(#6611): restore active turn token stamping

* test(#6611): make regeneration regressions CI portable

* test(#6611): preserve the tracked regeneration error fixture

* fix(#6611): preserve extracted session harness contracts

* fix(regeneration): preserve concurrent and imported turn ownership (#6611)

* fix(regeneration): honor imported turn ownership safely (#6611)

* test(regeneration): cover both session lock outcomes (#6611)

* test(regeneration): clear unused transaction fixtures (#6611)

* fix(#6611): consume public active-turn marker

* ci: re-trigger (Playwright install hung ~2.5h on GH runners, warm-cache retry)

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
Co-authored-by: n <a@n>
2026-08-19 11:11:53 +00:00
nesquena-hermesandn 963799158a Release: bound fuzzy matching for large sessions — source-aware workspace-prefix dedup (#7128, @Dandandad) (#7143)
Co-authored-by: n <a@n>
2026-08-19 03:44:51 +00:00
af36f409b0 perf: bound fuzzy matching for large sessions (#7128)
Co-authored-by: Dr. Carlos Daniel Carvalho <292283903+Dandandad@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-19 03:40:46 +00:00
nesquena-hermesandn bec67d60ea Release: decode JSON-array-string skills.disabled in Skills panel (#7120, @webtecnica) (#7141)
Co-authored-by: n <a@n>
2026-08-19 02:41:38 +00:00
webtecnica 20123b1268 fix(skills): decode JSON-array string skills.disabled in panel read and toggle write (#7120) (#7134) 2026-08-18 19:40:04 -07:00
asorourx 1768826e15 Keep static-serving path validator strict; scope dotfile allowance to archive members
The prior change loosened the shared _is_safe_relative_path() to permit dot-prefixed
segments. That validator also gates static serving, asset URLs and manifest paths, so
loosening it made hidden files reachable: GET /extensions/<id>/.env, /.git/config and
/.git/hooks/pre-commit return 200 — credential exposure plus a git-hook write/execute
vector.

Restore _is_safe_relative_path() to strict (reject every dot-prefixed segment) and add
a separate _is_safe_archive_member() wired only into the install/uninstall loops. It
permits a narrow allowlist of benign hidden LEAF files (.gitkeep, .gitignore,
.gitattributes, .env.example) so extensions can still ship them, while dot-directories
(.git/) and all other dotfiles (.env, .secret) stay rejected.

Adds regression coverage for both validators: traversal, NUL/backslash, dot-directories,
and the leaf allowlist edges.
2026-08-18 19:23:41 +00:00
63a562f6ed Release — bound concurrent full-session resolution (#7127) (#7130)
* perf: bound concurrent full-session resolution

* Release — bound concurrent full-session resolution (#7127)

* ci: re-trigger fresh run (Playwright-install hang on prior run)

* ci: re-trigger (Playwright-install CDN hang on 3 shards, 12/15 passed)

* ci: re-trigger (sustained Playwright-CDN hang, run 4)

* ci: supersede 2h20m-zombie Playwright-hung run (run 6)

---------

Co-authored-by: Dr. Carlos Daniel Carvalho <292283903+Dandandad@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-08-18 19:16:01 +00:00
dcb515611d Release — SSE Last-Event-ID resumption fix (#6886) (#7129)
* fix(sse): scope gateway SSE probe and honor Last-Event-ID on /api/chat/stream

Two narrowly-scoped cross-client SSE fixes from the hermes-android#58
follow-up:

1. The /api/sessions/gateway/stream?probe=1 response only describes the
   optional gateway/agent-sessions stream, but its payload gave clients no
   way to tell that. A 404 'agent sessions not enabled' result could be
   misread as SSE being unsupported server-wide, even though the persistent
   per-session stream (/api/session/stream) is always on and independent of
   the show_cli_sessions setting. Add additive scope markers to the probe
   payload (scope, session_stream_available, session_stream_path) and
   clarify the handler docstring.

2. /api/chat/stream emitted id: stream_id:seq on every journaled event
   but ignored the Last-Event-ID request header, so any spec-compliant SSE
   client that auto-reconnects (browser EventSource, Android/CLI clients)
   replayed the journal from seq 0 and double-rendered the transcript.
   _chat_stream_resume_after_seq() now resolves the resume cursor with
   explicit after_event_id/after_seq query params first and the
   Last-Event-ID header as fallback — the same resolution chain
   api/kanban_bridge.py already uses. Both the dead-stream journal replay
   and the live-subscribe gap check consume it; direct callers of the gap
   check that only pass query params keep historical behavior.

Also documents the cross-client SSE endpoint contract in
docs/sse-streams.md.

Related hermes-webui/hermes-android#58

* fix(sse): split cursor presence from validity on /api/chat/stream resume

Address PR #6886 review (gate-fail): the resume logic conflated "no
cursor" with "invalid cursor" into a single None, which broke replay on
three edges. Root cause and fix:

Cursor presence vs validity — _chat_stream_resume_cursor() now returns
(after_seq, resume_requested) instead of a bare seq. A client that
SUPPLIED any cursor (query params, replay=1, or Last-Event-ID) asked to
resume; an invalid / foreign-run / unparseable cursor resolves to
after_seq=None but keeps resume_requested=True, and is normalized to
replay-from-start so no journal events are silently skipped. A genuinely
cursor-less request stays a fresh subscribe.

1. Explicit-query precedence is now decided by parameter PRESENCE, not
   successful parsing: a malformed/foreign after_event_id/after_seq never
   falls through to let the Last-Event-ID header override it and skip
   events.
2. Invalid/foreign cursors no longer bypass replay + gap detection and
   silently drain a truncated buffer; an ahead-of-stream cursor is no
   longer clamped onto the snapshot's terminal fence (which emptied the
   dead-stream replay by filtering the terminal frame).
3. The runner-local observe path now receives the resolved header cursor
   (retaining explicit-query precedence) so header-only runner reconnects
   resume instead of duplicating events from the start.

Regression tests cover invalid-cursor replay-from-start, ahead-of-stream
terminal survival, header-only runner reconnect, and explicit-query
precedence; the stale clamp test is updated to the corrected contract.

Related hermes-webui/hermes-android#58

* fix(sse): consume resolved resume cursor at every consumer on /api/chat/stream

Address PR #6886 r2 review (gate-fail). The r1 presence/validity split
fixed the entry resolution, but the individual consumers still re-derived
ahead/equality/re-parse decisions, so the invariant leaked at the edges.
This makes every consumer honor the resolved (after_seq, resume_requested)
pair and normalizes ahead-of-stream in one place.

1. Equality off-by-one: a valid cursor EQUAL to the snapshot cutoff now
   enters the drain dedup bound (the event at the cursor was already
   delivered), instead of being treated as "ahead" and double-sending the
   retained buffer. Ahead is now strictly after_seq > snapshot_cutoff_seq.
2. Numeric ahead-of-stream cursors are normalized to replay-from-start:
   on the active path, after_seq > snapshot_cutoff_seq resets to 0 before
   gap-coverage/replay so a truncated buffer can no longer falsely declare
   coverage and lose the pre-tail events; on the dead-stream path the
   cursor is normalized against the journal summary's authoritative
   last_seq before _replay_run_journal, so an ahead cursor returns the
   full replay instead of an empty SSE body.
3. Runner-local now consumes the resolved cursor instead of re-parsing
   after_event_id or re-reading Last-Event-ID. The resolver also returns
   the opaque raw cursor so the runner uses the already-precedence-checked
   value; explicit opaque cursor= keeps precedence, invalid/foreign
   cursors pass no cursor (full replay), and the header is honored only
   when no explicit cursor was supplied.

Regression tests cover the equality dedup, active+dead ahead-of-stream
normalization (with after_seq-filtering mocks that assert the real replay
body), and the three runner precedence probes (malformed-explicit blocks
header, foreign passes no cursor, opaque cursor= wins).

Related hermes-webui/hermes-android#58

* fix(sse): carry runner-cursor provenance through /api/chat/stream resume

Address PR #6886 r3 review (gate-fail). The (after_seq, resume_requested,
raw_cursor) tuple carried the cursor VALUE but not its PROVENANCE, so the
runner handoff couldn't distinguish a valid opaque runner cursor from a
journal-shaped one from a malformed one — and gated the runner handoff on
journal-specific validity, which is wrong for the runner adapter (its
cursors are opaque).

The resolver now returns a distinct, centrally-resolved runner cursor:

- a valid after_seq PAIRS with whatever after_event_id shape was supplied
  — including an opaque runner id like event:2 that the journal parser
  reads as a foreign run — so the paired runner cursor resumes at the seq
  (never None / full replay, which would duplicate tokens and tool
  events). after_seq is read directly (via _parse_run_journal_after_seq_value)
  rather than through the after_event_id-first journal parser, which would
  swallow a paired opaque runner id.
- a header-only opaque runner id (Last-Event-ID: event:2) is preserved
  as-is so the runner resumes from it.
- a malformed or foreign explicit cursor WITHOUT a valid paired after_seq
  still yields no runner cursor, blocking Last-Event-ID and replaying from
  start — the r2 malformed-blocks-header guarantee is unchanged.

Regression tests pin the paired opaque id, header-only opaque id, and
malformed-with-valid-seq runner resumes, plus the unchanged
malformed-without-pair replay-from-start behavior.

Related hermes-webui/hermes-android#58

* fix(sse): normalize ahead cursor when the snapshot cutoff is unknown

Address PR #6886 r4 review (gate-fail). The ahead-cursor normalization
was guarded by `snapshot_cutoff_seq is not None`, so when the subscribe
snapshot had no parseable last_event_id (channel has not seen an
id-bearing frame yet), a positive ahead cursor (e.g. Last-Event-ID
run_1:999) was never normalized — it was installed verbatim as the live
dedup bound, filtering EVERY queued frame including the terminal
stream_end fence. The reconnect stalled on heartbeats with an empty SSE
body. origin/master ignored the header on this path and delivered the
events, so this was a new regression.

Fix: treat an unknown, non-truncated snapshot cutoff as fence 0 before
the ahead-cursor normalization, so a positive client cursor can't become
the live dedup bound with no cutoff to bound it — it fails closed to
replay-from-start, delivering the buffered events rather than silently
losing them. Frames dropped while the cutoff is unknown still cannot
prove gap coverage (cutoff None → not covered), so the recovery_control
fail-closed path is preserved.

Regression test: Last-Event-ID run_1:999 + snapshot_cutoff_seq=None
delivers queued events run_1:1 and run_1:2 exactly once, emits the
terminal stream_end, and releases the subscriber. Without the fix the
test stalls on heartbeats (pytest-timeout kills it), proving the stall.

Related hermes-webui/hermes-android#58

* Release — SSE Last-Event-ID resumption fix (#6886)

---------

Co-authored-by: Paladin173 <35980893+Paladin173@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-08-18 15:08:14 +00:00
59023b784c Release — approval no-run mirror masking fix (#7086) (#7125)
* fix(approval): stamp mirror token when creating no-run legacy mirror (#6008 follow-up)

In in-process (legacy) mode, a guarded command parks an _ApprovalEntry in
_gateway_queues whose .data has no run_id. submit_gateway_pending_mirror()
mirrors it into _pending for the UI in the `elif not exact_local_entry`
branch, but that branch never stamped _GATEWAY_MIRROR_TOKEN on the mirror.

_resolve_approval_legacy() binds a no-run mirror back to the parked
_ApprovalEntry via a token match (pending_token == gateway_token). Without
the mirror token the match fails, resolve_gateway_pending_local() is never
called and entry.event.set() never fires, so the agent thread stays blocked
even though the endpoint returns ok:true and the UI closes the card. Only a
~1.5s reconcile fallback re-creates the card with a token -> the FIRST click
on "Allow once" silently does nothing.

Stamp the token from the live gateway head at mirror-creation time, matching
the exact-local branch, so the first click unblocks the agent.

Adds a regression test reproducing the scenario.

* fix(approval): resolve non-head no-run mirrors by scanning all live producers

The first-click fix stamps the mirror token from the mirror's own producer.
But _resolve_approval_legacy() previously matched that token only against the
queue head (`_gateway_queues[sid][0]`), so a non-head mirror (multiple parked
producers, #7093) could not be resolved: the head's token differs from the
selected producer's, the match failed, and the agent thread stayed blocked.

Scan every live producer in _gateway_queues for a matching
_webui_mirror_token or approval_id, and stamp the resolved producer's
approval_id before delegating to resolve_gateway_pending_local(). Registered
producers keep head/tail ordering, so exact-producer isolation is preserved.

Adds tests/test_issue_legacy_first_click_repro.py::test_legacy_non_head_
producer_unblocks_only_its_producer, which reproduces the non-head scenario.

* fix(approval): fail closed on no-run mirror token binding (818fd2fd follow-up)

submit_gateway_pending_mirror() previously fell back to "first unclaimed
token" or the queue head when a no-run mirror's identity (request_id/
approval_id) didn't match any live producer. That let approving a
stale/foreign mirror silently resolve a different, unrelated live
producer than the one the user actually saw.

Bind the token only on an explicit identity match; otherwise leave the
mirror tokenless so it can't borrow another producer's slot. A tokenless
orphan is still created automatically for the legitimate #7093 case
(no live producer at all) via reconcile_gateway_pending_mirror_locked.

Adds two adversarial regression tests (stale request_id vs. a live
sibling, and an identity-less mirror with multiple producers parked) and
fixes an unrelated test that omitted request_id, unlike every real
_ApprovalEntry.

* fix(approval): don't let an unmatched no-run mirror mask a live producer

Follow-up to 63087a40. Failing closed on token binding stopped stale
mirror A from executing live producer B, but left A parked in _pending:
reconcile then deferred to the pre-existing no-run mirror and skipped
tokenizing B, so B never surfaced as the head. The real pending approval
became unactionable (409 in gateway mode, ok:true resolving nothing in
local mode) while an unresolvable card sat in front of it.

reconcile_gateway_pending_mirror_locked() now always tokenizes every live
no-run producer and derives live_token from the authoritative head
whenever _gateway_queues holds producers, so an unmatched tokenless
mirror is discarded instead of suppressing the real one. Retained
missing-producer mirrors survive only while no producer is live, keeping
the #7093 tokenless-orphan boundary intact.

Tests assert the displayed == resolved invariant: the authoritative
producer is the head, responding with the stale/identity-less id fails
without waking anything, and approving the displayed entry wakes exactly
its own producer. Three test_gateway_approval_runs_api.py cases that
encoded the old masking behavior are updated to the new invariant.

* docs(approval): correct reconcile comment on which mirrors survive

The comment added in 9c2ca104 claimed only the authoritative head's mirror
may survive while a producer is live. That contradicts the code it
documents and the non-head regression test: the `entry_token in
live_local_tokens` branch keeps a mirror bound to ANY live no-run
producer, which is exactly what makes a non-head mirror resolvable
(test_legacy_non_head_producer_unblocks_only_its_producer).

State the actual rule instead — a mirror survives only if bound to some
live producer's own token, the head's via live_token or a non-head
producer's via live_local_tokens — so unmatched and tokenless copies are
discarded without implying non-head mirrors are dropped too. Comment
only; no behavior change.

* Release — approval no-run mirror masking fix (#7086)

---------

Co-authored-by: VHSgunzo <wosganzo@gmail.com>
Co-authored-by: n <a@n>
2026-08-18 13:11:07 +00:00
cced4f3509 Release — approval-card YOLO handoff fix (#6945) (#7122)
* fix(approvals): handle gateway yolo in webui

* fix(approvals): preserve concurrent yolo enables

* fix(approvals): honor settled yolo state

* fix(approvals): commit yolo after relay

* fix(approvals): preserve gateway yolo relay state

* fix(approvals): close remaining yolo race windows

* fix(approvals): serialize gateway yolo handoff

* fix(approvals): linearize yolo disable

* fix(approvals): serialize initial yolo relay

* fix(approvals): serialize card yolo handoff

* fix(yolo): ignore stale card responses

* fix(yolo): reject stale session generations

* fix(yolo): bind responses to approval owner

* fix(yolo): drain pending approvals with exact ownership

* fix(approvals): show skip-all busy state

* fix: close approval lifecycle races

* test: update approval SSE handoff contract

* test: include displayed approval owner helpers

* fix: linearize yolo transition registration

* Release — approval-card YOLO handoff + stale-response ownership fence (#6945)

---------

Co-authored-by: dss539 <dss539@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-08-18 12:41:36 +00:00
b1c2bea92e Release — scoped extension Configure hook (#7111) (#7116)
* Add scoped extension Configure hook

* docs: fail-closed optional chaining in Configure consumer example + old-Core regression

The documented consumer used ext?.settings.registerConfigure(...), which throws
TypeError on an older E0 Core where settings exists but registerConfigure does not.
Use full optional chaining (ext?.settings?.registerConfigure?.()) and add a Node
regression proving the consumer survives a Core lacking the capability.

Co-authored-by: franksong2702 <franksong2702@users.noreply.github.com>

---------

Co-authored-by: Frank Song <franksong2702@gmail.com>
Co-authored-by: n <a@n>
Co-authored-by: franksong2702 <franksong2702@users.noreply.github.com>
2026-08-18 05:30:27 +00:00
5e8937981d Release — suppress Windows git console windows (#6534) (#7115)
* fix: suppress console windows for git subprocess calls on Windows

On Windows, every git.exe spawned by the WebUI (update checks,
version detection, workspace git info, rollback inspection,
worktree management) creates a visible console window. When
Windows Terminal is the default terminal, each short-lived git
process opens a new tab. Combined with the 60-second version-skew
monitor polling /api/settings and the boot-time update check, this
produces a flood of flashing terminal tabs.

Root cause: subprocess.run() calls in api/updates.py, api/rollback.py,
api/worktrees.py, and api/workspace.py did not pass
creationflags=CREATE_NO_WINDOW on Windows. api/workspace_git.py
already had the fix (#5692); these four files were missed.

Backend fix:
- Add _windows_hide_flags() helper to each affected module
  (mirrors the existing pattern in api/workspace_git.py)
- Pass creationflags=_windows_hide_flags() to all git
  subprocess.run() calls (10 call sites across 4 files)

Frontend fix (prevent timeout->retry cascades):
- Increase update-check timeout from 30s/60s to 300s (5 min)
- Increase /api/settings fetches from 30s to 60s
- Increase /api/git-info fetch from 30s to 60s

The backend git fetch to GitHub can take 15s per repo; with both
WebUI and Agent repos checked sequentially plus other git calls,
the total easily exceeds the old 30s default frontend timeout,
triggering the api() wrapper auto-retry (3 attempts) and
amplifying the git process flood.

* [verified] fix: address Windows git subprocess review

---------

Co-authored-by: ya1yi2fa3 <ya1yi2fa3@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-08-18 04:25:26 +00:00
509cdb8dfc Release — carry explicit_model_pick through /goal (#6705) (#7114)
* fix(session): sync provider change mid-session (#6703)

* fix(session): full chat-start parity for /goal explicit-pick (#6703)

Addresses review parity details on #6703:
- stamp model_explicit_pick_signature after explicit goal resolution
  (both kickoff paths), matching /api/chat/start #5979, so a first /goal
  launch after a deliberate custom-provider pick survives a cold
  streaming catalog instead of reverting to the profile default
- consume the one-shot pending session-model marker in cmdGoal,
  matching the chat/start path (messages.js), so a later /goal with an
  unchanged dropdown isn't treated as an explicit pick

* fix(session): consume pending model-pick marker only after successful kickoff (#6705)

* fix(session): node-driven behavioral tests for /goal explicit-pick marker (#6703)

* changelog: #6705 — carry explicit_model_pick through /goal

* ci: re-trigger tests on fresh runners (prior run's shards hung on GitHub-hosted runners)

* ci: re-trigger (playwright browser-install hung on GH runners again; warm cache should restore)

---------

Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: n <a@n>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-18 03:03:05 +00:00
83e4903274 Batch release — double-render fix (#7109) + inline-video preload (#7076) (#7113)
* perf(media): preload visible chat videos

* fix(media): ignore detached visibility callbacks

* fix(ui): require provable live owner for live-turn preservation (#6948)

A completed assistant turn could render twice: when the stream ended
(S.activeStreamId cleared) but INFLIGHT[sid] was not yet cleaned, the
preserve guard treated bare rendered content (.msg-body presence) as
proof of liveness and re-attached the stale live DOM node OVER the
rebuilt settled transcript — a second copy of the same message. Data
was always clean (state.db, sidecar, /api/session each hold one row);
the duplicate existed only in the rendered DOM.

Preserve the live turn across the wipe only while there is a provable
live owner: an active stream (S.activeStreamId, the #3877 mid-stream
case) or explicit live-assistant evidence in the current message
projection (S.messages carrying a client-side _live / _activityBurstId /
_liveSegmentSeq marker — the reconnect/terminal-projection case). A
settled transcript has neither, so the durable settled render wins and
appears exactly once. The #5390 blank-turn guard (对话消失) is preserved:
a dead empty shell has no live projection either and is still dropped.

* changelog: exp batch — #7109 double-render fix + #7076 video preload

---------

Co-authored-by: allenliang2022 <allenliang2022@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: n <a@n>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-18 00:11:28 +00:00
69d5647b60 Batch release exp-v0.52.237 — profile goal isolation (#6899) + hermes-webui entry point (#6742) (#7108)
* fix(goals): delegate profile evaluation to native manager

Use Hermes' context-local home override to run profile-scoped goal operations through the native GoalManager, preserving current judge, wait, and failure semantics while retaining the explicit-DB legacy fallback.

* docs(goals): describe profile ownership boundary

* fix(goals): gate native profile persistence capability

* feat(cli): add packaged hermes-webui CLI entry point (#6739)

* fix(cli): route hermes-webui entry through bootstrap:main for wheel install (#6742)

* docs(changelog): note #6899 profile goal isolation + #6742 hermes-webui entry point

* test(goals): guard hermes_cli import with importorskip for CI (agent not installed)

#6899's native-contract tests imported hermes_cli unconditionally, failing
CI with ModuleNotFoundError. Match the repo's established importorskip pattern
so they skip cleanly when the agent isn't installed and run when it is.
Co-authored-by: ticketclosed-wontfix

---------

Co-authored-by: Nick <202622897+ticketclosed-wontfix@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: n <a@n>
2026-08-17 10:45:59 +00:00
747b5a6913 Batch release exp-v0.52.236 — dotted model labels (#6628) + MCP SDK pin (#6616) (#7103)
* fix(models): normalize dotted Bedrock/Vertex model IDs in labels

Split out of #6607 at reviewer request — it was unrelated to that PR's MEDIA
path handling and needed its own tests.

`us.anthropic.claude-opus-5` carries a cross-region routing prefix plus a vendor
namespace, and `mistral.mistral-large-2407-v1:0` adds a provisioned-revision
suffix. None of it belongs in a human label, so the turn footer rendered
"Us.anthropic.claude Opus 5".

Strip only the two shapes these hosts actually publish, against a CLOSED
provider allow-list:

  <region>.<vendor>.<model>   us.anthropic.claude-opus-5
  <vendor>.<model>            mistral.mistral-large-2407-v1:0

A generic "drop leading letters-only dot segments" loop was rejected because it
rewrites arbitrary uncatalogued IDs: `deepseek.v3` rendered as "V3" (vendor name
silently deleted) and `foo.bar.baz` as "BAZ". Dropping a vendor is additionally
gated on the remainder still naming the model, so `deepseek.v3` — where the
vendor IS the name — is left byte-intact.

Backend and frontend are paired: tests/test_dotted_model_label.py drives one
table through both `_get_label_for_model()` and `_stripDottedModelPrefix()` and
fails on divergence, including version dots (`gpt-4.1`, `qwen3.6-35b`),
URI-scheme IDs, and unknown vendors.

The Python half is inlined inside `_get_label_for_model` on purpose: the #3429
harnesses extract that function's source and eval it in isolation, so a
module-level helper NameErrors there. The #3429 JS driver is updated to pull in
the new helper for the same reason.

* fix(models): recognize `global` as a Bedrock routing prefix

Re-review catch, and a real gap: the catalog ships six
`global.anthropic.claude-*` IDs (api/config.py:1901-1909) and the first-party
routing notes use `global.anthropic.claude-…` as the canonical Bedrock shape,
but `global` was missing from the region allow-list. All six therefore kept the
noise this change exists to remove:

    global.anthropic.claude-opus-4-7  ->  "Global.anthropic.claude Opus 4 7"

Added to the region set in both implementations, which now read identically.

The root cause is two lists that must agree — the region allow-list and the
shipped catalog — so the new guard is catalog-driven rather than another
hardcoded region list: it scrapes every dotted `<head>.<vendor>.<model>` ID out
of api/config.py and asserts none of them keeps a routing/vendor namespace in
its label. A future routing prefix added to the catalog without updating the
region set fails there, instead of shipping mislabeled.

Also pins all six `global.*` IDs plus `us-gov` in the shared strip table (no
suite covered `us-gov` before either), and verified by mutation: dropping
`global` from the Python set fails 3 tests, dropping it from only the JS set
fails 2 — so backend/frontend parity drift is caught, not just a total absence.

Note: the review cited tests/test_provider_prefix_label_normalization.py and
tests/test_ui_model_label_parity.py as the suites to extend; neither exists in
this tree (nothing under tests/ referenced `us-gov` at all), so the cases live
in tests/test_dotted_model_label.py alongside the existing paired coverage.

* fix(models): add missing Bedrock vendors; make the catalog guard actually broad

Two self-review findings (hostile critic pass). Both are the same class as the
`global` gap the reviewer already caught, which means the guard I added for that
gap was too weak to prevent a recurrence.

1. Missing vendors. `luma`, `twelvelabs` and `ibm` are real Bedrock
   foundation-model vendors and were absent from the allow-list, so genuine IDs
   shipped with the namespace intact:

       luma.ray-2                       -> "Luma.ray 2"
       us.twelvelabs.marengo-embed-2-7  -> "Us.twelvelabs.marengo Embed 2 7"
       us.ibm.granite-3-8b              -> "Us.ibm.granite 3 8B"

   Added those plus `nvidia` and `snowflake` to both implementations.

2. The catalog guard inspected 6 of 75 dotted IDs. Its scrape regex only matched
   three-segment `"id": "<region>.<vendor>.<model>"` literals with double quotes,
   so the entire two-segment `<vendor>.<model>` shape -- the OTHER documented
   shape this PR handles -- was invisible to it. A test that reads 8% of the
   corpus while claiming to cover the catalog is worse than no test, because it
   reads as proof.

   Rewritten to scrape any quoted `id` value regardless of quote style or segment
   count (75 dotted IDs now inspected), and to DERIVE the namespace heads from the
   production `_regions`/`_vendors` literals instead of retyping them, so a set
   that grows without a test update is still covered. Version dots (`qwen3.6-plus`,
   `gpt-5.4`) are correctly skipped as non-namespaces.

Also added an explicit per-vendor round-trip test. It asserts no dotted NAMESPACE
survives rather than that the vendor word is absent, because a vendor legitimately
reappears inside some model names (`mistral.mistral-large-2407` -> "Mistral Large
2407") -- my first version of that assertion was wrong for exactly that reason.

Verified by mutation: removing the new vendors fails the round-trip test; deleting
the strip entirely fails 6 tests.

* test(models): drive real getModelLabel() in the parity oracle

The paired test compared _get_label_for_model() against itself, so JS
getModelLabel() never ran and nothing asserted the two sides agreed. It
stayed green when the JS dotted strip was reverted to a no-op AND when the
whole post-strip retry chain was deleted -- blind to both.

Replaced with two tests driven through the real getModelLabel() under Node,
reusing the driver already proven in test_issue3429_uri_scheme_model_label.
Only sinks are stubbed (_dynamicModelLabels empty as it is pre-fetch,
_fmtOllamaLabel identity); every decision function is the shipped source.

The oracle asserts the actual contract -- no routing/vendor namespace leaks
into either label -- rather than string equality of the two labels. Label
divergence is pre-existing and by design: at base dd7f6ac3, claude-opus-5
(no dot, untouched here) already labels 'claude-opus-5' in JS vs
'Claude Opus 5' in Python, and openai/gpt-4o 'GPT-4o' vs 'GPT 4O'.
getModelLabel() checks _dynamicModelLabels first (ui.js:6991), populated
from the server label (ui.js:3483,3494,3606), so the backend wins once a
catalog loads; the JS formatter is the pre-fetch fallback.

Dropping the JS post-normalization instead would regress the picker to raw
ids before catalog load (us.anthropic.claude-sonnet-4-5 -> 'claude-sonnet-4-5'
instead of 'Sonnet 4.5').

Mutation-checked: JS strip no-op fails 2, retry chain deleted fails 1,
'global' dropped from backend _regions fails 1. Region set derived from
api/config.py source rather than retyped.

No production code changed.

* fix(test): pin mcp SDK to compatible 1.x range with bootstrap guard (#6602)

Change  to  in requirements-dev.txt and the
mcp_server.py docstring, and add a bootstrap guard in mcp_server.py
that checks Server.list_tools existence at module load time.

The guard fails fast with a clear error message if an incompatible
mcp SDK (2.x) is installed, instead of producing dozens of
secondary errors at test collection/setup. This completes the
mitigation started in PR #6564 by adding a minimum version floor
and an import-time compatibility check.

* docs(changelog): note #6628 dotted model labels, #6616 mcp SDK pin

---------

Co-authored-by: Sam Painter <samfp@amazon.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: n <a@n>
2026-08-17 06:37:52 +00:00
92c790ffe5 Batch release exp-v0.52.235 — cancelling-run reattach (#7096) + GPT-5.6 max reasoning (#7083) + test order-independence (#7101) (#7102)
* fix(tests): make mtime_invalidation and glm_5_3 tests order-independent (#7100)

* fix: expose max reasoning for GPT-5.6 models

* fix(recovery): do not reattach cancelling runs

ACTIVE_RUNS tracks worker lifecycle, which is deliberately broader than
"a turn a browser may attach to". cancel_stream() keeps the row as
phase="cancelling" while the worker unwinds so a successor turn cannot
start on top of it, but the client has already reached a terminal state
for that stream: its run journal ends in a terminal event.

The recovery lookups treated every same-session ACTIVE_RUNS row as
attachable. An idle session holding a cancelling row therefore received a
recovered server_turn_started on every /api/session/stream subscription:
the client attached, replayed the terminal event, tore the renderer down,
resubscribed, and the server replayed the same frame again. The result is
an endless attach/replay loop that rebuilds the transcript repeatedly.

Separate the two meanings instead of narrowing one call site:

- api/config.py gains active_run_is_attachable() and
  active_run_cancel_is_stale() as the shared predicates.
- active_stream_id_for_session() (browser recovery) returns attachable
  rows only; _session_has_active_turn() (busy check) keeps counting a
  fresh cancellation so a successor cannot overlap the unwinding worker.
- _live_active_stream_id() applies the same rule to the hidden-tab
  status poller, on both the STREAMS and ACTIVE_RUNS paths.
- routes._cancelled_run_is_stale() now delegates to the shared predicate
  rather than keeping a parallel copy of the staleness rule.
- A cancelling row past a bounded unwind window with no live STREAMS
  channel is reclaimed from ACTIVE_RUNS and its stream owner released, so
  a wedged worker cannot suppress background wakeups forever. Age alone
  does not reclaim a row that still owns a live channel.

Staleness anchors on cancelled_at, falling back to started_at, so a
long-running turn cancelled moments ago is never treated as an orphan.

tests/test_cancelling_run_not_attachable.py covers both directions of
each rule. The six behavioral tests fail on the unpatched tree and pass
with the fix; two targeted mutations (forcing the attachability predicate
true, and disabling the staleness reaper) each turn the suite red.

* docs(rfc): document cancellation attach/admission split

Records the runtime contract introduced by the recovery fix in the WebUI
run-state consistency RFC, so the distinction is discoverable instead of
living only in code comments.

- Adds ACTIVE_RUNS to the State Layers table as the worker-lifecycle
  registry, explicitly not the set of runs a browser may attach to.
- Adds invariant 9: lifecycle-busy is not client-attachable. Cancellation
  splits the two meanings, recovery paths must exclude cancelling rows,
  and admission checks must keep counting them.
- Documents the bounded cancellation-unwind window: reclamation needs
  both age and the absence of a live STREAMS channel, and staleness is
  anchored on the cancellation timestamp.
- Extends the review checklist with the admission-vs-attachment question
  and the evidence required when changing a reclamation window.

* docs(changelog): note #7096 cancelling-run reattach, #7083 GPT-5.6 max reasoning, #7101 test order-independence

---------

Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: Abdulrahman Elkenany <boudy.elkenany123@gmail.com>
Co-authored-by: allenliang2022 <allenliang2022@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-08-17 05:37:00 +00:00
e88aff4388 Batch release exp-v0.52.234 — O_BINARY (#7077), native-control font (#6871), mobile nav order (#6909) (#7099)
* fix(navigation): preserve mobile action mirror order

Reposition stale extension action mirrors during synchronization without moving already-correct elements, preserving native tab order and keyboard focus.

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna
Assisted-by: Claude Code:claude-opus-5

* fix(typography): route native controls through UI font

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna
Assisted-by: Claude Code:claude-opus-5

* fix(windows): open anchored files in binary mode

* docs(changelog): note #7077 O_BINARY, #6871 native-control font, #6909 mobile nav order

---------

Co-authored-by: starship-s <45587122+starship-s@users.noreply.github.com>
Co-authored-by: allenliang2022 <allenliang2022@users.noreply.github.com>
Co-authored-by: n <a@n>
2026-08-17 01:42:02 +00:00
f60805de01 fix(#6867): refresh artifacts after settled-session recovery (#7067)
* fix(#6867): preserve artifact ownership across compression rotation

* docs(changelog): note the #6867 artifacts-panel ownership-guard fix

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
Co-authored-by: n <a@n>
2026-08-17 00:53:49 +00:00
f1115355a6 fix(#6906): keep kanban actions reachable in short windows (#7065)
* fix(#6906): tighten Kanban modal viewport regression

* fix(#6906): preserve modal centering fallbacks

* docs(changelog): note the #6906 Kanban modal height-cap fix

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
Co-authored-by: n <a@n>
2026-08-16 23:45:51 +00:00
nesquena-hermesandn 53d96fa7d7 docs(changelog): note the #4948 stale-approval-card fix (shipped in #7093) (#7094)
Co-authored-by: n <a@n>
2026-08-16 23:00:26 +00:00
nesquena-hermesandn d9c5b33449 fix(approval): match the notified gateway-copy to its producer by request_id (#7093)
submit_gateway_pending_mirror() matches an incoming approval back to its
queued producer entry by object identity (entry.data is approval) then by
approval_id. On the LOCAL backend neither holds: the installed hermes-agent
core delivers the approval to WebUI's notify callback as a COPY
(notify_cb(dict(entry.data))), so identity fails, and a local _ApprovalEntry
stamps a request_id but no approval_id, so the approval_id fallback misses.
The mirror is then created with no token, reconcile keeps the orphan, and
_session_has_pending_approval stays True after the entry is dropped — the
stale-approval-card dead-end (#4948 local variant): clicking a card whose
approval is gone returns a bare ok:false the UI renders as "Approval response
not accepted." with a stuck card.

Add a request_id fallback (gated on the local not-run_id path) that reunites
the notified copy with its queued entry via the request_id the core stamps
once on the source and preserves through the copy.

The four HTTP tests in test_approval_unblock.py were unfaithful to production:
they built _ApprovalEntry(approval) (stamping request_id onto a copy in
entry.data, not onto the source approval) then submitted the pre-stamp source.
They now submit dict(entry.data), matching what the core actually passes. The
assertions are unchanged; with the fidelity fix but the production change
reverted, the tests fail — so they detect the real regression, not vacuously.

This class is invisible on CI, which never installs the agent core, so both
approval test files skipif-skip entirely and their green never exercised these
paths.

Co-authored-by: n <a@n>
2026-08-16 15:56:25 -07:00
d13e922120 Release batch B: Docker single-container UID fix + GLM-5.3 catalog (experimental) (#7090)
* feat: add GLM-5.3 to Z.AI model list

Add glm-5.3 as the newest zai entry in _PROVIDER_MODELS and a matching
zai/glm-5.3 entry in _FALLBACK_MODELS, and bump the Z.AI onboarding
default_model from glm-5.1 to glm-5.3 (Z.ai's current flagship; legacy
GLM-5.2/5.1 requests are routed to GLM-5.3 per docs.z.ai).

Reasoning gating needs no change: _zai_glm_classification() treats
GLM >= 5.2 as effort-ladder capable, so glm-5.3 is already covered and
pinned by tests/test_zai_reasoning_effort_gating.py.

New regression coverage in tests/test_glm_5_3_catalog.py: catalog
presence, newest-first ordering, fallback entry, onboarding default,
full reasoning_effort ladder, and the get_available_models() payload.

* test: restore _cfg_fingerprint in catalog test fixture

Review follow-up (#7017): the isolation fixture snapshot restored cfg,
_cfg_mtime, and _cfg_path but left _cfg_fingerprint pointing at the
temporary config loaded by the payload test. api/config.py uses that
fingerprint to distinguish in-memory overrides from changed files
(config.py:371), so a stale value could make later same-process tests
skip reloading a changed config. Snapshot and restore it like the rest.

* fix: keep Z.AI onboarding default at glm-5.1 until direct API serves GLM-5.3

Review follow-up (#7017): GLM-5.3 is live on Z.ai's Coding Plan
endpoint only; the direct api.z.ai endpoint the zai provider uses
still lists the GLM-5.3 API as coming soon. Defaulting new direct-API
users onto glm-5.3 would fail their first message, so the catalog
addition stays (opt-in) and the default stays glm-5.1. Bump the
default in a follow-up once the direct endpoint serves GLM-5.3.

* fix(docker): probe the configured state dir before /workspace for UID/GID (#7027)

UID/GID auto-detection probed /workspace before the configured
HERMES_WEBUI_STATE_DIR. In a stock single-container image none of the
priority-1 candidates exist, but /workspace does — owned by the image's
build-time 1024:1024. Detection therefore returned the image's own owner,
which carries no information about the host, while the one directory whose
owner *is* the host UID by definition — the state-dir bind mount — was never
probed. With a host-owned state mount and no explicit WANTED_UID the
container remapped to 1024, failed its own state-dir writability check, and
restart-looped.

The log line made this expensive to debug: 1024 is also the fallback default,
so "Auto-detected workspace UID: 1024" read as if detection had found nothing.

- probe ${HERMES_WEBUI_STATE_DIR:-/app/data} first, for both UID and GID
- keep /workspace as a lower-priority signal (unchanged for setups that
  actually bind-mount it), and keep the hermes-home probes from #668
- stop treating an explicitly supplied 1024 as "unset": the sentinel and a
  valid UID were the same number, so an operator who deliberately ran as 1024
  got it overwritten by detection. The explicit/detected origin is persisted
  next to the value because `su` drops the environment when the script
  re-enters as the runtime user, so the second pass would otherwise see an
  explicit choice as a detected one.

Tests: tests/test_7027_state_dir_uid_probe.py runs the real resolution block
under bash with `stat` stubbed, covering the non-1024 state mount, the
explicit 1024 override (including across the privilege drop), and the
pre-existing /workspace + hermes-home fallback paths. A new state-dir-uid job
in the Docker smoke workflow boots a real container on a host-owned state
mount and gates on /health.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Release batch B: Docker UID fix + GLM-5.3 catalog (experimental)

Two independently gate-passed contributor PRs, rebased fresh onto master and
re-gated as a combined stage (Codex SAFE TO SHIP; full suite green except the 8
pre-existing approval-test failures that fail identically on clean origin/master
— CI green on the same commit; tracked separately for a fix).

- #7027 (@jorgejiro) probe configured state-dir before /workspace for Docker
  UID/GID; persist explicit marker so a supplied 1024 survives root->su re-entry
  (fixes the single-container restart loop) (#7034)
- #7017 (@rh-id) add GLM-5.3 to the Z.AI model list (onboarding default stays glm-5.1)

Co-authored-by: jorgejiro <jorgejiro@users.noreply.github.com>
Co-authored-by: rh-id <rh-id@users.noreply.github.com>

---------

Co-authored-by: Ruby Hartono <58564005+rh-id@users.noreply.github.com>
Co-authored-by: Jorge <jorgejiro@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: n <a@n>
Co-authored-by: jorgejiro <jorgejiro@users.noreply.github.com>
Co-authored-by: rh-id <rh-id@users.noreply.github.com>
2026-08-16 21:52:14 +00:00
8686528b1b Release batch A: 6 low-risk gate-passed fixes (experimental) (#7089)
* fix(goal): suppress reserved SILENT sentinel at /api/goal ingress

A wake relay POSTing the exact [SILENT] suppression sentinel to
/api/goal lets it reach _start_chat_stream_for_session, which persists
it as pending_user_message. If 8701 restarts while the turn is pending,
session recovery materializes it as a visible _recovered user turn —
the same phantom-recovered-turn exposure #7018 closed for chat ingress.

Apply the shared _is_silent_control_message() guard immediately after
/api/goal session-ID validation and before session lookup or goal-state
mutation, returning the same 200 no-op. Matching stays exact and
case-sensitive; ordinary kickoff text is unaffected.

Add tests mirroring test_silent_control_suppression.py for the goal
path (args and text fields, before-lookup suppression, exact-match
semantics).

Closes #7019

* fix(sessions): keep evicted subagent parents in the import window

The sidebar nests a subagent row under its parent only when the parent row
is present in the same payload. The visible-window limit was applied as a
flat per-row recency slice, so a frozen orchestrator (which stops writing
while its leaves keep streaming) lost the recency race against its own
leaves and fell outside the window -- promoting those leaves to top-level
sidebar rows.

Re-add subagent parents that the oversampled candidate set already
projected, after the slice. No extra queries, no change to
CLI_VISIBLE_SESSION_LIMIT. webui ancestors are deliberately not imported
because that sidebar bucket already owns them.

Supersedes #7031.

* fix(sessions): document and pin the parent-recovery bound (greptile review)

Greptile flagged (P1) that a selected subagent child whose parent ranks below
the limit * 8 oversample is still promoted to a top-level row. That is real and
measured (the parent drops out at candidate #25 of 24), but it is the bound of
the design, not a regression: the walk reuses rows the projection already
fetched and never issues an extra query. Resolving arbitrarily old ancestors
needs an unbounded per-row lookup on the hot sidebar path -- the approach
rejected in #7031 -- so the bound is documented and pinned instead.

- Docstring: state that `limit` bounds the recency slice, not the row count,
  so callers must iterate rather than assume len(rows) <= limit; state that
  recovery is bounded by the oversampled candidate set.
- Inline comment: mark the bound at the exact lines Greptile flagged and point
  at `candidate_limit` as the knob if the window proves too tight.
- Tests: cover the candidate-window exhaustion Greptile said was untested --
  parent inside the oversample is recovered, parent beyond it stays unresolved
  -- plus the over-limit return contract and a parent-cycle guard.

No behaviour change; 7 tests pass.

* fix(share): add _hadAppearance guard to prevent fabricated appearance choice

Summary:
The inline appearance bootstrap in static/share.html writes hermes-theme and
hermes-skin to localStorage unconditionally on every page load, even when the
browser had no prior appearance state. This fabricates an explicit user choice
on first access via a shared link, making the server-side SETTINGS_DEFAULTS
unreachable for deployments that customise the default theme or skin.

Root Cause:
share.html:9 — the boot IIFE resolves a theme+skin and calls
localStorage.setItem() without guarding on whether the user had previously
chosen an appearance. The same bug was fixed in index.html by PR #6808
(commit tomtong2015) but share.html was left unchanged.

Change:
1. Added _hadAppearance guard before the two localStorage.setItem() calls:
   var _hadAppearance = localStorage.getItem('hermes-theme') !== null ||
                        localStorage.getItem('hermes-skin') !== null;
   if (_hadAppearance) { setItem('hermes-theme', t); setItem('hermes-skin', s); }
2. The first-paint DOM mutations (classList.add('dark'), dataset.skin) remain
   outside the guard — only persistence is protected.
3. Synced the skin allowlist with index.html: added neon-soft and neon-paint
   (zeus and verdigris were already present).

Verification:
- test_6808_appearance_bootstrap_no_fabricated_choice.py: 11/11 passed
  covering fresh-browser (no writes), pre-paint fallback, explicit state
  normalisation, and legacy migration survival.

Closes #7030

* fix(tests): pin rebuild budget in issue2513 custom-provider catalog test

test (3.13, 4) failed on this PR with:

  assert "@custom:alpha-proxy:alpha/remote" in alpha_ids
  E  AssertionError: assert '@custom:alpha-proxy:alpha/remote' in {'alpha/sticky'}
  WARNING api.config:config.py:8355 live provider-catalog rebuild exceeded
          4.0s budget - serving fallback, refreshing catalog out-of-band

Pre-existing wall-clock flake, not a regression from this PR: this branch
touches only api/agent_sessions.py and tests/test_subagent_parent_in_import_
window.py, and both api/config.py and this test file are byte-identical to
origin/master. The same shard passed on 3.11 and 3.12.

The test never pinned _LIVE_REBUILD_BUDGET_SECONDS, so it raced the global
4s budget in get_available_models(). On a starved runner the cold rebuild
overruns, the degraded fallback catalog is served, and the monkeypatched-
urlopen model alpha/remote is dropped - leaving only the config-declared
sticky model, exactly as CI observed.

Force the synchronous (unbounded) rebuild path, matching the existing
precedent in tests/test_issue2540_models_endpoint_error.py:20-24.

Verified: with the budget forced to 0.001s the unpatched test reproduces
the CI assertion verbatim; patched it passes at 0.001s, 0, and default.

* fix(#7013): preserve media deny coverage across platforms

* fix(docker): keep repository agents out of runtime context

* fix(docker): gate image runtime proof behind integration job

* test(models): stabilize custom provider catalog regression

* Release batch A: 6 low-risk gate-passed fixes (experimental)

Batched contributor fixes, each individually Codex-gated during the overnight
certifier cycles and re-verified clean-to-ship as a combined stage (Codex SAFE
TO SHIP on the combined diff; full suite green except 8 pre-existing approval
tests that fail identically on clean origin/master — CI green on same commit).

- #7019 (@webtecnica) suppress reserved [SILENT] sentinel at /api/goal
- #7031 (@carlotestor) keep evicted subagent parents in the sidebar import window
- #7030 (@webtecnica) share.html first-visit appearance guard + skin-id sync
- #6853 (@rodboev) exclude repo-root AGENTS.md from the Docker runtime image
- #7013 (@webtecnica) test-only: platform-neutral test portability (playwright
  import guard + media tests served from allowed roots)
- test-only (@carlotestor) pin custom-provider catalog rebuild budget (#7054)

Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
Co-authored-by: carlotestor <carlotestor@users.noreply.github.com>
Co-authored-by: rodboev <rodboev@users.noreply.github.com>

---------

Co-authored-by: webtecnica <webtecnica@gmail.com>
Co-authored-by: carlotestor <carlotestor@users.noreply.github.com>
Co-authored-by: Rod Boev <rod.boev@gmail.com>
Co-authored-by: n <a@n>
Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
Co-authored-by: rodboev <rodboev@users.noreply.github.com>
2026-08-16 21:20:56 +00:00
nesquena-hermesandnesquena-hermes dc3bf44e57 Release: SQLite source build loads on arm64 (ld.so.conf.d fix for #6900) (#7044) (#7045)
Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
2026-08-14 20:53:38 -07:00
nesquena-hermesandnesquena-hermes 6086446018 fix(docker): source-built SQLite must load on arm64 (ld.so.conf.d) — follow-up to #6900 (#7044)
* fix(docker): register /usr/local/lib in ld.so.conf.d so the source-built SQLite loads on arm64

The #6900 SQLite-from-source upgrade passed the amd64 docker-smoke but FAILED the
multi-arch release build on linux/arm64 with 'AssertionError: SQLite 3.46.1 still
vulnerable' — Python's sqlite3 kept loading the base image's /usr/lib multiarch
libsqlite3 (3.46.1) instead of the freshly-compiled /usr/local/lib copy (3.53.0),
because /usr/local/lib is NOT in the default ld.so search path on Debian arm64
(it happened to be picked up on amd64). Register /usr/local/lib via an
ld.so.conf.d entry before ldconfig so the new lib wins on every architecture.

The PR's own build-time version assertion is what surfaced this (correctly), and
it also serves as the verification that the fix works once the arm64 build passes.

Follow-up to #6900 (exp-v0.52.227, whose multi-arch image failed to build/push).

* test-fix: keep wal.html link within 500 chars of ARG (condense arm64 comment)

---------

Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
2026-08-14 20:51:42 -07:00
nesquena-hermesandnesquena-hermes 71584418a8 Release: upgrade bundled SQLite to 3.53.0 (WAL-reset fix, secure_delete preserved) (#6900, @qxxaa) (#7043)
Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
2026-08-14 20:16:20 -07:00
1aab1e4105 fix: compile SQLite 3.53.0 from source to fix WAL-reset corruption bug (#6900)
* fix: compile SQLite 3.53.0 from source to fix WAL-reset corruption bug

The python:3.12-slim base ships SQLite 3.46.1 (Debian Trixie), which is
vulnerable to the WAL-reset corruption bug discovered March 2026.
Debian has not backported the fix.

Compiles SQLite 3.53.0 from the official amalgamation tarball during the
Docker build. Installs to /usr/local/lib (takes ldconfig priority over
/usr/lib). Build tools (gcc, make, libc6-dev) are purged in the same
layer. A build-time Python assertion fails the image build if the linked
library is still vulnerable.

Version and year are build args for easy bumps.

https://sqlite.org/wal.html#walresetbug

* fix: enable FTS5/FTS4/R-Tree and pin tarball SHA-256

Address review feedback (nesquena-hermes):

1. Add --enable-fts5 --enable-fts4 --enable-rtree to ./configure so the
   compiled SQLite matches the distro package's feature set. Without
   FTS5, state.db session/message full-text search breaks silently.

2. Pin the amalgamation tarball SHA-256 as a build arg and verify with
   sha256sum -c before extraction.

3. Extend the build-time assertion to create and drop an FTS5 virtual
   table, proving the module is available - not just the version number.

4. Add regression tests for checksum verification, FTS5 configure flag,
   and FTS5 build-time vtable assertion.

* docker(sqlite): compile with SQLITE_SECURE_DELETE to preserve deleted-row erasure

Codex gate found a SILENT data-privacy regression: the Debian base image's SQLite is
built with SQLITE_SECURE_DELETE (PRAGMA secure_delete=1, deleted content overwritten),
but compiling 3.53.0 from the amalgamation without the flag drops it to 0 -> deleting a
session removes its rows but leaves transcript bytes recoverable in state.db. Reproduced
end-to-end via api/models.py delete path. Add CPPFLAGS=-DSQLITE_SECURE_DELETE and a
build-time assertion that PRAGMA secure_delete==1 (fails the image build if not).

Co-authored-by: qxxaa <qxxaa@users.noreply.github.com>

---------

Co-authored-by: qxxaa <qxxaa@users.noreply.github.com>
Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
2026-08-14 20:14:37 -07:00
nesquena-hermesandnesquena-hermes 7506884d29 Release: bound long-session memory to prevent tab-focus OOM (#7006, @webtecnica) (#7042)
Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
2026-08-14 19:44:38 -07:00
37c59baf68 fix(ui): prevent OOM on tab focus for long sessions (#7006)
* fix(ui): prevent OOM on tab focus for long sessions

Coordinate visibility-recovery and session-updated refresh paths so they
never start concurrent full-transcript loads, window the O(N) pre-render
scans to the virtual render window, and bound the render-window expansion.

Closes #6999

* fix(ui): full turn context + collision-free cache hash (#7006 review rework)

Addresses the maintainer's CHANGES_REQUESTED on PR #7006:

- Turn-context maps (_assistantTurnFinalVisibleContentMap /
  _assistantTurnVisibleContentMap) receive the FULL visWithIdx again.
  The virtualized render window (head+tail around a gap) cut through
  assistant runs and merged distinct turns across an omitted user
  boundary, which could duplicate the final answer in Worklog/Thinking
  or feed later-turn prose into the reasoning echo-strip. Turn context
  is input, not output: it must always be complete.

- _addBoundedHash now streams the FULL string form through the FNV-1a
  loop (no length+head+tail clip, no slice copies), and object fields
  (attachments, tc.args, compression-anchor keys) are walked key-by-key
  so no integral JSON.stringify() is allocated. Same-length middle-only
  edits of tool args, attachment metadata, tool snippets, or
  compression-anchor keys no longer collide in _sessionHtmlCache.

Regression tests:
- tests/test_issue6999_turn_context_virtual_gap.py: multi-assistant turn
  with a virtualized gap proves full-context echo-strip suppresses the
  exact final-answer echo, and documents the gapped head+tail failure.
- tests/test_issue2613_render_cache_signature.py: Node harness executing
  the real _messageRenderCacheSignature() proves same-length middle-only
  mutations change the signature for every structured field while the
  old clip scheme collides on the same data.

Closes #6999

* fix(ui): coalesce session-updated under refresh guard; insertion-order cache hash

Re-gate round 2 for #6999:
- session-updated frames arriving while the external-refresh guard is held
  are latched per SID (max announced count) and drained by the refresh
  owner's finally with ONE guarded follow-up, instead of being dropped
  (production does not guarantee a second event).
- _hashObjectInto walks keys in insertion order (matching Object.entries
  render projection) with an explicit per-key index discriminator, so
  opposite-insertion-order argument objects no longer share a cache
  signature; recursion depth is now threaded (children increment, not
  restart).

* test(jump-scroll): stub _loadingSessionId + _drainSessionUpdatedPendingCount in extraction harness

#7006 adds a reference to the pre-existing module var _loadingSessionId and a new
_drainSessionUpdatedPendingCount() helper inside refreshActiveSessionIfExternallyUpdated.
The jump-to-answer scroll-settle tests extract that function into a node harness, which
didn't declare them -> ReferenceError (19 parametrized failures). Add the missing stubs
to the harness preambles (no-op drain, since these tests exercise scroll ownership, not
the coalesce path -- real drain behavior is covered by test_issue6999_session_updated_coalesce.py).
Test-harness-only; production static/ untouched.

Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>

---------

Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
2026-08-14 19:42:15 -07:00
nesquena-hermesandnesquena-hermes 95f05c5147 Release: composer auto-dismisses transient Reconnected status pill (#7028, @starship-s) (#7041)
Co-authored-by: nesquena-hermes <agent@nesquena-hermes>
2026-08-14 18:34:28 -07:00
starship-sandnesquena-hermes 108841ba31 fix(composer): dismiss reconnected status (#7028)
Keep reconnect confirmations transient so they stop consuming mobile composer space, while a single owned timer prevents stale notices from clearing newer status text.

Assisted-by: Hermes Agent:gpt-5.6-sol
Assisted-by: Codex:gpt-5.6-luna

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-14 18:27:00 -07:00
1c4ea9c99a fix(appearance): don't let the boot script fabricate an explicit them… (#6808)
* fix(appearance): don't let the boot script fabricate an explicit theme choice

`syncSettings` in static/boot.js is careful to let the server win on a first
visit. Its comment says why:

    // empty (new-browser) state is indistinguishable from a user who chose
    // the defaults.  To avoid blocking server→client sync on first visit we
    // only let localStorage override the server when it carries an explicit
    // user-selectable theme value or a NON-DEFAULT skin.

The inline bootstrap in static/index.html defeats that. It resolves a theme with
`localStorage.getItem('hermes-theme')||'dark'` and then unconditionally writes the
result back:

    localStorage.setItem('hermes-theme',theme);localStorage.setItem('hermes-skin',skin);

On a brand-new browser that stores `hermes-theme=dark` before any request is made.
By the time syncSettings runs, `lsHasExplicitTheme` is true, so:

1. the server's `theme`/`skin` in SETTINGS_DEFAULTS (or settings.json) is ignored, and
2. the reconcile branch below POSTs the fabricated value back to /api/settings,
   overwriting the stored setting it was supposed to be reading.

The net effect is that the server-side appearance default is unreachable: a fresh
browser always lands on dark and then persists it server-side. Deployments that
set a different default in SETTINGS_DEFAULTS never see it applied.

Fix: only rewrite localStorage when there was already appearance state to
normalise. A fresh visit now leaves localStorage untouched, so `lsHasExplicitTheme`
is false and syncSettings takes the server value exactly as its comment intends.

Deliberately keeps the existing normalisation. The write is not pointless — it
canonicalises legacy names (`solarized` → `dark` + `poseidon`). Gating on "was
there any prior state" rather than per-key preserves that: with `hermes-theme=solarized`
and no `hermes-skin`, the derived skin is still written, so the mapping survives
the next load. A per-key guard would have silently dropped it.

Verified by driving the bootstrap in node against a stubbed localStorage. Behaviour
is byte-identical to before for every input that has prior state — legacy
`solarized`, an explicit `light`+`mono`, and a skin-only `slate` all produce the
same painted result and the same localStorage. The only case that changes is the
empty one, which is the bug:

    before   {} → localStorage {hermes-theme: dark, hermes-skin: default}   server locked out
    after    {} → localStorage {}                                            server wins

Painting is unchanged throughout: a fresh visit still renders dark, since that is
still the fallback.

* test(appearance): pin absent vs explicit appearance state in the bootstrap

Adds the regression coverage asked for in review. Nothing previously
distinguished "no appearance stored" from "user chose these values":
tests/test_1003_appearance_autosave.py covers picker autosave and
tests/test_bugfix_sweep.py covers guarded storage access, but neither fails if
the bootstrap starts fabricating a choice again.

Executes the REAL bootstrap, lifted out of static/index.html, in node against a
stubbed localStorage and documentElement — the same harness used to verify the
fix, now committed instead of run by hand. Covers all three states the review
called out, plus the pre-paint assertion:

  - both keys absent   -> no setItem at all, AND the dark class is still applied
                          (the fix changes persistence only, never the paint)
  - either key present -> both values normalised and persisted
  - legacy names       -> solarized/slate/monokai/nord/oled still migrate, and
                          the DERIVED skin is still written

That last group is why the guard is on the pair rather than per key: with
`hermes-theme=solarized` and no `hermes-skin`, the write of the derived skin is
the only thing that persists the mapping, so a per-key guard would silently drop
it on the next load.

Verified as a regression test, not a tautology: against the pre-fix
static/index.html, test_fresh_browser_writes_nothing and
test_guard_is_on_the_pair_not_per_key both FAIL, while the other nine still pass
— the unfixed bootstrap does still paint dark and does still persist
migrations, so only the buggy behaviour differs.

One harness note worth keeping: the driver and the extracted script are passed
as FILES (`node <driver> <script>`). Passing the script via `node -e ... --
<script>` shifts process.argv, so the driver eval'd the JSON scenario instead of
the bootstrap — and since `eval('{}')` is a valid empty block rather than a
syntax error, the empty-store case passed while executing nothing. _run() now
asserts the bootstrap produced some effect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(appearance): keep `var themes=` intact so the skin tests still find the block

CI caught this; review could not. `tests/test_sienna_skin.py` (x2) and
`tests/test_catppuccin_skin.py` locate the inline bootstrap with

    init_script_idx = INDEX_HTML.find("var themes=")
    end_idx = INDEX_HTML.find("</script>", init_script_idx)
    init_block = INDEX_HTML[init_script_idx:end_idx]

Declaring `_hadAppearance` as the first `var` in that chain turned `var themes=`
into `,themes=`, so `find()` returned -1, `init_block` came out EMPTY, and three
assertions failed against the empty string — including "Default theme must remain
'dark'", which read as though this PR had changed the default. It had not: the
default is untouched, the block simply could not be located any more.

Splitting the declaration (`;var themes=` instead of `,themes=`) restores the
literal. Both are `var` in the same function scope, so behaviour is identical —
the regression module still passes unchanged, including the empty-store case
(no writes, dark class still applied) and all five legacy migrations.

Failed on shards 3 and 4 across 3.11/3.12/3.13 while the same shards passed on
master, which is what pinned it to this change rather than flake.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Hermes Agent <hermes@aip.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-14 18:51:41 +00:00
nesquena-hermesandnesquena-hermes ae075ba188 docs(changelog): boot script no longer fabricates appearance choice on fresh browser (#6808) (#7029)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 18:41:06 +00:00
nesquena-hermesandnesquena-hermes dd44aa9830 docs(changelog): ja-locale profile wording consistency (#6762) (#7026)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 16:48:28 +00:00
asakuraandnesquena-hermes 155838f2c3 fix(i18n): unify ja locale profile wording (プロフィール → プロファイル) (#6762)
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-14 16:43:19 +00:00
nesquena-hermesandnesquena-hermes 0dc9cc381c docs(changelog): suppress reserved [SILENT] control turns at ingress (#7018) (#7020)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 12:41:19 +00:00
6d6560f286 fix(chat): suppress reserved SILENT control turns at ingress (#7018)
Cron agents use the exact final response `[SILENT]` as a delivery-suppression
sentinel. If a wake relay accidentally POSTs that sentinel to `/api/chat/start`
and 8701 restarts while the turn is pending, session repair materializes it as
a visible `{role: user, _recovered: true}` message. The pending value can then
be recovered again on later restarts, creating repeated `[SILENT]` turns.

Treat the exact normalized sentinel as a successful no-op at both server-side
turn entry points: the HTTP `/api/chat/start` handler and
`start_session_turn`. Both checks run before session lookup, runtime barriers,
or pending-state mutation. Matching is deliberately exact and case-sensitive,
so `[silent]`, prose containing `[SILENT]`, and ordinary user text are not
suppressed.

Add regression tests proving both paths return HTTP/status 200 without session
lookup, plus negative cases for non-exact text. The new tests fail 3/3 before
the fix. Targeted and neighbouring chat-start suites pass 21/21.

Co-authored-by: allenliang2022 <allenliang2022@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-14 12:39:20 +00:00
nesquena-hermesandnesquena-hermes 91356b7c3f docs(changelog): kanban sidebar-eviction fix (#6998) (#7016)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 10:41:12 +00:00
9b8a9e05d1 fix(sidebar): prevent kanban runs from hiding CLI sessions (#6998)
* fix: prevent kanban rows from hiding CLI sessions

Keep worker-heavy Kanban history out of the bounded interactive session window and recover it through a separate capped pass. Add regression coverage and update query-pass contract tests.

* docs: describe imported session projection bounds

---------

Co-authored-by: zicochaos <23367003+zicochaos@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-14 10:39:36 +00:00
nesquena-hermesandnesquena-hermes b58ea15e31 docs(changelog): translate workspace-sort keys in 14 locales (#6910) (#7014)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 05:40:23 +00:00
webtecnicaandnesquena-hermes 3d18a8ca5a i18n: translate workspace-sort keys in 14 non-English locales (#6902) (#6910)
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-13 22:39:25 -07:00
nesquena-hermesandnesquena-hermes 69fb748f73 docs(changelog): streaming session-scope render guard (#6502) (#7012)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 05:03:33 +00:00
6261b626a7 fix(streaming): scope streaming message rendering per-session to prevent cross-session render leak (#6449) (#6502)
When the user switches from a streaming session (A) to another session (B),
A's streaming text could render in B's window because:

1. The token handler guards against rendering when the session is not active,
   but the rAF/setTimeout window between _scheduleRender() and _doRender()
   can outlive a session switch.
2. _doRender() writes to assistantBody without checking whether the stream's
   session is still the active frontend pane.

Fix: add _isActiveSession() guards to:
- _scheduleRender() — drop scheduled renders when the session changed
- _doRender() — bail before any DOM write if the session switched
- _flushPendingSegmentRender() — defense-in-depth for any future callers

These guard at the same level the token/interim_assistant event handlers
already do, closing the rAF gap that could leak streamed text into a
different session's view.

Co-authored-by: webtecnica <contato@webtecnica.com.br>
Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-13 22:02:33 -07:00
nesquena-hermesandnesquena-hermes fe8fc067aa docs(changelog): first-response jump scroll-settle fix (#6621) (#7011)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 04:18:49 +00:00
3a95d82531 fix(chat): prevent first response jump from snapping back (#6621)
* fix(chat): keep response jump position after session load

Cancel pending load-time bottom settling before an explicit response jump takes scroll ownership. Add a regression test covering the first-click race.

* fix: gate response jump ownership by destination

* fix: keep response jumps programmatic through smooth scroll

Native smooth-scroll frames from response jumps could reach the manual scroll listener and claim reader ownership before the final destination was known.

- own response-jump scrolling with a generation- and session-scoped lifecycle
- reconcile sticky ownership from final 80px tail geometry
- preserve current low-delta wheel takeover semantics when integrating with master
- cover visible assistant, user-row, and virtualized paths through the production scroll listener

* fix(scroll): hold reader off-bottom during jump owner + cancel jump on explicit End (#6621 gate fixes)

Codex gate found two defects in the jump-scroll ownership mechanism:
- Jump during an active stream snapped back to bottom on the next token: the
  owner suppressed the scroll listener but left the pre-jump pinned state, so
  scrollIfPinned() reclaimed the bottom. Now _beginMessageJumpScroll unpins for
  the ownership window and scrollIfPinned() no-ops while an owner exists.
- An explicit End click could be undone by the pending jump reconcile restoring
  the stale unpinned snapshot. scrollToBottom() now cancels the active jump
  owner first; _cancelMessageJumpScroll restores the preserved snapshot so a
  non-reconcile cancel doesn't leak the transient unpinned state.

* fix(scroll): drop undefined currentSid global ref in _messageJumpSessionId (#6621 brick-class scope-undef gate)

The PR's _messageJumpSessionId() referenced a bare 'currentSid' global that
does not exist in ui.js (everywhere else it's a local const from
S.session.session_id). test_static_js_scope_undef flagged it brick-class
(#3696). The S.session.session_id fallback already IS the canonical accessor;
removed the dead first line and updated the harness to drive the session-change
case via S.session.session_id.

* fix(scroll): round-2 gate fixes for jump-owner state transitions (#6621)

Re-gate found three more state-transition edge cases from the temp-unpin window:
- wheel-up interrupting an active jump owner after the programmatic latch
  expires now explicitly establishes the unpinned reader-owned state (was
  cancelling the jump but leaving pinned -> next token snapped to bottom).
- _resetStreamScrollFollow() now cancels the jump owner FIRST, before its
  pinned-state assignments, so a stream starting mid-jump can't have its pin
  undone by the snapshot restore -> auto-follow stays enabled.
- _finishMessageJumpScroll() flushes a deferred external-session refresh after
  reconciliation when the terminal state is pinned to the tail, so a refresh
  deferred during the temp-unpin window isn't stranded.

* fix(scroll): make jump-owner guards typeof-safe + widen static test window (#6621)

The two _messageJumpScrollOwner guards (scrollIfPinned, scrollToBottom) are
pulled into other scroll test harnesses (test_issue6414, test_issue4856) that
stub the scroll env without declaring the new global; bare references threw
ReferenceError. Guard with typeof (matches the PR's own jumpScrollOwned check).
Also widen test_low_delta_wheel_intent_is_tracked_separately's source-slice
window 1400->1800 to still contain the (unchanged) _lastMessageWheelIntentMs
line after the +9-line wheel-up-during-jump block was inserted.

* test(scroll): add streaming-frame-hold + wheel-during-jump regression tests (#6621 Fable finding ii)

Fable UX gate flagged that the PR's tests model only the stale load-time settle
callback, not the streaming case. Add two node-harness tests:
- a streaming render frame (scrollIfPinned) fired inside the jump-owner window
  must not snap the reader to the bottom (holds at target, 0 bottom-writes).
- a gentle wheel-up during the owner window after the programmatic latch stales
  hands ownership to the reader UNPINNED at their position, never pinned
  mid-transcript.

* fix(scroll): widen jump-owner takeover to downward + touch input (#6621 Fable S3)

Fable's re-review (on the fixed code) confirmed the streaming blocker + wheel-up
concern resolved, and probe-proved one remaining narrow regression: a gentle
wheel-DOWN or touch scroll during the owner window cancelled the owner and
restored the pinned snapshot but did NOT re-unpin (the takeover was gated on
wheelUp only), so the next streaming token yanked the reader to the bottom.
Widen the takeover to any message-pane scroll input during the owner window.
Add a wheel-down regression test; widen the #4970 static-slice window to 2000
to still contain the (unchanged) _lastMessageWheelIntentMs line.

---------

Co-authored-by: pxxD1998 <214340659+pxxD1998@users.noreply.github.com>
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-13 21:17:42 -07:00
nesquena-hermesandnesquena-hermes 5a68684fc7 docs(changelog): 402 exhausted-TTL fast recovery (#6626) (#7010)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-14 03:16:43 +00:00
webtecnicaandnesquena-hermes 7a5e6bc27a fix(providers): reduce pool-exhaustion TTL for HTTP 402 from 1h to 2min (#6626)
* fix(providers): reduce pool-exhaustion TTL for HTTP 402 from 1h to 2min

Mirrors the agent-side fix: when a provider returns HTTP 402 (Payment
Required), the credential pool marks the key as exhausted with a 1-hour
TTL — even though recharging credits restores the key immediately.

Both copies of _entry_exhausted_ttl_seconds() now handle HTTP 402 with
a 2-minute TTL, matching the new EXHAUSTED_TTL_402_SECONDS constant in
agent/credential_pool.py.

* fix(providers): tie HTTP 402 TTL to installed runtime contract (#6626)

* test(providers): behavior-test 402 TTL 119/120s boundary on both helper copies (#6626)

Add executed boundary tests for the HTTP 402 pool-exhaustion TTL:

- Host helper: with the runtime TTL monkeypatched to 120, an exhausted
  entry is still unavailable at 119s and becomes available at 120s, for
  both the integer 402 and the persisted string "402" error codes.
- Embedded _ACCOUNT_USAGE_SUBPROCESS_CODE copy: execute the subprocess
  code under the existing quota-subprocess harness and assert the same
  TTL and 119/120s eligibility boundary against the embedded helper,
  instead of counting source-string occurrences.
- Keep the legacy 3600s fallback and last_error_reset_at precedence
  assertions.

* test(providers): make 402 TTL parity tests self-contained for standalone CI

The parity tests imported agent.credential_pool directly, but the WebUI
test suite runs standalone (Hermes Agent runtime is not installed in CI),
causing ModuleNotFoundError on every shard since 08-02. Simulate the
runtime contract with a fake agent.credential_pool module (same pattern
already used by the embedded-subprocess helper test) so both the 120s
cooldown and the legacy 1h fallback paths are covered.

---------

Co-authored-by: nesquena-hermes <nesquena+hermes@gmail.com>
2026-08-13 20:15:30 -07:00
nesquena-hermesandnesquena-hermes 7c1239dbe7 docs(changelog): stamp v0.52.113 stable section (promotion) (#7009)
Co-authored-by: nesquena-hermes <nesquena-hermes@users.noreply.github.com>
2026-08-13 19:43:40 -07:00
ahmadalguydi b5acb01833 fix(insights): preserve compact period selector width 2026-08-12 22:07:35 +03:00
asorourx cbe4502c46 Allow legitimate dotfiles in extension archives
_is_safe_relative_path rejected any path segment starting with '.', which
blocks hidden leaf files like .gitkeep / .env.example — not just the ./..
traversal segments. This breaks install of any extension shipping a dotfile,
including the gallery's own profile-avatars (screenshots/.gitkeep) -> "Unsafe
archive member".

Drop the over-broad startswith('.') clause; the '.'/'..' guard plus the
caller's resolved.relative_to(root) containment check already cover traversal.
Verified: .gitkeep is accepted while ../evil and a/./b stay blocked.

Fixes #6619
2026-08-06 03:41:50 +00:00
Hermes AgentandClaude Opus 5 a4491c568e fix(passkeys): derive the RPID from Origin, not Host
When a reverse proxy rewrites Host to its own upstream address — which
happens whenever a separate SPA front-end proxies /api to this backend —
rp_context() derived an RPID that never matched the origin the browser was
actually on, so WebAuthn's RPID-origin check rejected every registration
and login.

Prefer the Origin header's hostname, which proxies do forward, and fall
back to Host for direct access.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 08:54:18 -05:00
ahmadalguydi aaebb62758 fix(insights): preserve period selector chevron 2026-08-05 15:54:55 +03:00
webtecnica c103ab4d32 Add #6247: doc-ref the proposal and merged implementation
Refs #6247 — the public conversation lifecycle regression gate proposal.
Refs #6251 — the merged first implementation slice (normal Live-to-Final row).
2026-07-19 10:31:12 -03:00
igorroncevic 037c6cf67a fix: harden child session attach ordering 2026-07-11 00:41:43 +02:00
igorroncevic 2f5c911e2f hermes-webui: sort child sessions so read-only delegates are last 2026-07-10 18:13:15 +02:00
235 changed files with 39031 additions and 2533 deletions
+1
View File
@@ -5,3 +5,4 @@ __pycache__
*.pyo
tests/
.env*
/AGENTS.md
+170
View File
@@ -87,6 +87,176 @@ jobs:
cache-from: type=gha
cache-to: type=gha,mode=max
# Startup-health gate for UID/GID auto-detection (#7027).
#
# The compose variants above all mount a hermes-home volume, so they never
# exercise the plain single-container shape from the issue: a state directory
# bind-mounted from a host directory owned by a non-1024 user, /workspace left
# as the image's own build-time 1024, and no explicit WANTED_UID/WANTED_GID.
# Before the fix the auto-detector read /workspace, remapped to 1024, failed
# its own state-dir writability check and restart-looped. Source-level
# invariants can't catch that — only actually booting the container can.
state-dir-uid:
name: State-dir UID detection (#7027)
runs-on: ubuntu-latest
needs: [compose-config, build-image]
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: docker/setup-buildx-action@v3
- name: Restore Docker image from cache
uses: docker/build-push-action@v6
with:
context: .
load: true
tags: ghcr.io/nesquena/hermes-webui:latest
cache-from: type=gha
- name: Boot on a host-owned state mount with no explicit IDs
run: |
set -euo pipefail
CONTAINER="hermes-smoke-uid-${{ github.run_id }}-${{ github.run_attempt }}"
HOST_UID=1001
STATE_DIR="$(mktemp -d -t hermes-smoke-uidstate-XXXXXX)"
sudo chown -R "$HOST_UID:$HOST_UID" "$STATE_DIR"
cleanup() {
local rc=$?
echo "::group::Cleanup (rc=$rc)"
docker logs "$CONTAINER" 2>&1 | tail -200 || true
docker rm -f "$CONTAINER" || true
sudo rm -rf "$STATE_DIR" || true
echo "::endgroup::"
return $rc
}
trap cleanup EXIT
# Deliberately NOT mounting /workspace: the point of this gate is that
# the stock image's own /workspace (owned 1024:1024) must not win over
# the host-owned state mount.
echo "::group::docker run"
docker run -d --name "$CONTAINER" \
-p 127.0.0.1:8788:8787 \
-e HERMES_WEBUI_STATE_DIR=/app/data \
-v "$STATE_DIR":/app/data \
ghcr.io/nesquena/hermes-webui:latest
echo "::endgroup::"
# ----- The product gate: the container must reach /health -----
echo "::group::Probe /health"
attempts=0
max_attempts=60
until curl --fail --silent --max-time 5 http://127.0.0.1:8788/health > /dev/null; do
attempts=$((attempts + 1))
if [ "$attempts" -ge "$max_attempts" ]; then
echo "❌ WebUI /health never returned 200 after $max_attempts attempts (~5m)"
echo " Startup logs follow in the cleanup group below."
exit 1
fi
if [ "$(docker inspect -f '{{.State.Running}}' "$CONTAINER")" != "true" ]; then
echo "❌ Container exited before /health came up — this is the #7027 restart loop"
exit 1
fi
sleep 5
done
echo "✅ /health = 200 after $attempts attempts"
echo "::endgroup::"
echo "::group::Verify the detected identity"
LOGS="$(docker logs "$CONTAINER" 2>&1)"
# The identity must have been detected from the state dir, not /workspace.
if ! echo "$LOGS" | grep -qE -- "-- Auto-detected UID: ${HOST_UID} \(from /app/data\)"; then
echo "❌ UID was not auto-detected from the state directory (#7027)"
echo "$LOGS" | grep -E -- "-- (Auto-detected|WANTED_)" || true
exit 1
fi
if ! echo "$LOGS" | grep -qE -- "-- Auto-detected GID: ${HOST_UID} \(from /app/data\)"; then
echo "❌ GID was not auto-detected from the state directory (#7027)"
echo "$LOGS" | grep -E -- "-- (Auto-detected|WANTED_)" || true
exit 1
fi
# And the runtime user must really carry it.
ACTUAL_UID="$(docker exec "$CONTAINER" id -u hermeswebui)"
ACTUAL_GID="$(docker exec "$CONTAINER" id -g hermeswebui)"
if [ "$ACTUAL_UID" != "$HOST_UID" ] || [ "$ACTUAL_GID" != "$HOST_UID" ]; then
echo "❌ hermeswebui runs as ${ACTUAL_UID}:${ACTUAL_GID}, expected ${HOST_UID}:${HOST_UID}"
exit 1
fi
echo "✅ hermeswebui runs as ${ACTUAL_UID}:${ACTUAL_GID}, detected from the state mount"
echo "::endgroup::"
echo "::group::Startup log scan"
BAD_PATTERNS='!! ERROR|!! Exiting script|Failed to verify state directory|Permission denied'
if echo "$LOGS" | grep -E -i "$BAD_PATTERNS"; then
echo "❌ Startup logs contain known-bad pattern (see above)"
exit 1
fi
echo "✅ No known-bad patterns in startup logs"
echo "::endgroup::"
- name: Explicit WANTED_UID=1024 must survive auto-detection
run: |
set -euo pipefail
CONTAINER="hermes-smoke-uid1024-${{ github.run_id }}-${{ github.run_attempt }}"
STATE_DIR="$(mktemp -d -t hermes-smoke-uid1024-XXXXXX)"
HERMES_DIR="$(mktemp -d -t hermes-smoke-uid1024home-XXXXXX)"
# Both mounts are owned by a *different* UID than the one requested.
# The hermes-home mount is what makes this discriminating: it is a
# probe candidate the pre-fix script already read, so if the explicit
# 1024 is still treated as "unset" the container remaps to 1001 and
# the assertion below fails. World-writable state dir so the only
# thing under test is which identity wins, not mount permissions.
sudo chown -R 1001:1001 "$STATE_DIR" "$HERMES_DIR"
sudo chmod 0777 "$STATE_DIR"
cleanup() {
local rc=$?
echo "::group::Cleanup (rc=$rc)"
docker logs "$CONTAINER" 2>&1 | tail -100 || true
docker rm -f "$CONTAINER" || true
sudo rm -rf "$STATE_DIR" "$HERMES_DIR" || true
echo "::endgroup::"
return $rc
}
trap cleanup EXIT
docker run -d --name "$CONTAINER" \
-p 127.0.0.1:8789:8787 \
-e HERMES_WEBUI_STATE_DIR=/app/data \
-e WANTED_UID=1024 \
-e WANTED_GID=1024 \
-v "$STATE_DIR":/app/data \
-v "$HERMES_DIR":/home/hermeswebui/.hermes \
ghcr.io/nesquena/hermes-webui:latest
attempts=0
max_attempts=60
until curl --fail --silent --max-time 5 http://127.0.0.1:8789/health > /dev/null; do
attempts=$((attempts + 1))
if [ "$attempts" -ge "$max_attempts" ]; then
echo "❌ WebUI /health never returned 200 with explicit WANTED_UID=1024"
exit 1
fi
if [ "$(docker inspect -f '{{.State.Running}}' "$CONTAINER")" != "true" ]; then
echo "❌ Container exited before /health came up"
exit 1
fi
sleep 5
done
ACTUAL_UID="$(docker exec "$CONTAINER" id -u hermeswebui)"
if [ "$ACTUAL_UID" != "1024" ]; then
echo "❌ explicit WANTED_UID=1024 was overwritten by auto-detection — got ${ACTUAL_UID} (#7027)"
exit 1
fi
echo "✅ explicit WANTED_UID=1024 preserved (sentinel no longer shadows the valid value)"
smoke:
name: Smoke ${{ matrix.variant }}
runs-on: ubuntu-latest
+42
View File
@@ -59,6 +59,7 @@ actions. The topbar remains focused on conversation context and the workspace/fi
auth.py Optional password authentication, signed cookies, passkeys/WebAuthn
config.py Discovery, globals, model detection, reloadable config
helpers.py HTTP helpers: j(), bad(), require(), safe_resolve(), security headers
goals.py Persistent-goal commands and profile-scoped native GoalManager bridge
models.py Session model + CRUD, per-session profile tracking, CLI/state.db bridge
profiles.py Profile state management, hermes_cli wrapper
onboarding.py First-run onboarding status, real provider config writes, OAuth linking, readiness detection
@@ -235,6 +236,24 @@ Session is a plain Python class (not a dataclass, not SQLAlchemy):
title_from(): takes messages list, finds first user message, returns first 64 chars.
Called after run_conversation() completes to set the session title retroactively.
#### Imported `state.db` sidebar projection
`api.models.get_cli_sessions()` projects conversations from the active Hermes
profile's `state.db` into sidebar-shaped rows. The default projection keeps
interactive sources (CLI, TUI, ACP, messaging, and similar user-facing sessions)
in a bounded 20-row candidate window. Background sources use independent recovery
passes so a high-volume worker source cannot consume that interactive window:
- Cron: up to `CRON_PROJECT_CHIP_LIMIT` rows.
- Webhook: up to `WEBHOOK_PROJECT_CHIP_LIMIT` rows.
- Kanban: up to `KANBAN_PROJECT_CHIP_LIMIT` rows.
Source-specific views still use their dedicated bounds, and the later sidebar
visibility stage decides whether recovered background rows are shown. In
`all_profiles=True` mode the per-profile source bounds are disabled before rows
are merged; cross-profile scoping, visibility, deduplication, and final route
limits remain downstream responsibilities.
### 4.3 SSE Streaming Engine
This is the most architecturally interesting part. Two endpoints cooperate:
@@ -376,6 +395,29 @@ read_file_content(workspace, rel):
- Reads as UTF-8 with errors='replace' (binary files show replacement chars)
- Returns {path, content, size, lines}
### 4.8 Persistent Goal Profile Boundary
`api/goals.py` exposes the WebUI `/goal` command payloads and post-turn evaluation hook.
Hermes Agent's native `GoalManager` is the authoritative owner of goal evaluation,
continuation decisions, wait barriers, failure counters, contracts, subgoals, and
`state.db` persistence.
For a profile-scoped WebUI session, the bridge delegates only when the Agent exposes
both the context-local `set_hermes_home_override()` API and call-time default
`SessionDB` path resolution. The bridge probes the resolved default path under the
selected context before constructing the native manager, then binds that profile's
Hermes home before every native call. The override is reset in a `finally` block after
every operation, so concurrent sessions using the same session ID under different
profiles cannot cross-read or cross-write goal state. Goal snapshot rollback uses the
same scoped native persistence path.
Older Hermes Agent versions that lack either capability continue through
`_LegacyProfileGoalManager`, which pins persistence to the selected profile's explicit
`state.db` path. This includes intermediate versions whose context API is present but
whose default `SessionDB()` path remains frozen at module import. Keep this fallback
compatibility-only: new goal semantics belong in Hermes Agent's native manager rather
than a second WebUI implementation.
---
## 5. Frontend Architecture: Current State
+292 -65
View File
@@ -3,8 +3,216 @@
## [Unreleased]
### Changed
- **The chat composer grows natively instead of being resized by JavaScript on every keystroke.** `autoResize()` measured `scrollHeight` and wrote `style.height` on each input event — a forced synchronous reflow on the most-typed-in surface in the app. Browsers that support CSS `field-sizing: content` (Chromium today; also Firefox 152 and Safari 26.2) now own the geometry directly, gated on `CSS.supports()`, and the existing JavaScript path is untouched for every other engine. Measured behaviour is identical across both paths: 44px resting height, no jump when the first character is typed or the last deleted, growth to the 200px ceiling, then internal scrolling. Because `field-sizing` deliberately includes placeholder text in content sizing, `:placeholder-shown` pins fixed sizing while the composer is empty so a long placeholder can't inflate it. Thanks @starship-s. (#6760, #5514)
### Fixed
- **Structured (multimodal) message content stored in `state.db` renders as its real content instead of a placeholder, without dropping or misattributing turns.** When the session projection decodes the `state.db` content sentinel, list-valued content now gets an out-of-band identity — a `("structured_content", canonical)` tuple that no scalar string can compare equal to — shared across the merge, dedup, content, and visible reconciliation keys, so an in-band string can't collide with a rich row and an image-only row can no longer vanish from the projection. The expensive canonical serialisation is memoised per content object for the duration of a single `merge_session_messages_append_only()` call (scoped through a `ContextVar`, keyed by object identity with the object held alive so an id can't be reused for a different list mid-call), restoring one serialisation per list per call. Empty/falsy content keys exactly as before. Thanks @totalitarian. (#7492)
- **An empty composer no longer inflates to fit its own placeholder on the JavaScript sizing path.** The fallback resizer measured `scrollHeight` for an empty textarea, which reflects the *placeholder* — so a long busy or compression hint (which wraps to two or three lines) grew the empty composer to roughly 71px, and only after a resize that happened while empty, making the resting height depend on history rather than content. The inline height is now cleared when the value is empty so CSS `min-height` defines the resting size, matching the native path. (#6760)
- **Delegated subagent sessions nest under the conversation that spawned them instead of landing at the top of the sidebar.** The sidebar forces a *cross-surface* child row to top level when its parent row belongs to another surface (a Telegram/messaging parent), so an independent continuation doesn't get stacked under a foreign thread. A delegated subagent is also technically cross-source relative to its WebUI parent, so it was swept up by that rule and surfaced as a bare top-level row with no indication of what spawned it. Delegated subagents are now exempt from that branch and attach to their visible parent, while independent cross-surface continuations still stay top level. The delegated role is resolved in strict precedence order from the first non-blank of `raw_source` → `source_tag` → `source` and additionally requires the backend's child-session flag, so a stale lower-priority field can't promote an independent row into a delegated one. Thanks @lowlandsheperd. (#7263)
- **Remote (SSH/Docker) workspaces keep their target-side POSIX paths across session restore, streaming, uploads and project context.** When the terminal backend runs off-host, a workspace path such as `/srv/remote-alice` belongs to the *target* machine — but several resolvers normalized it against the WebUI host's filesystem, so a host-side expansion (macOS firmlinks turning it into `/System/Volumes/Data/srv/remote-alice`, or a local `~` expansion) could be persisted back onto the session and then used as the per-turn runtime CWD. The workspace, streaming, upload and project-context resolvers now receive the owning **session's profile** explicitly instead of resolving against whatever profile happens to be ambient, so remote paths stay target-side and a session owned by one profile can never resolve through another's. Local (non-remote) workspaces normalize exactly as before, and symlink/`..`/outside-root rejection is unchanged. Thanks @alfred-rootson. (#7168, closes #7131)
- **Signal gateway conversations are classified as messaging sessions instead of falling into the default/other bucket.** The messaging-source allowlists never learned about Signal, so a session started on the Signal gateway was normalized to `other`, failed the client-side external-session check the session list uses to render and refresh gateway sessions, and never showed up in the sidebar. `signal` (labeled "Signal") is now in the messaging-source sets on both sides — the Python normalizer (`MESSAGING_SOURCES` / `SOURCE_LABELS`), which also feeds the server-side messaging allowlist, and the client-side classifier (`_MESSAGING_RAW_SOURCES` / `_MESSAGING_SOURCE_LABELS`) — so Signal sessions group, label, and list exactly like Telegram, Discord, Slack, WeCom, and Matrix. No other source's classification changes. Thanks @webtecnica. (#7571)
- **Opening a session no longer stalls just because Codex refreshed its model cache in the background.** The WebUI keys its 24-hour `/api/models` cache partly on Codex's `~/.codex/models_cache.json`, but that file was fingerprinted by `mtime` + size — and Codex rewrites it on its own refresh timer, bumping those stat fields even when the model catalog is byte-identical (only a volatile `fetched_at` timestamp changed). Every Codex refresh therefore invalidated the cache, and the next session open paid a full live rebuild whose serial provider probes stalled the open (#7540). The Codex cache is now fingerprinted by its *content* with the refresh-timestamp fields (`fetched_at`, `updated_at`) stripped, so a timestamp-only rewrite keeps the cache while any genuine catalog change (models, `etag`, `client_version`) still invalidates it. A malformed or pathologically deep cache file degrades safely to the old stat fingerprint rather than erroring. Thanks @webtecnica. (#7556, closes #7540)
- **A timed-out command-approval card can be dismissed again instead of getting stuck on screen forever.** When the agent raised a local approval with no gateway notifier registered, the raw pending entry was stored without an `approval_id`; the frontend needs that id to own the card, so a card that timed out (or otherwise went unanswered) could be neither approved nor dismissed and the poll re-served it indefinitely. Reconciliation now mints a stable `gwlocal-mirrorless:` id on the raw local entry itself — the same tokenization contract already used for gateway producers — so the id survives repeated polls and the client's dismiss marker sticks. Gateway/run-bearing entries are untouched, and the synthetic prefix can't be mistaken for gateway authority (routing keys on exact ids, `run_id`, and mirror tokens, never the prefix). Thanks @CharlesMcquade. (#7510)
- **Manual "Regenerate title" now works for conversations that open with several queued user messages.** For a session whose transcript begins with consecutive user turns (no assistant reply before the second user message), the first-exchange walker stopped at the second user row with no assistant text, so "Regenerate title" skipped the LLM call, silently persisted the local fallback (a truncated first message), and returned `200` with that same wrong title on every retry — leaving the session permanently stuck. Manual regeneration now scans past the opening user rows to the first complete user+assistant pair so it reaches the aux model. The automatic in-stream title path is deliberately unchanged. Thanks @Raylands. (#7545, closes #7543)
- **Custom providers that Hermes auto-discovered now show their full live `/v1/models` catalog again instead of just the saved models.** When a provider entry carries `models_discovered: true` alongside a `models:` mapping of per-model metadata (vision, context length), that mapping is discovery metadata — not a hand-curated allowlist — but the WebUI treated it as one and suppressed the live catalog, collapsing the model picker to only the already-saved IDs. Such a provider now defers to its live `/v1/models` catalog (matching the Hermes Agent's own `_models_config_is_allowlist` semantics), while an explicit `discover_models: false` opt-out (including the Agent-compatible string forms `false`/`no`/`0`) still pins the configured list, and a transient empty live probe falls back to the configured models rather than dropping the provider from the picker. Thanks @evgenyponomarev. (#7409, closes #7404)
- **Ctrl-C and `ctl.sh`-launched daemons now shut down gracefully instead of leaving the server running.** The WebUI installed a SIGTERM handler that drains in-flight work through `serve_forever()`'s `finally`, but SIGINT fell through to Python's default, so a daemon started via `./ctl.sh start` could only be stopped with `kill <pid>` and the Settings "Stop server" button did not stop it. SIGINT now shares the same idempotent graceful-shutdown handler as SIGTERM. Thanks @aniruddhaadak80. (#7315, closes #7078)
- **Static assets are served with correct MIME types instead of a `text/plain` catch-all.** `_serve_static()` defaulted every extension outside its small built-in allowlist to `text/plain; charset=utf-8`, mislabeling PDFs, APKs, WASM and other binaries. Unknown extensions now resolve through Python's standard `mimetypes` database, with an explicit deterministic entry for `.apk` (whose type varies across host MIME databases), and fail closed to `application/octet-stream` for unknown or pre-encoded suffixes (`.svgz`/`.tgz`/`.js.gz`) so still-compressed bytes are never advertised with their decoded media type. Path containment, gzip/ETag/cache handling, and the text charset path are unchanged. Thanks @Oidaladn0. (#7372)
- **Chat-attachment uploads are created atomically so a raced duplicate or symlink can't silently overwrite or redirect the write.** `handle_upload` deduped filenames with an `exists()` check and then wrote with a plain `write_bytes`, so two concurrent uploads of the same name both passed the check and the last writer silently won — leaving one client a `200` naming bytes no longer on disk. The final create now uses the existing descriptor-anchored `open_anchored_create_fd` (`O_CREAT|O_EXCL|O_NOFOLLOW`, mirroring the #3398 workspace-upload hardening): a raced duplicate returns `409` instead of overwriting, and a symlink raced into the attachment path cannot redirect the write outside the session inbox. The normal success path, size/MIME handling, and `-1` dedup suffixing are unchanged. Thanks @zicochaos. (#7337)
- **Cron jobs can be delivered to your messaging platforms again.** `GET /api/crons/delivery-options` had been returning only `local` and `origin`, so Telegram, Discord, Slack, Feishu and every other platform was silently missing from the cron delivery picker — a scheduled job could not be pointed anywhere. The Agent moved its delivery-platform allowlist to a new module, and the endpoint still imported the old path inside a bare `except` that fell back to an empty set, so the import error was swallowed and the feature degraded with nothing logged. The allowlist is now resolved across both module paths so the picker survives future relocations, with a regression test that fails loudly instead of emptying out. (#7525)
- **Auxiliary-model settings no longer silently discard a configured provider that is missing from the model catalog.** The provider dropdown was built only from providers the catalog returned, so when a task's configured provider was absent (for example its group exposes no models), nothing matched, the select fell back to its first entry (`auto`), and the next Apply persisted `auto` — quietly throwing away the user's choice on a row they never touched. The configured provider is now kept selectable so it round-trips through Apply unchanged. Thanks @webtecnica. (#7519, closes #7486)
- **A custom provider whose model endpoint is unreachable no longer disappears from the auxiliary model picker.** A named provider whose `/v1/models` probe fails still reaches the client as a group with no models plus an endpoint-error marker, and the main model picker renders it with an unreachable-endpoint hint — but the auxiliary picker filtered out every group without models, so the provider vanished from those selects entirely. The auxiliary picker now keeps it and shows the same hint instead of an empty dropdown. Thanks @webtecnica. (#7522, closes #7521)
- **The OpenRouter setup wizard now offers Z.AI models under OpenRouter's own namespace instead of IDs that 404.** `_FALLBACK_MODELS` is authored for the *direct* provider endpoints, so its nine Z.AI entries carry the provider-native `zai/` prefix — and the OpenRouter onboarding setup reused that list verbatim. OpenRouter serves the same models as `z-ai/…`, so anyone who picked a Z.AI model during OpenRouter onboarding was handed an ID the provider does not recognise and their first message failed. The IDs are now translated at the OpenRouter projection boundary only; the fallback catalog and the direct Z.AI setup keep their own namespace. Thanks @webtecnica. (#7518, closes #7514)
- **Recovered "thinking" cards no longer multiply on every turn after context compaction.** When a compacted history left the display transcript longer than the model-facing result, a recovered reasoning-only row could be re-inserted *after* the new reply each turn — doubling the recovered block geometrically (1, 2, 4, 8…). Settlement now fixes the active-turn boundary before restoring reasoning metadata and only restores a historical reasoning row directly in front of its own API-safe successor (requiring a unique stable-id or exact-identity match), failing closed to the current-turn boundary when turn ownership can't be proven. The synchronous chat path mirrors the streaming path, so live and reloaded transcripts stay consistent. Thanks @carlotestor. (#7443, refs #7388)
- **Command-backed provider credentials no longer break the agent cache or discard conversation state between turns.** When a provider used `key_cmd`, the runtime supplied a lazy callable token source, but the cache signature tried to call `.encode()` on it; every turn therefore missed the cache and rebuilt the agent. Callable credentials now use a stable, secret-free cache marker while the existing runtime refresh path still receives the callable itself, so token resolution stays lazy and agent context is reused. Thanks @ArushKhasru. (#7398, closes #7396)
- **Agent turns now run in the workspace you selected instead of the server's install directory.** Because the WebUI runs the agent in-process, anything that resolved a default working directory from the process (`os.getcwd()`) landed in the Hermes install tree rather than the chosen workspace — so conductor children could record the wrong working directory and even pick up the install tree's own `AGENTS.md` as workspace doctrine. Each turn now binds its workspace to a task-local runtime cwd, so concurrent turns keep their own workspace and the binding is restored cleanly when the turn ends (a malformed workspace path degrades gracefully instead of failing the turn). Thanks @samfoy. (#7351)
- **A stale WebUI runtime after an Agent update is now reported and rejected without an unsafe automatic restart, and its update-marker read can no longer hang or exhaust memory.** When the running Agent revision no longer matches the on-disk checkout, WebUI keeps returning the manual-restart `409` and now surfaces richer update diagnostics while leaving the actual restart to the operator (it never restarts itself into a possibly half-applied checkout). The shared update marker — attacker-adjacent state any process with write access to the Agent home can create — is read defensively: the read never follows a symlink, never blocks on a FIFO/device, and never reads an unbounded file, classifying anything that is not a small regular file as unknown (fail-closed), with a portable fallback on platforms lacking the atomic open flags. Thanks @hrdwdmrbl. (#7160)
- **Sidebar session polling no longer retains full transcripts and run-journal detail in the response cache.** Cached rows are now projected to the canonical sidebar field allowlist before storage, heavy metadata is stripped across index and full-scan paths, nested request values are isolated, and invalidated rebuilds are rejected atomically instead of being reinserted after a concurrent clear. Thanks @martindell. (#7489)
- **The workspace switcher now lists workspaces in your configured order instead of re-sorting them alphabetically.** The chat-header workspace dropdown alphabetized its entries client-side, so it disagreed with the drag-and-drop order shown in the Workspaces settings panel (and never kept the default "Home" workspace first). The dropdown now renders the server's stored order verbatim — the same order as the settings view — with no client-side re-sort. Thanks @CharlesMcquade. (#7317)
- **Auxiliary-model settings now persist provider-native model IDs instead of silently dropping the provider prefix.** When an auxiliary model (title/summary/etc.) was chosen as a provider-qualified ID, settings canonicalization stripped the `@provider:` prefix and could persist a bare name that resolved under the wrong provider group. The picker now preserves the provider-native ID, dedupes qualified/unqualified variants, and fails closed on an ambiguous mismatch rather than persisting a wrong upstream model. Thanks @lowlandsheperd. (#7261)
- **Hardened a streaming race that could read a stream queue mid-teardown.** Several code paths read the in-memory SSE stream registry with a bare, unlocked `STREAMS.get()`, so a read that raced a concurrent teardown (cancel/finish popping the queue under `STREAMS_LOCK`) could observe-and-use a queue the registry had already released. Registry reads now go through a lock-disciplined `peek_stream()` helper that takes `STREAMS_LOCK`, and a self-enforcing test forbids reintroducing a bare read. Callers keep their existing empty/None-guard fallbacks, so behavior on the common non-race path is unchanged. Thanks @zicochaos. (#7339)
- **Automatic session titles no longer keep an internal placeholder for workspace-prefixed multimodal first turns, and internal scaffolding no longer leaks into generated titles.** Title eligibility compared persisted content that still carried WebUI workspace/attachment scaffolding against a sanitized visible title, so a workspace-prefixed structured multimodal first turn could never be recognized as provisional and stayed stuck on its placeholder title; the same internal scaffolding could also reach the title provider's input. Title generation now sanitizes title-facing structured text (stripping versioned workspace sentinels and generated attachment suffixes) across eligibility, first-exchange generation, manual prefer-latest generation, and adaptive refresh, while preserving genuine user content and native image parts. Persisted and model-facing message content is unchanged (titles only read it), and a real user-chosen title is never misclassified as provisional. Thanks @starship-s. (#7306)
- **The composer footer no longer jitters the transcript while a reply streams.** During SSE streaming, the composer footer's fit-stage probe transiently applied and reverted layout classes, which nudged the transcript viewport a few pixels backward on each frame (a visible clamp/jump on the crown-jewel chat surface). The fit probe now runs synchronously with the transcript geometry frozen and the caller's exact styles restored in a `finally`, so footer sizing is measured without moving the reader's scroll position; adaptive stage selection is unchanged. Thanks @ruizanthony. (#7275)
- **A Windows self-restart now relaunches the actual server instead of replaying whatever process launched it.** The Windows restart path replayed the launching process's `sys.argv`, so a server started indirectly (e.g. under pytest) would restart *that* process rather than the server. It now resolves the canonical replacement command per packaging mode — a source or pip install relaunches `server.py` directly (preferring windowless `pythonw.exe` when present), while frozen/packaged argv is left unchanged — and preserves host, port, state, workspace, and agent environment across the restart. The POSIX (exec) restart path is unchanged. Thanks @rodboev. (#7270, closes #7195)
- **A malformed or non-object JSON request body now returns a clear `400 Invalid JSON body` instead of a misleading "missing field" error.** `read_body()` used to silently turn any JSON parse failure into `{}`, so a client that sent broken JSON got a confusing `Missing required field` 400, and a non-object body (a JSON array or scalar) passed straight through. It now rejects malformed JSON and non-object bodies with a clean 400 at the PATCH/DELETE/PUT/POST/TTS entry points (oversized bodies still return 413). Empty and whitespace-only bodies are still treated as `{}`, so body-optional endpoints (e.g. `DELETE /api/mcp/servers/{name}`) are unaffected. Thanks @zicochaos. (#7336)
- **A pre-`source`-column `state.db` no longer floods `errors.log` with the same warning every ~15 seconds.** When an older profile's `state.db` lacks the `source` column, the session-listing path logs a "no 'source' column" warning and returns an empty list — but the sidebar poll (behind a 5s cache) re-ran it every ~15s, so a single obsolete-schema profile DB produced a duplicate warning line for the life of the process (an operator saw 276 identical lines in ~70 min). The warning is now emitted once per resolved `state.db` path per process lifetime; the degraded-path behavior is otherwise unchanged, and a restart warns again. Thanks @rodrigogs. (#7476, closes #634)
- **Automatic session-title generation now publishes the conversation context, so relay backends route the title call to the right target.** The aux-client title generator (`generate_title_raw_via_aux` → `call_llm`) ran with no Agent conversation context published, so relay targets that derive routing from it (e.g. OpenCode Go) mis-routed the title request (#7470). The title call now republishes the WebUI session id as the Agent's ambient conversation context for the duration of the aux call and resets it on every exit path (success, early return, and exception), using the real runtime's `ContextVar` so concurrent title calls stay isolated. An older or absent Agent runtime without the context setter keeps the previous behavior rather than failing title generation. Thanks @webtecnica. (#7474, closes #7470)
- **Live progress text in Transparent Stream no longer tears in the middle of a word.** During a streaming turn, a live progress row could split a word across a block boundary and render the tail on its own line (e.g. `Ces deu` / `x fichiers passent.`). The streamed text was always intact — it was purely a client-side render artifact in the live fade/reconcile path. The renderer now adopts the parser-owned DOM once after a parser-owner replacement (keeping the per-token fast path for normal ticks) and handles the source-space cursor at block boundaries, so live prose stays whole with no lost, duplicated, or reordered text. Verified with the crown-jewel stream gate (no flicker, no alternating, no prose reversals across transparent/worklog/hidden modes). Thanks @ruizanthony. (#7082)
- **A slash-style model ID picked under a session provider that differs from the profile default now routes to the right provider instead of 404ing.** When you selected a model whose ID contains a `/` (e.g. a Nous portal row `upstage/solar-pro4:free`) while the session's provider differed from the profile default, `model_with_provider_context()` dropped the provider hint, so the request went to the *default* provider's base URL (e.g. `api.x.ai` under an xAI-OAuth default) and 404'd. The resolver now emits an explicit `@provider:model` hint when the session provider is a known static/portal provider or a `custom:<slug>` that resolves to a unique `custom_providers[]` entry, while still keeping the bare ID for unknown/ambiguous slugs so existing custom/proxy base-URL routing stays in charge. Ambiguous custom slugs fail closed to a persisted error rather than crashing the worker. Thanks @Manny7717. (#7356, closes #7333)
- **Editing and resubmitting a message can no longer truncate most of a long session's history if you click "Send edit" more than once.** `submitEdit` was re-entrant: its only guard was `S.busy`, which the send path doesn't set until after two multi-second awaits (`_ensureAllMessagesLoaded` on a long session, then the truncate round-trip). On a laggy instance that left a 10s+ window in which a second click started another full truncate — and because the first call sets `_oldestIdx = 0`, the second call recomputed `keep_count` from a still-window-relative index, producing a far smaller keep count that could delete most of a 2000-message session. A `_submitEditInFlight` guard is now acquired synchronously before the awaits and released in a `finally` (so an early return or throw can't wedge editing off), while the absolute keep count is still captured up front so a single legitimate submit is unaffected. The Enter-to-submit path clicks the same button, so it's covered too. Thanks @alanjds. (#7277)
- **Read-only child sessions stay grouped after their writable siblings in the sidebar lineage view.** When a session lineage is collapsed, read-only child sessions (e.g. shared/observed sessions) could interleave with writable children instead of sorting after them, making the group order unstable. The comparator now orders writable children first and read-only children after, within each lineage, with a stable total order that tolerates missing order keys and both the `read_only` and `is_read_only` shapes. Session open/selection is unaffected (still keyed by `session_id`, independent of row position). Thanks @igorroncevic. (#5888)
- **Inline event-handler arguments across the UI are now escaped for the JavaScript context, closing an attribute-injection class.** `esc()` only escapes for HTML, but the browser decodes attribute entities *before* it parses an inline handler's JavaScript, so a value interpolated as `onclick="fn('${esc(id)}')"` let a quote in the data break out of the JS string literal (the #3797 bug class). A shared `jsArg()` helper (`esc(JSON.stringify(String(v)))` — `JSON.stringify` supplies its own quotes, `esc()` then makes the result attribute-safe) now wraps **every** inline-handler string argument across cron/kanban/checkpoint/gateway/passkey/notes actions, session handoff hints, and onboarding device-code copy. The one-off `_kanbanJsArg` that previously fixed this for kanban buttons alone is removed in a clean cutover. Thanks @zicochaos. (#7335, closes #3797)
- **Nested lists that mix bullets and numbers render as a real hierarchy, and a very deeply nested list no longer blanks the conversation.** The renderer used to draw one marker family and then re-parse its own generated HTML with the other, so a list mixing `-` and `1.` lost its nesting. It now builds the nested `<ul>`/`<ol>` tree in a single pass. That tree is serialized with an explicit work stack rather than recursion — the first version of this fix recursed once per nesting level and threw a `RangeError` on a pathologically deep list, and because rendering happens after the transcript is cleared the uncaught error blanked the whole session. Output is byte-identical to the previous renderer across 50,000 randomized trees. Thanks @webtecnica. (#6716, closes #6700)
- **When a skill isn't found, the error reply no longer implies the skill list is complete when it was truncated.** `/api/skills/content` capped the `available_skills` courtesy list at 20 names but presented it as the whole set, so a caller (agent or human) that couldn't find its skill there would wrongly conclude it isn't installed. The reply now includes `total_skills` and `available_skills_truncated`, and the hint reads "Showing 20 of N skills…" when the list is cut, so a missing skill is never mistaken for an absent one. Thanks @elight. (#7428, closes #7426)
- **Session title generation and the handoff summary work again on the Anthropic and Codex out-of-band paths.** A Hermes Agent refactor unified provider dispatch and removed the two response normalizers these paths called (`agent._normalize_codex_response` and `agent.anthropic_adapter.normalize_anthropic_response`), so on a current Agent both paths raised as soon as they ran — background title generation and the handoff summary silently produced nothing on those providers. Both call sites now go through the Agent's current transport API (`_get_transport(...).normalize_response(...)`), matching the Agent's own canonical helper including the Anthropic OAuth tool-prefix strip. Empty-response and incomplete-summary handling are unchanged. Thanks @zicochaos. (#7402)
- **A long final assistant message no longer stalls every other active stream for seconds.** The echo-suffix folding that removes a duplicated tail from streamed text re-normalized the entire remaining window for *every* candidate cut position — quadratic in the window size, so a ~6,000-character conclusion burned seconds of CPU while holding the GIL, freezing every other stream in the process. The window is now folded once and the cut point located by walking backwards across the echo, which is linear and produces byte-identical output on every input (verified by ~380,000 differential comparisons across Unicode whitespace, window sizes, straddling and repeated echoes, with zero divergences). Thanks @ruizanthony. (#7288)
- **Extensions that ship a `.gitkeep` (or `.gitignore`/`.gitattributes`/`.env.example`) install again instead of failing with "Unsafe archive member".** The archive validator rejected *any* dot-prefixed path segment, which also blocked the benign placeholder files extensions legitimately carry — so an extension with an empty `assets/` directory couldn't be installed at all. Archive install/uninstall now use a dedicated validator that permits a short allowlist of hidden files **as the final path segment only**; dot-directories (`.git/`, `.ssh/`) and every other dotfile (`.env`, secrets) stay rejected, as do traversal, absolute paths, encoded separators, NUL/backslash, and an allowlisted name used as a directory. The shared strict validator is unchanged and still gates static serving, asset URLs, and manifest paths, so an installed hidden file lands on disk but can never be fetched over HTTP or injected into a page (`.env.example` is installable but not servable; `.env` remains rejected outright). Thanks @asorourx. (#6620, closes #6619)
- **Passkey sign-in works behind a reverse proxy that rewrites `Host`.** WebAuthn scopes a credential to an "RPID" that must match the origin the browser is actually on, but the RPID was derived from the `Host` header — which a proxy fronting a separate SPA rewrites to its own upstream address, so registration and login failed with an RP ID mismatch. The RPID now comes from the request `Origin` (accepted only as a well-formed `http(s)` origin with a hostname, rebuilt from its parsed parts, otherwise falling back to `Host`). This is not a trust decision: the authenticator still only releases a credential whose RPID is a registrable suffix of the real page origin, `clientDataJSON.origin` must equal the origin stored with the challenge, the authenticator's `rpIdHash` must match that stored RPID, and the assertion signature is still verified against the stored credential public key. Thanks @cabluvsmkm2009-source. (#6772)
- **Trusted-header auth can be told to accept pipe-separated groups from the identity proxy.** An Authentik outpost (depending on its proxy provider / property mapping) can emit a group header like `admins|developpeur`; the parser only split on commas and newlines, so the whole string was read as one group name, matched nothing in the configured group→profile map, and the session silently fell back to the unbound `default` profile despite a legitimate mapped membership. Setting `HERMES_WEBUI_TRUSTED_GROUPS_PIPE_SEPARATOR=1` now also treats `|` as a separator (comma and newline stay the default). It is opt-in on purpose: a group *name* can legitimately contain a literal `|`, so splitting on it unconditionally could silently change an existing deployment's profile binding. Repeated or adjacent separators collapse instead of producing empty group names. Thanks @SamsGuamejy. (#7331)
- **Jumping to the start of a long conversation while it is still streaming no longer opens a blank transcript, and an idle long session no longer re-renders itself several times a second.** With transcript virtualization on, the scroll listener rescheduled a virtualized render on every scroll event — including its own programmatic scrolls — so a settled long session sat in a ~5-6 renders/second loop that pinned a CPU core and broke text selection. The listener now skips rescheduling while a programmatic scroll is in flight, and the render window is restored after a settled render. "Jump to session start" during an active stream deliberately skips the full re-render (it would drop live Activity cards), so it now mounts the render window explicitly — without that, the jump landed at the top of an all-spacer transcript with no messages rendered. Thanks @LeonardoLGDS. (#7427, closes #6799)
- **The last character before a tool call is no longer dropped from streaming assistant text.** When a model emitted a `<function_calls>` block immediately after prose, the strippers that remove the XML block also trimmed the text, silently eating the final character of the sentence before it. Only leading whitespace is trimmed now, so the prose survives intact. Thanks @bsgdigital. (#7207)
- **A brand-new chat no longer shows a phantom "Compressing context" divider on its very first message.** The server appends a placeholder lifecycle row whenever a stream has events but nothing visible to project yet, and the worklog classified any such row reporting a bare `running` phase as the start of a context compression — so a session that never compressed anything grew a permanent, reload-surviving compression divider. Classification now requires an explicit compressing phase or matching text. Thanks @MuhammadUsamaMX. (#7447)
- **Searching the model picker matches the model names you actually see, and a punctuation-only search no longer returns every model.** OpenRouter curated entries were listed by raw id, so typing the display name (e.g. "Ox Alpha") found nothing; picker rows now carry the friendly name from the local OpenRouter metadata cache (read from disk only — never a network call — falling back to the raw id when unknown). Search also folds spaces, dots, hyphens and underscores, so "ox alpha", "ox-alpha" and "ox.alpha" all match — while a search consisting only of separators correctly matches nothing instead of everything. Thanks @webtecnica. (#7385, closes #7228)
- **The Insights period selector shows its dropdown chevron again.** An inline `background:` shorthand on the control reset the shared `select` background image, hiding the chevron and removing the padding reserved for it, so the filter looked like static text. The inline style moved into a class so the shared select styling stays authoritative, and the control keeps its compact width. Thanks @tomatotomata. (#6768, closes #6725)
- **Docker: secrets ending in `=` (like a base64 shared secret or password) are no longer truncated when the container reloads its environment after dropping privileges, so shared-secret auth and `HERMES_WEBUI_PASSWORD` matching keep working.** `docker_init.bash`'s `load_env()` reloaded the saved env snapshot with `while IFS='=' read -r key value`, and Bash silently drops a single trailing `=` from the last field — so any value ending in one `=` (e.g. `openssl rand -base64 32` output, which is 44 chars ending in `=`) came back one character short. The reload now reads each line whole and splits at the first `=` only, preserving embedded and trailing `=`; blank/comment lines are skipped and lines without `=` keep their prior empty-value behavior. Thanks @webtecnica. (#7433, closes #7432)
- **WebUI session titles are written to `state.db` again on the latest Hermes Agent, so `hermes sessions list` isn't blank for WebUI sessions.** The agent's `SessionDB.set_auto_title_if_empty` was renamed to `set_auto_title(session_id, title, *, source)` in a state-module refactor (released in agent v2026.9.7); the WebUI's background title sync still called the old name, whose `AttributeError` was silently swallowed — so auto-generated titles stopped persisting. The sync now calls `set_auto_title` with LLM provenance, preserving the exact prior behavior (only fills a NULL/lower-authority title, never clobbers a manual rename, de-duplicates colliding titles with a `#N` lineage suffix). Also repoints the gateway-approval test suite's `tools.approval._ApprovalEntry` reference to its post-split location (`tools.approval_gateway_wait`) via a test-only compatibility shim, so the agent-coupled tests run green across both the pre- and post-split agent.
- **Regenerating a response no longer refetches the entire conversation, so regenerate on a long chat is fast instead of stalling (and silently cancelling).** Clicking "regenerate" used to pull the whole transcript back from the server before it could rebuild the request, which on a large session caused a long UI freeze and could time out and silently drop the regeneration. The regeneration authority now reads only a bounded, sidecar-anchored tail of recent turns instead of the full transcript — and it is exact by construction: it takes the tail's prefix proof, keys, and rows from a single WAL-consistent database snapshot (a `data_version` check forces a full read if the database is written mid-snapshot), reuses the canonical projection and durable ordering so the bounded rows match the full reader byte-for-byte, and falls back to a full read whenever the skipped prefix isn't provably identical or the tail contains a duplicated turn key. If any of those invariants can't be met it reads the whole transcript exactly as before, so a concurrent edit or an unusual session can never produce a stale or reordered regeneration. Thanks @webtecnica. (#7204, closes #6826)
- **Loading the session sidebar is faster on installs with many delegated-subagent sessions — the subagent classification no longer runs a separate database probe per row.** Building `/api/sessions` used to open a `state.db` connection for every session that might be a delegated child, an N+1 pattern that grew with session count. The classification now reads the authoritative `state.db` source for all candidate rows through the existing batched overlay (one read-only connection, chunked 500-ID `IN` queries), and a rich-read failure recovers the remaining IDs in a single bounded source-only pass rather than falling back to per-row queries. Behavior is unchanged — a delegated child whose index row still says `webui`/`fork` while its `state.db` source is `subagent` is classified read-only exactly as before. Thanks @starship-s. (#7231)
- **Reopening a conversation that contains pasted or generated images no longer shows the same turn twice.** A native-image user turn is stored two ways — the rich WebUI sidecar row (with the actual image) and the Hermes Agent state row (the same turn as model-facing text, with each image replaced by `[screenshot]`) — and transcript reconciliation treated them as two different messages, so reopening the chat rendered a duplicate `[screenshot]` bubble under the real one. Reconciliation now recognizes the scalar state row as a mirror of the rich image row (matching on the exact `[screenshot]` projection, full-precision timestamp, role/tool shape, and provenance) and keeps only the rich row, so the image and its metadata survive and the duplicate is gone. It fails closed: a literal `[screenshot]` typed by the user, an ambiguous or conflicting identity match, a timestamp mismatch, or malformed content all preserve both rows rather than risk collapsing two genuinely different turns. Thanks @starship-s. (#7230)
- **Reopening a long conversation is now much faster — the transcript is served from a bounded cache instead of re-loading the whole session database each time, without ever showing stale content.** Reopening a large chat used to pay the full `state.db` load + merge cost even when nothing had changed; the WebUI now derives a bounded, indexed signature of the session's database (DB/WAL/SHM) plus its sidecar/lineage provenance, and reuses a cached merge when that signature is unchanged. The cache is fail-closed by construction: it is bypassed for active or pending-turn sessions, older-message paging (`msg_before`) always reads the full uncapped transcript, a state edit changes the per-target signature (so an edit to one session is never masked by an unrelated active stream), incomplete or unverifiable lineage provenance is evicted rather than reused, and any uncertainty falls back to the full load. Thanks @ruizanthony. (#7212)
- **The three-container Docker stack can now expose the agent gateway and dashboard for remote access without editing the compose file — and stays loopback-only by default.** The agent gateway (`8642`) and dashboard (`9119`) host-port publishes are now `${HERMES_AGENT_BIND:-127.0.0.1}` / `${HERMES_DASHBOARD_BIND:-127.0.0.1}`: unset, they bind to `127.0.0.1` exactly as before (the dashboard runs `--insecure`, so a non-loopback default would publish it on every interface); set `HERMES_AGENT_BIND=0.0.0.0` (and put the dashboard behind auth / a reverse proxy) in `.env` to reach them from another machine. The compose comments also document the real cause of a "Hermes agent is not responding" WebUI error — the gateway listener only opens when `API_SERVER_KEY` is set. Thanks @mercael91. (#7145, closes #6876)
- **Local model IDs that carry a colon tag (e.g. Ollama `qwen3.8:27b-mtp-q8_0`) are no longer truncated when selected.** The WebUI wraps a chosen model in its internal `@provider:model` form, and a naive colon split dropped everything after the model's own colon, so the tag/quantization suffix was lost. Model identity now resolves through the shared `@provider:model` parser, which preserves built-in, plugin, custom, `host:port`, tagged, and unqualified forms (an explicit `model_provider` still wins), keeping WebUI in parity with the models you actually run in the Hermes CLI. Thanks @darthzen. (#7183, closes #7182)
- **Session exports now work when the WebUI is reverse-proxied under a subpath.** The JSON/HTML/Markdown export links were built as root-relative URLs (`/api/session/export?…`), so a deployment mounted under `https://host/<prefix>/` resolved them at the host root and bypassed the deployment prefix (export 404s / wrong target). Export URLs now resolve in the browser against `document.baseURI`, so both root-mount and subpath-mount deployments build the correct link. Server API unchanged. Thanks @lowlandsheperd. (#7189)
- **Cancelled or interrupted compression no longer loses or misattributes transcript work.** WebUI now fences the transcript with a durable revision so a stale or partially-written snapshot can't overwrite the authoritative current turn: current-turn authority no longer reuses shifted historical indices, stale-snapshot outcomes settle terminally (no automatic retry loop), and partial output stays associated with the turn that produced it. On Hermes Agent builds without the new revision parameter the WebUI omits it and follows the unchanged legacy path (the fence activates once the companion Agent build ships), so existing installs are unaffected. Thanks @ruizanthony. (#6554)
- **The dashboard "loopback-only" warning no longer fires when you've configured a public dashboard URL.** The warning was keyed only off the browser's own origin, so a reverse-proxied deployment with a configured public `status.browser_url` still saw the false "only reachable on this machine" alert. It now shows only when the WebUI page origin is non-loopback **and** the resolved dashboard target is actually loopback (a shared classifier covers `127.0.0.0/8`, `::1`, IPv4-mapped-IPv6 loopback, and `localhost`/`*.localhost`), so a genuine loopback-only bind is still flagged. Thanks @webtecnica. (#6573)
- **Opening a new chat from a conversation now starts in that conversation's workspace, not the profile-global default.** With several conversations running in parallel on different workspaces, "New Chat" used to jump to the profile-wide default workspace instead of inheriting the workspace of the conversation you started it from. It now inherits the current conversation's workspace (an explicit profile-switch workspace still wins first, and a blank page still falls back to the profile default). If that inherited workspace directory was since deleted, the server recovers to a valid fallback instead of failing to open the chat — and only that verified same-profile inheritance is recovered, never a foreign or explicitly-typed path. Thanks @ruizanthony. (#7180)
- **Tool-heavy turns no longer rewrite the entire session file on every tool call.** The periodic checkpoint used to re-serialize the whole session sidecar each time a tool call completed, even when nothing it persists had actually changed — so a turn with many tool calls wrote the same bytes to disk over and over. The checkpoint now skips the rewrite when the state it would persist is unchanged (with a periodic refresh and a fail-safe that always writes if anything is uncertain), cutting redundant disk I/O on busy turns without weakening crash/reload durability. Thanks @ruizanthony. (#7172)
- **Clarification prompts now honor your configured `agent.clarify_timeout`, including `0` for "wait indefinitely."** The WebUI clarify bridge kept its own divergent timeout (it read only the legacy `clarify.timeout` key, fell back to 120s, and treated `0` as invalid), so the agent's "unlimited" setting was ignored and prompts could expire out from under you. It now reuses the agent's own timeout resolver, so the two stay in lockstep: `0` (or any non-positive value) means no expiry — the card waits until you answer or the run is cancelled, with no countdown — and positive values still expire and clear as before. Thanks @BBaFSE. (#7163)
- **Custom providers that share a model name are now routed by the provider you actually selected, so requests can't reach the wrong endpoint or use the wrong API key.** When two custom (OpenAI-compatible) providers exposed the same model id, resolution could pair one provider's endpoint with another's credential; provider slugs are now derived through a single canonical identity so overlaps are detected and each request resolves its endpoint and key from the same selected provider. Ambiguous configurations now fail closed with an actionable error instead of silently mis-routing. Thanks @rouVling. (#7107)
- **Dragging several files or folders into the workspace at once no longer silently drops all but the first.** The OS-drop handler read the browser-owned drop entries lazily during async traversal, but the browser only keeps those entries valid during the synchronous drop event — so every root after the first became unreadable and was skipped. The handler now snapshots all dropped files and folder entries up front, then traverses the snapshot, so multi-root drops upload completely. Thanks @nightcityblade. (#7142)
- **A message you sent at the moment a reply finished no longer shows up twice (out of order) after a reload.** The Hermes Agent stamps persisted messages with a `_row_id`, but the WebUI replay-merge only recognized its other row-id aliases, so the same stored row looked new and was appended again after the settled answer. Replay now recognizes the `_row_id` alias at the shared identity chokepoint, so the row is matched to its existing turn and kept in order. Thanks @ruizanthony. (#7133)
- **Remote (SSH/Docker) workspaces no longer fail with "Path does not exist" when the WebUI host is macOS.** A macOS host resolving a target-side Linux workspace path (e.g. `terminal.cwd = /home/<user>`) ran it through the local filesystem, where macOS rewrites the synthetic `/home` firmlink to `/System/Volumes/Data/home` and then rejects the non-existent local path. Remote terminal profiles now preserve the target-side POSIX path verbatim instead of resolving it against the host filesystem — after the same containment checks (POSIX normalization rejects `..` traversal, NUL bytes, and blocked system roots, and the path must stay within `terminal.cwd`). Local workspaces still resolve and are contained exactly as before, and the fix is host-agnostic (helps a Linux host with a remote profile too). Thanks @alfred-rootson. (#7132)
- **Regenerating a response is now atomic — a concurrent message sent at the same moment can no longer be silently erased, and regeneration is no longer wrongly blocked on completed imported turns.** Starting a regeneration snapshots and validates session state under the session lock before mutating it, so a message that another tab/device sends during the same window survives (the regeneration cleanly returns a conflict instead of rolling back the winner), and completed WebUI-owned turns imported from a CLI/TUI/desktop session can be regenerated while malformed, read-only, or foreign final turns stay correctly rejected. The active-turn marker is now carried through the public projection so a mid-stream reload renders the pending prompt once instead of twice. Thanks @rodboev. (#6677)
- **Large sessions no longer replay a stale duplicate prompt (or drop a real turn) when reconciling the sidecar against the model transcript.** Matching one logical turn across the two stores relied on an O(n²) fuzzy fallback that was bounded to the first 1,000 keys for performance — but above that bound a workspace-prefixed prompt (state.db stores the model-facing `[Workspace::v1: …]` wrapper while the WebUI sidecar owns the bare text) stopped deduping, so a 1,001-key session replayed the prefixed prompt after its answer. The workspace prefix is now folded into the exact key so protocol-equivalent turns match in O(1) above the cap, and — critically — that normalization is source-aware: only the model-facing state.db key strips the wrapper, while the sidecar-owned visible key preserves literal text verbatim, so a user who genuinely types `[Workspace: …]` isn't collapsed into a different turn. Thanks @Dandandad. (#7128)
- **The Skills panel no longer shows every skill as enabled (and silently drop your other disabled skills on a toggle) when `skills.disabled` is stored as a JSON-array string.** `hermes config set skills.disabled …` and JSON-mode config saves persist the list as a quoted string (`'["skill-a","skill-b"]'` or the Python-literal `"['a']"`), which the panel read as a single literal skill name — so no skill matched, all rendered enabled, and toggling one rewrote `skills.disabled` to a one-entry list, discarding the rest. The panel read and the toggle write now decode the JSON-array/Python-literal string form into a real list (reusing the agent's shared parser when present so the two surfaces cannot drift, with a safe `ast.literal_eval` fallback), while a genuine scalar skill name still means one entry. Thanks @webtecnica. (#7120)
- **Loading several large sessions at once can no longer spike memory into an out-of-memory crash.** Full sidecar resolution (the expensive JSON parse for a large transcript) is now bounded by a small semaphore, and concurrent requests for the same session elect a single leader that parses once while the others wait for its cached result instead of each parsing their own copy — so a burst of tabs or reconnects against big sessions can't multiply transient memory. Metadata-only reads and cache hits stay on the lock-free fast path, and a failed parse always releases its slot (no capacity leak). Thanks @Dandandad. (#7127)
- **Reconnecting to a live chat stream now resumes from the client's last received event instead of losing the tail or replaying the whole turn.** The `/api/chat/stream` endpoint honors the SSE `Last-Event-ID` header (and `after_event_id` / `after_seq` / `replay=1`) with a resume cursor that carries provenance, so a dropped connection re-subscribes at the right place: an in-range cursor replays only the gap, an unparseable/foreign/ahead cursor fails closed to replay-from-start (never silently skipping frames), and a genuinely truncated buffer still emits `recovery_control`. A prior edge where an ahead cursor with no known snapshot cutoff installed itself as the live dedup bound — discarding every queued frame including the terminal `stream_end` and stalling the reconnect on heartbeats — is fixed by treating an unknown cutoff as fence 0 before normalization. Thanks @Paladin173. (#6886)
- **A parked no-run approval can no longer mask the real pending approval when a gateway producer is live.** No-run "legacy" approval mirrors are now stamped with a token derived from their own live producer, and whenever a session's gateway queue has live producers the reconciler tokenizes from the authoritative head and discards any unmatched or tokenless mirror instead of letting it linger — so a stale card can't sit in front of the genuine pending approval (previously gateway mode returned 409 and permanently retained the stale mirror, and local mode reported success while resolving nothing). Tokenless-orphan retention is preserved for the no-producer case, and responding to a stale id fails without waking a different producer. Thanks @VHSgunzo. (#7086)
- **Toggling YOLO from an approval card no longer hangs the gateway run or leaks across a session switch.** The approval card's YOLO branch now takes the same single `gateway_yolo_handoff(sid)` as the Runs dispatcher (revalidating the exact mirror after acquisition and holding it through settlement, with no nested-lock path), and the client captures one immutable owner tuple (card-owner session, approval id, load generation) at click time — failing closed without a POST when the visible card is not the active session, and fencing every post-response mutation (card hide, pending clears, stale-clear requery, controls, toast, status) on that exact tuple so a stale card can't project its result into the successor session. Thanks @dss539. (#6945)
- **Git operations no longer flash a console window on Windows.** Every git subprocess boundary (agent runtime, rollback, updates, workspace, worktrees) now routes through a shared `windows_hide_flags()` helper that passes `CREATE_NO_WINDOW` on Win32 and stays a no-op on POSIX, so short-lived git commands no longer pop a visible console window on Windows while keeping captured stdout/stderr intact. Thanks @ya1yi2fa3. (#6534)
- **A `/goal` launch no longer silently switches providers mid-session after a deliberate model pick.** When a session had a deliberate cross-provider model/provider choice, running `/goal` omitted the `explicit_model_pick` marker that `/api/chat/start` sends, so the model resolver treated the persisted pick as stale and "repaired" it back to the profile default — silently changing providers while the goal ran. Both the frontend `/goal` command and the backend `_handle_goal_command` (initial kickoff and kickoff-prompt fallback paths) now carry the explicit-pick marker and stamp the pick signature, matching `/api/chat/start`; ordinary default sessions are unaffected (they still omit the marker, so a default pick stays repairable). Thanks @webtecnica. (#6705)
- **A completed assistant turn no longer renders twice in the feed.** In `renderMessages()`, the live-turn preserve guard kept the in-flight assistant node across the DOM rebuild whenever it held any rendered content, so when a stream ended but its inflight bookkeeping was not yet cleared the just-completed turn was re-attached over the rebuilt settled transcript (a first copy without the model label, then a second with it). The stored data was always clean (state.db, sidecar, and `/api/session` each hold one row) — the duplicate existed only in the DOM. Preservation now requires a provable live owner (`S.activeStreamId` or explicit live-projection markers) instead of bare DOM content; the mid-stream flicker guard and blank-turn guard are preserved. Thanks @webtecnica. (#7109)
- **Visible inline chat videos no longer appear stalled after metadata.** Every inline video stayed at `preload="metadata"`, so a currently-visible mobile player could look frozen because the browser stops loading after metadata. A shared IntersectionObserver now promotes a near-viewport video to `preload="auto"` (and immediately on play), keeps off-screen history videos at metadata to avoid buffering every historical video, and unobserves nodes removed before they intersect. No media URL, caching, or controls change. Thanks @allenliang2022. (#7076)
- **Profile goal contract no longer drifts across profiles on older agent builds.** On Hermes Agent builds where the context-override API was exposed but `SessionDB()` still used the import-time default DB path (v2026.5.28–v2026.7.20), a goal mutation in one profile could overwrite another profile's goal. Native profile-goal persistence is now gated on the agent supporting *both* context-home resolution *and* call-time DB-path selection; otherwise it falls back to the legacy per-profile manager. Profiles stay isolated. Thanks @ticketclosed-won. (#6899)
- **Packaged `hermes-webui` console entry point.** `pip install`-style wheel installs now expose a `hermes-webui` command that boots the WebUI through `bootstrap:main` — the same launcher path as `python server.py`, so agent-interpreter discovery and dependency validation run before the server starts (a direct `server:main` entry started the UI but broke every chat with a missing-dependency error). Thanks @webtecnica. (#6742)
- **Dotted Bedrock/Vertex model IDs get clean labels in the model picker.** Model IDs with dotted vendor prefixes (e.g. `anthropic.claude-*`) rendered awkwardly in the picker. They now get normalized display labels while the exact model IDs and provider grouping stay authoritative — distinct regional/vendor models remain separate entries, and live discovery still precedes the static fallback. Thanks @samfoy. (#6628)
- **MCP dev dependency pinned to a compatible range.** The MCP SDK dev dependency was unpinned, so an incompatible 2.x install broke `mcp_server.py` with an `AttributeError`. It is now pinned to a compatible 1.x range with a bootstrap guard that exits with an actionable message on an incompatible version; the SDK remains optional (imports degrade gracefully when it is absent). Thanks @webtecnica. (#6616)
- **Cancelling runs are no longer reattached on recovery.** A background run in the middle of cancellation could be resurrected by a reconnect/recovery reattach, producing a dying run that could double-deliver or churn. Recovery paths now exclude runs whose phase is `cancelling` (they stay busy but are not reattached), while genuinely-live runs still reattach; stale-cancelling reclamation is bounded (age ≥ 180s and absent from active streams) and idempotent. Thanks @allenliang2022. (#7096)
- **GPT-5.6 models expose their full reasoning range.** The four GPT-5.6 model IDs support the `max` reasoning-effort level, but the capability filter capped them below that. The ceiling is now lifted for exactly those IDs on the OpenAI-family lanes (older GPT-5 and o-series ceilings unchanged), and `max` is offered only when the authoritative configured/live catalog includes it. Thanks @boudywho. (#7083)
- **Test suite: two model-catalog tests are now order-independent.** `test_mtime_invalidation` depended on sibling-test cache state and `test_glm_5_3_in_models_payload_for_zai_provider` depended on the installed agent-core catalog version, so both failed spuriously under isolation/sharding. They now set up their own preconditions. Thanks @webtecnica. (#7101)
- **Binary files in the workspace no longer truncate when served on Windows.** The shared anchored file-open helper omitted the platform's `O_BINARY` flag, so on Windows a binary workspace file (video, image, archive) was read in text mode and truncated at its first `0x1A` (Ctrl-Z) byte — a 5.2 MB video could arrive as a few hundred bytes, breaking Range requests and playback. Non-directory anchored opens now set `O_BINARY` where the platform provides it (a no-op on Linux/macOS). Thanks @allenliang2022. (#7077)
- **Unstyled native form controls now use the UI font.** Bare `button`, `input`, `select`, and `textarea` elements fell back to the browser default font instead of the app's UI font, and the markdown table filter lost the UI font after its `font: inherit` shorthand. Both now route through `--font-ui`; mono, conversation, and skin-scoped controls are unchanged. Thanks @starship-s. (#6871)
- **Mobile navigation action shortcuts keep their rail order.** In the mobile sidebar, mirrored navigation actions could land out of order relative to the desktop rail after a native-tab or profile reorder. They now reconcile into rail-source order with no needless DOM moves. Thanks @starship-s. (#6909)
- **The Workspace Artifacts panel no longer keeps a stale entry after a session recovers, and no longer flashes another session's artifacts into the panel you're reading.** After a settled-session recovery the Artifacts panel could keep showing a moved or deleted artifact as a dead link until you manually switched sessions, and a background session finishing could momentarily repaint the foreground session's panel. Artifact re-projection is now scoped to the current session and pane (a background session's completion no longer touches the panel you're viewing), and the current session's panel refreshes correctly after recovery. Thanks @rodboev. (#6867)
- **Kanban task and board dialogs stay usable in short windows.** On a short window (short-wide landscape, small laptop, or a phone in landscape) the Kanban create/edit-task and create-board dialogs were taller than the viewport with no height cap and no internal scroll, so the dialog overflowed both the top and bottom edges and its action buttons (Create / Save / Cancel) were unreachable. The dialogs now cap their height to the viewport (`calc(100dvh - 48px)`, with a `100vh` fallback) and scroll their content internally, and the overlay uses safe centering so the top stays reachable when the dialog is taller than the window. Normal-height dialogs are unchanged. Thanks @rodboev. (#6906)
- **A stale approval card no longer dead-ends with "Approval response not accepted."** On the local backend, clicking a dangerous-command approval card whose approval had already been resolved or cleared could leave the card stuck showing an error, because WebUI failed to match the resolved approval back to its producer (the agent core delivers the approval as a copy that carries a `request_id` but no `approval_id`, so the existing identity/`approval_id` matches both missed and a tokenless mirror was orphaned). WebUI now also matches on the core's per-approval `request_id`, so a resolved/stale card clears gracefully. (#4948)
- **Docker single-container: an explicitly configured UID/GID is no longer lost, fixing a restart loop.** The entrypoint now probes the configured state-dir bind mount before the image-owned `/workspace` when auto-detecting the container UID/GID, and persists an `explicit` marker so a deliberately supplied `WANTED_UID`/`WANTED_GID` (e.g. `1024`) survives the root→user re-entry instead of being re-detected and overwritten. This fixes the documented single-container restart loop; the default stays `1024`, the two- and three-container topologies are unaffected, and a missing legacy marker falls back safely. Thanks @jorgejiro. (#7027)
- **The Z.AI model list now includes GLM-5.3.** GLM-5.3 was added to the Z.AI provider catalog so it's selectable in the model picker. The Z.AI onboarding default stays at GLM-5.1 until the direct API serves 5.3. Thanks @rh-id. (#7017)
- **A restart can no longer turn a suppressed `[SILENT]` control message into a phantom recovered turn on the `/goal` path.** The `/api/goal` endpoint now recognizes the exact reserved `[SILENT]` delivery-suppression sentinel and treats it as a successful no-op before any session lookup or goal-state mutation — matching the guard the chat-start endpoint already applies. Previously a wake relay that accidentally POSTed the sentinel to `/goal` could, after a restart, materialize it as a visible recovered user turn. Matching is exact and case-sensitive, so ordinary goals are unaffected. Thanks @webtecnica. (#7019)
- **A busy subagent tree no longer promotes an active child to a misleading top-level sidebar row.** The session sidebar builds from a bounded oversampled candidate window; after the recency slice, subagent parents that were present in the candidate set could be dropped, leaving their still-running child leaves rendered as if they were their own top-level conversations. Those parents are now re-added after the slice (bounded, cycle-safe, subagent-scoped so WebUI ancestors stay excluded), so the delegated-session tree nests correctly. Thanks @carlotestor. (#7031)
- **A shared conversation link can no longer fabricate appearance state that overrides an operator's deployment default.** The public share page (`share.html`) now applies the same first-visit appearance guard as the main app bootstrap (and picks up two skin IDs it was missing), so opening a shared link on a fresh browser no longer writes an explicit theme/skin into local storage that would then be treated as a deliberate user choice. Public-share rendering is unchanged. Thanks @webtecnica. (#7030)
- **Docker: the repository's own agent-instructions file is no longer baked into the runtime image.** A root-anchored `/AGENTS.md` entry was added to `.dockerignore` so the repo-root `AGENTS.md` (maintainer/build guidance, not user content) doesn't get copied into the Docker runtime and injected into the agent's project context. Nested workspace `AGENTS.md` files remain available exactly as before. Thanks @rodboev. (#6853)
- **Docker: the bundled SQLite is upgraded to 3.53.0 to fix a WAL-reset corruption bug (and still erases deleted rows).** The `python:3.12-slim` base image ships SQLite 3.46.1, which is vulnerable to the March 2026 WAL-reset corruption bug, and Debian hasn't backported the fix — so the image now compiles SQLite 3.53.0 from the official amalgamation (SHA-256 pinned and verified before extraction, FTS5/FTS4/R-Tree enabled, build tools purged afterward). The source build is also compiled with `SQLITE_SECURE_DELETE` and asserts `PRAGMA secure_delete = 1` at build time, so deleting a conversation still overwrites its content on disk exactly as the base image did (a from-source build without that flag would have left deleted transcript bytes recoverable in `state.db`). Only affects the Docker image. The source-built library is registered in `ld.so.conf.d` so it loads ahead of the base image's copy on every architecture (including arm64, where `/usr/local/lib` isn't on the default loader path). Thanks @qxxaa. (#6900, #7044)
- **Foregrounding a long or busy conversation no longer risks running the browser out of memory.** Switching back to a tab with a very long active chat could make the browser re-materialize the entire transcript's structured data at once — on the largest sessions this ballooned memory usage into the multi-gigabyte range and could OOM-kill the tab or the whole machine. The catch-up load on tab focus is now coalesced, the render window growth is bounded, and full structured-value allocations for off-screen turns are dropped, so foregrounding a long chat stays lightweight. A session-updated notification that arrives while a refresh is already in flight is now latched and reconciled once the refresh finishes (instead of being dropped), and the render cache signature now distinguishes tool-argument order the way the rendered output does (so differently-ordered arguments can't reuse stale cached markup). Thanks @webtecnica. (#7006, #6999)
- **The composer no longer keeps a "Reconnected" status pill sitting next to the send button.** When a live response reconnected after a brief drop, the small "Reconnected" notice was written into the composer's status slot and stayed there indefinitely — on a narrow phone layout it crowded (and could slightly clip) the send button and the other composer controls. Transient stream notices now auto-dismiss on their own: "Reconnected" clears after one second and a provider-fallback notice after four, so the notice flashes briefly and then gets out of the way. The status setter uses a single owned timer that cancels any prior countdown before arming a new one, so a newer status can't be wiped early by an older one, and a status set without a timeout (a persistent one) is left in place exactly as before. Thanks @starship-s. (#7028)
- **Self-hosted appearance defaults are reachable again on a fresh browser.** The inline theme bootstrap wrote the resolved theme/skin back to local storage on every load — including a brand-new browser with no saved preference. That made the settings sync treat the fallback (`dark`/`default`) as an explicit user choice and ignore the server-side appearance default, so operators who changed the default appearance for their instance never saw it take effect on first visit. The bootstrap now persists only when a saved preference already exists; a fresh browser keeps the exact same first paint but lets the server default win. Thanks @tomtong2015. (#6808)
- **Japanese locale: consistent wording for "profile".** The Japanese UI used two different katakana spellings for "profile" — the cron profile selector said プロフィール (the personal/social-profile loanword) while the rest of the app (profiles tab, kanban, settings) said プロファイル (the technical/config term). The two cron strings are now unified to プロファイル, matching the rest of the Japanese locale and the Hermes docs. Thanks @0809android. (#6762)
- **A stray `[SILENT]` control message can no longer turn into a phantom chat turn.** Cron agents use the exact response `[SILENT]` as a delivery-suppression sentinel; if an external wake relay accidentally POSTed that sentinel to the chat API and the agent restarted while the turn was pending, session recovery could materialize it as a visible (recovered) user message that reappeared on every restart. Both chat-ingress entry points now recognize the exact reserved sentinel and treat it as a successful no-op before any session lookup or pending-state change. Matching is exact and case-sensitive, so ordinary messages are unaffected. Thanks @allenliang2022. (#7018)
- **A busy Kanban board no longer pushes your regular conversations out of the sidebar.** The sidebar builds its session list from a bounded 20-row candidate window, and Kanban runs were competing in that window — so on a profile with 20+ recent Kanban runs, those rows could fill the window and evict your CLI / TUI / ACP / messaging conversations from the list. Kanban now gets its own dedicated, bounded recovery pass (the same pattern cron and webhook sessions already use), so a high-volume Kanban worker can't crowd out interactive conversations. Explicitly filtering the sidebar to Kanban still shows them normally. Thanks @zicochaos. (#6998)
- **The workspace file-sort menu is now fully translated in all 14 non-English languages.** The six sort options (Sort by, Name A→Z / Z→A, Date created / modified, and the "creation time unavailable" note) were still showing English placeholder text in the non-English locales after they were added in #6091; they're now translated into German, Spanish, French, Italian, Portuguese, Czech, Turkish, Polish, Russian, Japanese, Korean, Chinese (Simplified and Traditional), and Vietnamese. Thanks @webtecnica. (#6910)
- **Streaming text can no longer render into the wrong conversation when you switch sessions mid-response.** If one session was still streaming and you switched to another, a render that had already been scheduled could fire after the switch and briefly write the first session's tokens into the second session's pane. The live-stream render path now re-checks that its session is still the active one right before it writes to the DOM (including inside the deferred animation-frame/timeout window), so a scheduled render is dropped instead of leaking into the wrong pane; the streamed text is preserved and re-renders correctly when you switch back. Thanks @webtecnica. (#6502)
- **Jumping to the first response no longer snaps back to the bottom while the answer is still streaming.** When you tapped "jump to answer" on the first response, the view scrolled to the answer but then got yanked back down as more of the response streamed in, so you couldn't read from the top. The jump now takes ownership of the scroll position and holds it where you put it — it only releases when you explicitly scroll down, press End, or switch sessions. Normal bottom-following (when you haven't jumped) is unchanged. Thanks @pxxD1998. (#6621)
- **A provider that hit its credit limit (HTTP 402) recovers in ~2 minutes after you top up, instead of staying greyed out for an hour.** When a pay-as-you-go provider returned 402 (Payment Required), the WebUI marked it unavailable for a full hour, so even after you added credits it stayed unusable for up to 60 minutes. The WebUI now derives the 402 cooldown from the installed agent runtime's own credential-pool contract rather than hard-coding it, so its "is this provider usable yet" decision always matches what the runtime will actually lease (no window where the UI offers a provider the runtime still refuses). On a runtime that ships the shorter 402 cooldown the wait drops to ~120s; on an older runtime it safely keeps the previous one-hour behavior. Thanks @webtecnica. (#6626)
- **Editing your transcript (truncate, retry, or undo) no longer risks the deleted messages coming back.** When you intentionally shrink a conversation, the on-disk session ends up with fewer messages than its `.bak` sidecar — and the session-recovery inspector used to read that as suspected data loss and recommend restoring the backup, resurrecting the messages you just removed. The session now stamps a one-time provenance marker whenever you deliberately shrink a transcript, and recovery treats a valid, current marker as "this shrink was on purpose" and leaves it alone. The marker is validated strictly (a malformed or missing marker fails closed to the original data-loss safeguard), and a genuine loss — where the backup is newer than your last intentional edit — still triggers recovery exactly as before. Thanks @franksong2702. (#6954)
- **Multi-container Docker: the WebUI can now reach the agent gateway API without hand-editing compose files.** The two- and three-container Compose topologies now forward the gateway API server settings end to end — `API_SERVER_ENABLED`/`API_SERVER_HOST`/`API_SERVER_KEY` into the agent container and the matching `HERMES_WEBUI_GATEWAY_API_KEY` into the WebUI container — so setting `API_SERVER_KEY` in `.env` is enough to enable Tasks/cron and gateway-backed features (the agent only binds port 8642 when the key is a usable value). Password vars are now `${HERMES_WEBUI_PASSWORD:-}` substitutions (set once in `.env`, no compose edit), and a new `.gitattributes` forces LF on shell scripts so a Windows `core.autocrlf` checkout no longer breaks the image build (`exec …init.bash: no such file or directory`). Thanks @Yang-qwq. (#6917)
@@ -28,62 +236,17 @@
- **Multi-container Docker deployments can reach the agent API again.** In the two- and three-container Compose topologies, the agent's gateway API bound to loopback only, so the WebUI (and dashboard) containers — separate hosts on the Docker network — couldn't reach `http://hermes-agent:8642` even with `HERMES_API_URL` and `API_SERVER_KEY` set correctly, breaking cron/Tasks and gateway-backed features. The agent service now sets `API_SERVER_HOST=0.0.0.0` so it's reachable across the Compose network; the host port publish stays `127.0.0.1`-only, so this does not expose the API beyond the Docker network and host loopback. Thanks @bragov4ik. (#6968)
- **Two WebUI sessions with the same auto-generated title no longer leave the second one blank.** Follow-up to the `state.db` title sync: hermes-agent enforces a uniqueness rule on session titles, so when two sessions independently generated a byte-identical auto-title, the second sync raised (and silently swallowed) a collision error and left that session untitled in `hermes sessions list`. The sync now catches the collision, derives a de-duplicated variant via the existing lineage helper (`"My Session"` → `"My Session #2"`), and retries — and because the retry goes through the same provenance guard, a title you set manually is still never overwritten. Thanks @webtecnica. (#6964)
- **Custom-provider models keep their exact identity across catalog refreshes.** When a session used a model served through a custom provider whose upstream id contains a namespace (e.g. `z-ai/glm-5.2` under a `tokenrouter` custom endpoint), a catalog refresh could mis-resolve the picker selection — matching a *different* model whose bare suffix collided, or (in an interim fix) fabricating a new bare-id option instead of re-selecting the real one. Model resolution now matches the exact routed id first, and only falls back to a provider-hinted bare-id repair (preserving the older #6195 behavior for legacy sessions that stored just the suffix) when there is exactly one unambiguous candidate — otherwise it declines to guess. The resolved value is always an existing dropdown option, so re-renders no longer inject duplicates or lose the selection. Thanks @josephcy95. (#6944)
- **WebUI no longer crashes a send when running against an older hermes-agent.** The streaming path passed `persist_user_timestamp=` to `agent.run_conversation()` unconditionally, so a WebUI paired with an older agent whose `run_conversation()` predates that keyword raised `TypeError: unexpected keyword argument` and the message failed. The three call sites now check the callable's signature first (via a small `_supports_kwarg` introspection helper) and only pass the timestamp when the agent accepts it — the same defensive pattern already used for `moa_config`. Modern agents are unaffected (the timestamp is still passed with the identical value); older agents degrade gracefully instead of erroring. Thanks @harsh2hell. (#6935)
- **`hermes sessions list` no longer shows blank titles for WebUI sessions.** Background title generation wrote the auto-title to the WebUI session sidecar but never to hermes-agent's `state.db`, so the CLI/TUI session list showed those WebUI sessions untitled. The title is now bridged into `state.db` after it's generated (and on adaptive refresh) via `set_auto_title_if_empty`, which respects title provenance — a name you set manually via CLI/Gateway/TUI (`user` source) is never overwritten by the auto-title, an already-auto-titled row stays put, and an existing session's `source` is preserved (the sync creates the row with INSERT-OR-IGNORE semantics). Thanks @liuguangyong93. (#6892)
- **No more phantom "Compressing context" barrier after switching sessions.** The compression UI state is per-session, but `loadSession()` did not clear it on a switch, so a compression card from a prior session could leak across the load and appear as a phantom "Compressing context" barrier on a fresh session that never triggered compression. Switching sessions now clears the compression UI (state, elapsed timer, composer session-lock, and barrier element) alongside the other per-session resets, before the transcript loads. A compression genuinely running on the session you switch *to* is unaffected — it re-surfaces from that session's own live event stream (and manual compression from its server job status). Thanks @happy5318. (#6938)
- **Served media/files now revalidate with an ETag, and a mid-transfer error can never corrupt the response.** File serving (`_serve_file_bytes`) now sends an `ETag` and honors `If-None-Range`/`If-None-Match` so an unchanged image or attachment returns a cheap `304` instead of re-sending the whole body. The transfer path is also hardened: once the status line and headers are committed, any body-write failure — a client disconnect *or* a non-disconnect error like a truncated read, generic `OSError`, or `PermissionError` — is contained and logged at debug instead of escaping to the request handler, where it would have emitted a second `500` status line after the committed `200`/`206` and corrupted the HTTP stream. Thanks @happy5318. (#6922)
- **Layout test helper no longer flags scroll-reachable fields as off-viewport.** The shared `assert_layout_sane` degenerate check reported a false "interactive element is off-viewport" for an input that sits below the fold at rest but is reachable by scrolling an `overflow-y:auto` ancestor (e.g. a tall dialog on a short landscape viewport). It now exempts elements whose off-viewport edge is absorbed by a scrollable ancestor, while still flagging genuinely-unreachable elements (no scroll escape). The Kanban board settings modal layout test also drops its synthetic 480×320 short-landscape case (a viewport no real device uses, whose sub-fold geometry was environment-marginal and CI-flaky); real desktop/tablet/narrow viewports still assert the layout. (internal test infra)
### Changed
- **The System/health panel now reports WebUI runtime memory diagnostics.** Beyond host CPU/RAM/disk, the system-health response now observes the process-owned caches and stream structures named in #6351 (session-list cache, agent cache, live-stream structures) by reading their existing owners non-blockingly — a collector that never waits on an owner lock, so it can't stall a request. This surfaces sizes/bounds that were previously invisible for diagnosing runaway memory. Thanks @rodboev. (#6363)
- **Slash-command autocomplete now searches skills by keyword.** Typing `/` and a term now matches plain skills on a case-insensitive keyword in their **name or description** (previously only a name prefix), so a skill is discoverable by what it does, not just its exact slug. Built-in, agent/plugin, and bundle commands keep prefix matching and always take precedence over a same-slug skill, and skill suggestions are held until the bundle-metadata request settles so a slow response can't briefly surface a shadowed skill. Thanks @MinhoJJang. (#6932)
- **Extensions can observe turn lifecycle events via `handle.events.on(...)`.** A registered extension (see `window.hermesExt.register`) can now subscribe to session turn-lifecycle events — `turn:start`, `turn:complete`, `turn:error`, and `turn:cancel` — through a frozen `events` accessor on its handle, each delivering a frozen scalar event (`sessionId`, `streamId`, `status`, and start/end timestamps; no message content). Core owns the authoritative live-session stream and emits each event exactly once per real transition (a per-session/stream state machine suppresses duplicate starts and post-terminal events, including on reconnect/replay), with per-listener error isolation so a misbehaving extension cannot break streaming. Subscriptions return an idempotent unsubscribe. This is the B1 observer surface on top of the E0 identity handle — a stable page-local facade instead of extensions polling private DOM state. Thanks @franksong2702. (#6924)
- **Extensions can claim a cooperative page identity via `window.hermesExt.register(id)`.** An injected extension can now claim a stable browser-page identity together with its existing settings and storage accessors: `register(id)` returns a frozen `{id, settings, storage}` handle for an ID present in the effective-enabled, Core-sanitized manifest inventory captured at page boot, and returns `null` (fail-closed, no namespace side effects) for empty, malformed, unknown, or untrusted IDs. Re-registering the same ID in one page returns the same handle; boot-time trust is immutable across status refreshes (a refresh cannot elevate an ID absent at boot). This is cooperative attribution for extensions sharing the page — explicitly not a sandbox, capability, permission, isolation mechanism, or security boundary. Thanks @franksong2702. (#6795)
- **Refine selected transcript text directly in the composer.** Selecting text inside a chat message now shows a "Refine" action alongside "Reply with selection" in the floating selection toolbar. Refine seeds the selected text into the composer as an editable draft (a quoted block plus a localized "Refine instruction:" line, caret ready) so you can shape a follow-up before sending — the ChatGPT/Notion-style quote-to-edit flow. The transfer is one-way and marker-free (no internal sentinel reaches the composer or the sent message), preserves any existing draft and attachments, collapses to 44×44 icon-only buttons on phones, and covers all 15 locales. Thanks @rodboev. (#6410)
- **Agent replay is now byte-accurate without ever exposing internal provider text to the browser.** The exact provider-facing reply text (`api_content`) is preserved through the internal session store so agent-to-agent replay is faithful, while every client-facing surface (HTTP responses, SSE journal + runner streams, shares) strips it and its provenance aliases through the public projection. Includes a maintainer hardening pass: session redaction is selective (transcript fields are scrubbed; operational fields such as the workspace path pass through unchanged, so a workspace path that merely contains credential-shaped text is no longer masked into an invalid path), and the runner-backed SSE relay now projects payloads before emit so `api_content` cannot leak through that path. Thanks @franksong2702. (#6757)
- **Transparent Stream copy buttons now confirm success in place.** The tool-event and thinking copy controls flash a check for ~1.5s after a successful copy — the same feedback the normal chat copy buttons already give — then revert to the copy glyph. The confirmation is now clearly visible without hover (important on touch), stays correct across live-row reconciliation, cleans up if its row is torn down mid-feedback, and is never serialized into a restored session snapshot. Thanks @Stacey2911. (#6037)
- **Kanban boards can now set a default workspace path.** The board settings modal gains a "Default workspace path" field (a type-or-pick combobox sourced from your saved workspaces, so you can choose a known path or type a custom one — validated server-side). Tasks on the board without an explicit workspace path inherit this default. The board switcher chip is now always shown (even with a single board) and its menu item is now "Board settings…", enabled for the default board. Thanks @rodboev. (#6055, #5932)
- **The workspace file tree can now be sorted, and the "Show hidden files" toggle moved into a menu.** The workspace pane's ⋮ options menu gains a "Sort by" group — Name (A→Z), Name (Z→A), Date created, Date modified — and your choice persists across sessions. "Date created" is disabled with an inline note on platforms/servers that don't report creation time. The pre-existing "Show hidden files" toggle moves from a permanent inline row into that same menu, reclaiming vertical space on every workspace view; a small indicator on the heading still reflects when hidden files are shown. Thanks @rodboev. (#6091, #6066)
- **Kanban worker sessions can now be hidden from the default sidebar.** A new "Show kanban sessions" toggle (nested under "Show non-WebUI sessions", default off) keeps automated kanban worker runs — which can flood the list — out of the default sidebar, extending the same background-session hide already used for cron and webhook sessions. Turning it on shows them; an explicit kanban source filter always reveals them regardless of the toggle, and a project-assigned kanban session still surfaces under its project chip. Thanks @vduke. (#6780)
- **Typography is now driven by three explicit font tokens (`--font-ui`, `--font-conversation`, `--font-mono`), so a theme can style conversation prose separately from UI chrome.** The shell chrome/controls use `--font-ui`, message prose uses `--font-conversation` (which defaults to `--font-ui`, so there is no change to the default appearance — assistant/user prose stays on the system sans, not a serif), and code/logs/terminal use `--font-mono`. The embedded terminal now also re-fits its glyph metrics after web fonts finish loading (guarded against double-fit and disposed terminals) so its layout stays correct. This is groundwork that lets skins opt into a distinct reading font without changing the default for everyone. Thanks @starship-s. (#5918)
- **Active in-flight recovery snapshots are now sent in a compact transport form.** When you reload or reattach to a session with a run in progress, the `/api/session` payload no longer carries multiple redundant copies of each tool result and row value — it projects the live run-journal snapshot into a smaller, display-equivalent form (one authoritative source per value) while preserving the live message's timestamp so the reconstructed turn keeps stable identity. The durable journal and the canonical in-process snapshot are unchanged, so recovery renders identically with less payload. Thanks @allenliang2022. (#6858)
- **Background-process completions and watch matches now render as compact, collapsible summary cards.** Routine wakeups no longer fill the transcript with an always-expanded agent-facing prompt: the collapsed row shows the command, exit status or matched watch pattern, and timestamp, while one click reveals the complete command and verbatim output. Failed processes keep a red accent, long watch patterns remain reachable on touch and keyboard when expanded, and unfamiliar future event shapes safely retain the existing raw-notice fallback. Thanks @b3nw. (#6350, #6345)
- **The repository is now pip-installable with standard packaging metadata.** `pyproject.toml` gains a `[build-system]` + `[project]` section (setuptools + setuptools-scm) so packagers and distros can `pip install .` and get a proper wheel (`api` + `static` bundled). The normal `bootstrap.py` / `start.sh` / `ctl.sh` source-checkout launch path is unchanged. setuptools-scm writes its version to a separate `api/_scm_version.py` (not the Docker-owned `api/_version.py`), and the runtime version detector reads it as an explicit fallback (normalizing the PEP 440 value to a `v…` form) so an installed wheel reports a real version instead of `unknown` — Docker and git-checkout version resolution are byte-identical to before. Thanks @rodboev. (#6337, #2695)
- **Internal: removed a dead `rowIndex` parameter from the settled-scene row projector.** `attachLiveStream`'s `pushRow` helper carried an unused positional index left over from before final-segment eligibility moved to a row-identity `WeakSet`; dropping it is a pure no-op refactor (differential execution confirms byte-identical settled scenes — ordering, dedupe, final-prefix suppression, and sequence assignment all unchanged). Thanks @webtecnica. (#6258)
- **Internal: auxiliary-title test suites no longer leak a synthetic `agent` module across the process.** `test_title_aux_routing.py` and `test_2235_initial_aux_title.py` installed stub `agent` / `agent.auxiliary_client` modules via `sys.modules.setdefault` at import time and never removed them, so any later test importing `agent.context_compressor` in the same process got the stub and failed (breaking `test_issue4685_post_compression_context_metering.py` by run order). Both suites now route the install through a shared `tests/_aux_client_helpers.py` chokepoint scoped with `unittest.mock.patch.dict`, plus four tests covering restore-by-identity, absent-key cleanup, present-but-`None` preservation, and restore-on-raise. Test-only, no production change. Thanks @rodboev. (#6661, #6630)
- **Internal: clarified the transcript-virtualization boot comment (opt-in default retained).** The comment above the boot-time `_virtualizeTranscript` apply now records that the #4346 Phase B footer-jitter fix resolved the original scroll-up flicker root cause, while virtualization stays opt-in (default OFF) until battle-tested further. Zero behavior change — the apply predicate, settings-load fallback, checkbox, and server default all remain opt-in/OFF and in agreement. Thanks @webtecnica. (#6318, #6155)
- **Internal: added regression coverage for the empty-`tool_calls` drop on the metadata-restoration mirror path.** `_api_safe_message_positions()` must apply the same `tool_calls: []` drop as the main API sanitizer (#5737) so metadata restoration never treats a `tool_calls: []` row as a distinct message; a new test pins that empty arrays lose the key while populated tool-call chains (and their tool results) survive untouched. Test-only, no production change. Thanks @webtecnica. (#6803, #6796)
- **Internal: added a two-direction regression for identical-prompt retry settlement.** The active-turn ownership repair that resolves #6571 (settlement materializes the exact pending turn before checking for a terminal assistant answer) had no issue-specific coverage because the original reproduction omitted the runtime turn identity. A new test pins both directions through the real `_merged_transcript_lacks_final_assistant_answer()`: a successful identical-prompt retry keeps its terminal answer (not a `no_response`), while an otherwise-identical retry with no terminal answer still qualifies as a genuine failure. Test-only, no production change (a mutation bite — nulling the merge's `active_turn_identity` provenance — flips the successful case red, confirming the guard is live). (#6631, #6571)
- **The busy-time send behavior is now called "Default message mode," and new installs default to Steer.** The Settings → Preferences control formerly labeled "Busy input mode" is renamed to "Default message mode," and a fresh install now defaults to **Steer** (inject a mid-turn correction without interrupting) instead of Queue. Your existing choice is preserved — if you ever saved settings, your current mode (Queue/Interrupt/Steer) is migrated as-is and unchanged; only never-configured installs pick up the new Steer default. The saved preference still survives a reload or a brief server outage (the localStorage mirror from the previous release is intact). Thanks @rodboev. (#5162, #5145)
### Added
- **Each settled assistant turn's footer now shows the model that actually served it.** In transparent-stream mode the turn footer gains a compact model chip (e.g. `8s · claude-opus-4-8 · TTFT 640ms · …`), read *after* the turn completes so a mid-turn fallback shows the real model that answered rather than the one originally requested. The label renders exactly once per turn — a gateway/failover turn keeps its existing routing chip and the additive chip is suppressed — and the footer wraps gracefully at narrow widths and on mobile instead of stretching. The used-model is persisted with the session so it survives a reload. Thanks @franksong2702. (#6113, #6068)
### Fixed
- **Screen readers no longer announce multiple sidebar panels as "expanded" at once.** Switching between rail panels (Chat → Kanban → Skills → …) set `aria-expanded="true"` on each newly-active rail button without clearing it on the previous one, so the attribute accumulated and assistive tech announced several open disclosures simultaneously. The active-panel sync now clears `aria-expanded` on all rail nav buttons before marking the single active one, so exactly one button reflects the sidebar's open/collapsed state. Thanks @rodrigogs. (#6656)
- **User-uploaded images now follow the session's actual model, not the global default.** The routing decision that picks whether an uploaded image is embedded natively (as `image_url` parts) or routed through the text/vision-analyze pipeline was resolved from the global config default, ignoring the model the current session actually uses. A session switched to a text-only model would still have images forwarded natively to a model that can't see them, and a session on a vision-capable model could be degraded to text if the global default happened to be text-only. The session's resolved provider/model (including the pre-canonicalization `custom:slug` identity) is now threaded through the routing decision, and the credential-heal retry path preserves that identity so a text-only custom model keeps stripping historical images across an auth retry. No behavior change when no session identity is present (global default, as before). Thanks @happy5318. (#6882)
@@ -224,22 +387,6 @@
- **Async delegation (`delegate_task`) results now wake the exact browser tab that started them, exactly once, even across a WebUI restart.** Two problems on the async-completion delivery path are fixed together. (1) *Misroute:* a detached completion was routed by a mutable session-key index, so a result could wake the wrong session; it now routes by the immutable `origin_ui_session_id` return address captured from the commissioning turn, which is authoritative over the index on both delivery paths (the background wakeup and the next-turn drain). (2) *Duplicate on restart:* the completion is now delivered through a durable claim/complete/release lifecycle and acknowledged only after the wakeup turn is accepted, so the restart-restore sweep no longer re-enqueues an already-delivered result. Every failure path (formatter, routing, turn-start, ACK) releases the claim and keeps the result retryable rather than dropping it, and an unrouteable async event is handed to a bounded retry instead of being silently stranded. Thanks @carlotestor (#6185) and @sysophelper-droid (#6159). (#6185, #6159, #6002, #6225)
- **The Artifacts tab now shows each artifact's file name first, with the directory path secondary.** Artifact rows previously rendered the full path on one line, so the file name (the part you scan for) was truncated at the end of a long path. The name now renders prominently with the directory shown smaller and muted beneath it. Thanks @webtecnica. (#6161)
- **A session row's streaming indicator now reflects that session's own activity, not a child session's.** The sidebar "streaming" state was derived from a combined flag that was also true when a *child* (subagent) session was streaming, so a row could show as busy on its child's behalf. It now keys off the session's own streaming state. Thanks @webtecnica. (#6165)
- **`prefers-reduced-motion` now also suppresses the message-row color transition.** Message rows were missing from the reduced-motion transition-disable rule, so a color change (e.g. on theme or state change) still animated for users who asked for reduced motion. `.msg-row` is now included. Thanks @webtecnica. (#6166)
- **Deliberately picking a model whose bare id also exists under the default provider no longer reverts to the default on re-render.** When a model id (e.g. `glm-5.2`) is offered by both the default provider and another provider, and the provider hint was momentarily unavailable during a dropdown rebuild, the picker's normalized-match fallback snapped to the first (default) group — silently reverting a deliberate non-default pick. The resolver now returns no match for an ambiguous bare id with no provider hint, so the caller injects the correctly provider-scoped option instead of asserting the wrong provider; hinted, uniquely-named, and exact-match lookups are unchanged. Thanks @webtecnica. (#6199, #6195)
- **Model-emitted base64 images now render as images instead of dumping raw base64 text.** A `data:image/*` URI in an assistant message — whether as `![alt](data:image/png;base64,...)`, a raw `<img src="data:...">`, or a `MEDIA:data:...` reference — previously rendered as a wall of base64 text (or was swallowed entirely), and `![alt](file:///x.png)` produced a stray `!<a>` anchor. Data-image URIs now render inline under a strict allowlist (raster formats plus base64-only SVG, a safe payload charset, and a 2 MB cap), and `![](file://)` routes through the media pipeline. Every other `data:` scheme (`data:text/html`, etc.) stays inert text and is never embedded; the sanitizer permits `data:` for `<img>` only, and the streaming renderer enforces the same predicate so an image can't render mid-stream then vanish at settle. Base64 SVG renders via `<img src>` (scripts can't execute in that context). Thanks @ai-ag2026. (#6209)
- **A single configured model no longer appears multiple times in the model picker.** A configured model was registered under its bare id, `provider/model`, AND `@provider:model` forms, so the picker treated equivalent routing ids as separate entries and showed the same model more than once. Dedup now collapses only genuinely-equivalent entries (same normalized id **and** same provider), so two different providers offering the same bare model id (e.g. both offer `gpt-4o`) stay distinct while a truly-one-model-registered-three-ways case collapses to one badge — applied to both the live and static catalog builders. A picked provider badge resolves, sends, and persists as that provider (no snap-back to a first-match provider on a bare-id collision), and a colon-bearing model id (e.g. `model-a:free`) synthesized as a missing-catalog fallback reads its authoritative `data-provider` rather than mis-parsing at the last colon. Thanks @happy5318. (#6221, #3691)
- **Reopening a WebUI tab after a large active session is fast again, even when the session has few messages but large tool outputs.** The reconnect display-path tail optimization only fired based on message *count*, so a session with a handful of messages but multi-MB tool-call JSON forced a slow full-scan merge (bootstrap could take tens of seconds). The optimization now also fires when the session sidecar file exceeds 500 KB, regardless of message count — the same messages are returned (verified: identical message set, count, offset, and truncation metadata), just via the fast path. Thanks @webtecnica. (#6260)
- **OIDC allowlist values containing spaces (e.g. a group named `Hermes Users`) are no longer split into separate entries.** The allowlist parser split comma-separated values on all whitespace, so a multi-word group/user name became two spurious entries. Allowlist parsing now splits only on commas/newlines (multi-word values stay intact and match their exact signed claim — strictly tightening, no security loosening), while OAuth **scope** parsing keeps its space-delimited behavior per RFC 6749 §3.3. A blank allowlist entry no longer bricks an OIDC-only deployment. Thanks @webtecnica. (#6244)
- **Concurrent remote-gateway health checks no longer stampede the gateway.** When many requests needed a remote-gateway health probe at once, each fired its own probe (a thundering herd). Probes are now single-flighted: one "leader" thread probes while latecomers wait (bounded) and share the result, guarded by a single condition/lock with a `finally` that always releases waiters even if the probe errors. Thanks @ai-ag2026. (#5798, #5455)
- **Appending run-journal events no longer gets slower as a session's journal grows.** The next sequence number was recomputed by scanning existing entries on every append (O(n²) over a session's lifetime); it's now cached per journal path (guarded by a dedicated lock, evicted on journal delete), making each append O(1). Thanks @ai-ag2026. (#5799)
@@ -460,6 +607,54 @@
- **No more 57–70 s cold-startup stalls from the profile skills-stats thundering herd.** At container boot the frontend fires several profile-data requests at once; with `ThreadingHTTPServer` (one thread per request) they all missed the empty skills-stats cache simultaneously and each walked + parsed every profile's skill tree, stacking thousands of concurrent `stat()` calls under Docker's overlay2 filesystem. `_get_profile_skills_stats()` now serializes per-profile with double-checked locking (concurrent misses on one profile collapse to a single compute; independent profiles still compute in parallel), and `list_profiles_api()` single-flights the row build under `_LIST_PROFILES_CACHE_LOCK` so one thread builds while the rest wait for the cached result. The every-call cheap mtime probe (the #4783 out-of-band change-detection contract) is unchanged. (#5364)
### Changed
- **The System/health panel now reports WebUI runtime memory diagnostics.** Beyond host CPU/RAM/disk, the system-health response now observes the process-owned caches and stream structures named in #6351 (session-list cache, agent cache, live-stream structures) by reading their existing owners non-blockingly — a collector that never waits on an owner lock, so it can't stall a request. This surfaces sizes/bounds that were previously invisible for diagnosing runaway memory. Thanks @rodboev. (#6363)
- **Slash-command autocomplete now searches skills by keyword.** Typing `/` and a term now matches plain skills on a case-insensitive keyword in their **name or description** (previously only a name prefix), so a skill is discoverable by what it does, not just its exact slug. Built-in, agent/plugin, and bundle commands keep prefix matching and always take precedence over a same-slug skill, and skill suggestions are held until the bundle-metadata request settles so a slow response can't briefly surface a shadowed skill. Thanks @MinhoJJang. (#6932)
- **Extensions can observe turn lifecycle events via `handle.events.on(...)`.** A registered extension (see `window.hermesExt.register`) can now subscribe to session turn-lifecycle events — `turn:start`, `turn:complete`, `turn:error`, and `turn:cancel` — through a frozen `events` accessor on its handle, each delivering a frozen scalar event (`sessionId`, `streamId`, `status`, and start/end timestamps; no message content). Core owns the authoritative live-session stream and emits each event exactly once per real transition (a per-session/stream state machine suppresses duplicate starts and post-terminal events, including on reconnect/replay), with per-listener error isolation so a misbehaving extension cannot break streaming. Subscriptions return an idempotent unsubscribe. This is the B1 observer surface on top of the E0 identity handle — a stable page-local facade instead of extensions polling private DOM state. Thanks @franksong2702. (#6924)
- **Extensions can claim a cooperative page identity via `window.hermesExt.register(id)`.** An injected extension can now claim a stable browser-page identity together with its existing settings and storage accessors: `register(id)` returns a frozen `{id, settings, storage}` handle for an ID present in the effective-enabled, Core-sanitized manifest inventory captured at page boot, and returns `null` (fail-closed, no namespace side effects) for empty, malformed, unknown, or untrusted IDs. Re-registering the same ID in one page returns the same handle; boot-time trust is immutable across status refreshes (a refresh cannot elevate an ID absent at boot). This is cooperative attribution for extensions sharing the page — explicitly not a sandbox, capability, permission, isolation mechanism, or security boundary. Thanks @franksong2702. (#6795)
- **Refine selected transcript text directly in the composer.** Selecting text inside a chat message now shows a "Refine" action alongside "Reply with selection" in the floating selection toolbar. Refine seeds the selected text into the composer as an editable draft (a quoted block plus a localized "Refine instruction:" line, caret ready) so you can shape a follow-up before sending — the ChatGPT/Notion-style quote-to-edit flow. The transfer is one-way and marker-free (no internal sentinel reaches the composer or the sent message), preserves any existing draft and attachments, collapses to 44×44 icon-only buttons on phones, and covers all 15 locales. Thanks @rodboev. (#6410)
- **Agent replay is now byte-accurate without ever exposing internal provider text to the browser.** The exact provider-facing reply text (`api_content`) is preserved through the internal session store so agent-to-agent replay is faithful, while every client-facing surface (HTTP responses, SSE journal + runner streams, shares) strips it and its provenance aliases through the public projection. Includes a maintainer hardening pass: session redaction is selective (transcript fields are scrubbed; operational fields such as the workspace path pass through unchanged, so a workspace path that merely contains credential-shaped text is no longer masked into an invalid path), and the runner-backed SSE relay now projects payloads before emit so `api_content` cannot leak through that path. Thanks @franksong2702. (#6757)
- **Transparent Stream copy buttons now confirm success in place.** The tool-event and thinking copy controls flash a check for ~1.5s after a successful copy — the same feedback the normal chat copy buttons already give — then revert to the copy glyph. The confirmation is now clearly visible without hover (important on touch), stays correct across live-row reconciliation, cleans up if its row is torn down mid-feedback, and is never serialized into a restored session snapshot. Thanks @Stacey2911. (#6037)
- **Kanban boards can now set a default workspace path.** The board settings modal gains a "Default workspace path" field (a type-or-pick combobox sourced from your saved workspaces, so you can choose a known path or type a custom one — validated server-side). Tasks on the board without an explicit workspace path inherit this default. The board switcher chip is now always shown (even with a single board) and its menu item is now "Board settings…", enabled for the default board. Thanks @rodboev. (#6055, #5932)
- **The workspace file tree can now be sorted, and the "Show hidden files" toggle moved into a menu.** The workspace pane's ⋮ options menu gains a "Sort by" group — Name (A→Z), Name (Z→A), Date created, Date modified — and your choice persists across sessions. "Date created" is disabled with an inline note on platforms/servers that don't report creation time. The pre-existing "Show hidden files" toggle moves from a permanent inline row into that same menu, reclaiming vertical space on every workspace view; a small indicator on the heading still reflects when hidden files are shown. Thanks @rodboev. (#6091, #6066)
- **Kanban worker sessions can now be hidden from the default sidebar.** A new "Show kanban sessions" toggle (nested under "Show non-WebUI sessions", default off) keeps automated kanban worker runs — which can flood the list — out of the default sidebar, extending the same background-session hide already used for cron and webhook sessions. Turning it on shows them; an explicit kanban source filter always reveals them regardless of the toggle, and a project-assigned kanban session still surfaces under its project chip. Thanks @vduke. (#6780)
- **Typography is now driven by three explicit font tokens (`--font-ui`, `--font-conversation`, `--font-mono`), so a theme can style conversation prose separately from UI chrome.** The shell chrome/controls use `--font-ui`, message prose uses `--font-conversation` (which defaults to `--font-ui`, so there is no change to the default appearance — assistant/user prose stays on the system sans, not a serif), and code/logs/terminal use `--font-mono`. The embedded terminal now also re-fits its glyph metrics after web fonts finish loading (guarded against double-fit and disposed terminals) so its layout stays correct. This is groundwork that lets skins opt into a distinct reading font without changing the default for everyone. Thanks @starship-s. (#5918)
- **Active in-flight recovery snapshots are now sent in a compact transport form.** When you reload or reattach to a session with a run in progress, the `/api/session` payload no longer carries multiple redundant copies of each tool result and row value — it projects the live run-journal snapshot into a smaller, display-equivalent form (one authoritative source per value) while preserving the live message's timestamp so the reconstructed turn keeps stable identity. The durable journal and the canonical in-process snapshot are unchanged, so recovery renders identically with less payload. Thanks @allenliang2022. (#6858)
- **Background-process completions and watch matches now render as compact, collapsible summary cards.** Routine wakeups no longer fill the transcript with an always-expanded agent-facing prompt: the collapsed row shows the command, exit status or matched watch pattern, and timestamp, while one click reveals the complete command and verbatim output. Failed processes keep a red accent, long watch patterns remain reachable on touch and keyboard when expanded, and unfamiliar future event shapes safely retain the existing raw-notice fallback. Thanks @b3nw. (#6350, #6345)
- **The repository is now pip-installable with standard packaging metadata.** `pyproject.toml` gains a `[build-system]` + `[project]` section (setuptools + setuptools-scm) so packagers and distros can `pip install .` and get a proper wheel (`api` + `static` bundled). The normal `bootstrap.py` / `start.sh` / `ctl.sh` source-checkout launch path is unchanged. setuptools-scm writes its version to a separate `api/_scm_version.py` (not the Docker-owned `api/_version.py`), and the runtime version detector reads it as an explicit fallback (normalizing the PEP 440 value to a `v…` form) so an installed wheel reports a real version instead of `unknown` — Docker and git-checkout version resolution are byte-identical to before. Thanks @rodboev. (#6337, #2695)
- **Internal: auxiliary-title test suites no longer leak a synthetic `agent` module across the process.** `test_title_aux_routing.py` and `test_2235_initial_aux_title.py` installed stub `agent` / `agent.auxiliary_client` modules via `sys.modules.setdefault` at import time and never removed them, so any later test importing `agent.context_compressor` in the same process got the stub and failed (breaking `test_issue4685_post_compression_context_metering.py` by run order). Both suites now route the install through a shared `tests/_aux_client_helpers.py` chokepoint scoped with `unittest.mock.patch.dict`, plus four tests covering restore-by-identity, absent-key cleanup, present-but-`None` preservation, and restore-on-raise. Test-only, no production change. Thanks @rodboev. (#6661, #6630)
- **Internal: clarified the transcript-virtualization boot comment (opt-in default retained).** The comment above the boot-time `_virtualizeTranscript` apply now records that the #4346 Phase B footer-jitter fix resolved the original scroll-up flicker root cause, while virtualization stays opt-in (default OFF) until battle-tested further. Zero behavior change — the apply predicate, settings-load fallback, checkbox, and server default all remain opt-in/OFF and in agreement. Thanks @webtecnica. (#6318, #6155)
- **Internal: added regression coverage for the empty-`tool_calls` drop on the metadata-restoration mirror path.** `_api_safe_message_positions()` must apply the same `tool_calls: []` drop as the main API sanitizer (#5737) so metadata restoration never treats a `tool_calls: []` row as a distinct message; a new test pins that empty arrays lose the key while populated tool-call chains (and their tool results) survive untouched. Test-only, no production change. Thanks @webtecnica. (#6803, #6796)
- **Internal: added a two-direction regression for identical-prompt retry settlement.** The active-turn ownership repair that resolves #6571 (settlement materializes the exact pending turn before checking for a terminal assistant answer) had no issue-specific coverage because the original reproduction omitted the runtime turn identity. A new test pins both directions through the real `_merged_transcript_lacks_final_assistant_answer()`: a successful identical-prompt retry keeps its terminal answer (not a `no_response`), while an otherwise-identical retry with no terminal answer still qualifies as a genuine failure. Test-only, no production change (a mutation bite — nulling the merge's `active_turn_identity` provenance — flips the successful case red, confirming the guard is live). (#6631, #6571)
- **The busy-time send behavior is now called "Default message mode," and new installs default to Steer.** The Settings → Preferences control formerly labeled "Busy input mode" is renamed to "Default message mode," and a fresh install now defaults to **Steer** (inject a mid-turn correction without interrupting) instead of Queue. Your existing choice is preserved — if you ever saved settings, your current mode (Queue/Interrupt/Steer) is migrated as-is and unchanged; only never-configured installs pick up the new Steer default. The saved preference still survives a reload or a brief server outage (the localStorage mirror from the previous release is intact). Thanks @rodboev. (#5162, #5145)
### Added
- **GLM-5.3 Flash (`zai/glm-5.3-flash`) is now selectable in the Z.AI model list.** Z.ai's fast/cheap native-multimodal sibling of GLM-5.3 is added to the direct Z.AI catalog and the provider fallback list, following the same base-then-flash convention as the other GLM entries and inheriting the ≥5.2 reasoning-effort ladder. The onboarding default deliberately stays GLM-5.1. Thanks @rh-id. (#7323)
- **Extensions can register a native "Configure" entry point under Settings → Extensions → Installed.** For complex editors that don't fit scalar `settings_schema` fields, a boot-trusted extension can call `ext.settings.registerConfigure(handler)` on the E0 handle to add one native Configure button for its current effective-enabled Installed row, while keeping the editor UI, validation, state, and persistence fully extension-owned. Registration is trust-gated and quarantine-aware, the first handler per extension wins (duplicates return `null`), and it returns an idempotent unregister; the button never appears in Diagnostics and is removed immediately on disable/uninstall. Thanks @franksong2702. (#7111)
- **Each settled assistant turn's footer now shows the model that actually served it.** In transparent-stream mode the turn footer gains a compact model chip (e.g. `8s · claude-opus-4-8 · TTFT 640ms · …`), read *after* the turn completes so a mid-turn fallback shows the real model that answered rather than the one originally requested. The label renders exactly once per turn — a gateway/failover turn keeps its existing routing chip and the additive chip is suppressed — and the footer wraps gracefully at narrow widths and on mobile instead of stretching. The used-model is persisted with the session so it survives a reload. Thanks @franksong2702. (#6113, #6068)
### Performance
- **Fork a scheduled-job (cron) conversation into your own editable chat.** Cron-run sessions are read-only (the scheduler owns them), so you couldn't continue one. You can now branch/fork a cron session into a new WebUI-owned conversation — the original read-only session is never modified, and the fork picks up its history so you can carry on. Only genuine cron sessions qualify (verified by the server-side source, not a guessable id), and every other read-only session stays unbranchable. Thanks @rodboev. (#5555, #5477)
@@ -490,6 +685,10 @@
### Documentation
- **Remote access docs now lead with Tailscale Serve**, which keeps the WebUI bound to loopback while giving your tailnet an HTTPS MagicDNS hostname, including the Linux `serve config denied` permission fix (`tailscale set --operator=$USER`). Binding to `0.0.0.0` on the tailnet is documented as the fallback, with password auth called out as required. Thanks @evgenyponomarev. (#7430, closes #7419)
- The conversation-lifecycle regression gate test now records the proposal → implementation link in its header. Thanks @webtecnica. (#6330)
- **Documented the reverse-proxy basic-auth caveat for installed PWAs.** README, onboarding, and troubleshooting now explain that proxy basic auth can block the same-origin service-worker/shell update fetches an installed PWA needs (leaving it on a blank screen after an update), and give recovery steps — prefer WebUI's built-in password, or scope proxy auth so `sw.js`/manifest/shell requests complete. Thanks @rodboev. (#5673, #2781)
- **Documented the `HERMES_WEBUI_*` environment variables and refreshed stale version/test/locale references.** README, ARCHITECTURE, TESTING, ROADMAP, SPRINTS, and the `.env` example files now describe the supported runtime env vars (host/port/state-dir/agent-dir/home) and no longer carry stale hardcoded version/test-count/locale figures. Thanks @mo7al876any. (#5536)
@@ -498,6 +697,8 @@
### Internal
- **Extracted the 585-line inline `GET /api/session` handler out of the giant `handle_get` dispatcher into a dedicated `_handle_session_get` function.** Behavior-preserving refactor: the extracted body is AST-identical to the original inline block, the route condition and its precedence relative to sibling routes (`/api/sessions`, `/api/session/lineage/report`, `/api/projects`) are unchanged, and response projection, redaction, and error handling are byte-for-byte the same. Makes the most-hit session route far easier to read and maintain. Thanks @zicochaos. (#7340)
- **Hardened `test_tls_support::test_tls_startup_failure_fallback_to_http` against stdout-preamble noise.** The test asserted the server prints "TLS setup failed" by reading only the first 2000 bytes of its stdout. When the test interpreter lacks optional deps (`requests`/`websockets`), startup emits a burst of plugin/tool import warnings plus the startup-config banner *before* the TLS line, pushing the marker past that fixed window → spurious failure (distinct from the agent-package-isolation cause fixed separately below). It now drains all currently-available stdout (bounded by a 5s deadline, short-circuiting once the marker is seen) instead of a single fixed-size read. Verified failing→passing in a fresh venv without the optional deps. Test-only.
- **Fixed the last 2 chronic local full-suite test-isolation flakes.** `test_tls_support::test_tls_startup_failure_fallback_to_http` and `test_v050259_sessiondb_fd_leak::test_session_db_close_is_idempotent` failed only in a full local suite run (they skip on CI, where the agent package isn't co-located). Root cause was the same "simulate agent package unavailable" antipattern shipped-fixed for the profile cluster: `test_custom_provider_prefix_collisions` clobbered `sys.modules['agent']` at collection time (never restored), and several helpers emptied real package `__path__` lists in place. Fixed at the source — the collection-time shim is now guarded on `importlib.util.find_spec` (only installed when the real package is genuinely absent), the in-place `__path__` mutations use `monkeypatch.setattr(..., raising=False)`, and `test_issue1574`'s fake-agent helper restores the env/`sys.path` it mutates. Full local suite now 0 failed (was 2); `test_title_aux_routing` unaffected. Test-only.
@@ -512,6 +713,32 @@
- **Added regression coverage for messaging clear-watermark semantics.** New test locks the watermark behavior so a future change can't silently regress it. Thanks @rodboev. (#5589, #5572)
## [v0.52.113] — 2026-08-14
### Fixed
- **Custom-provider models keep their exact identity across catalog refreshes.** When a session used a model served through a custom provider whose upstream id contains a namespace (e.g. `z-ai/glm-5.2` under a `tokenrouter` custom endpoint), a catalog refresh could mis-resolve the picker selection — matching a *different* model whose bare suffix collided, or (in an interim fix) fabricating a new bare-id option instead of re-selecting the real one. Model resolution now matches the exact routed id first, and only falls back to a provider-hinted bare-id repair (preserving the older #6195 behavior for legacy sessions that stored just the suffix) when there is exactly one unambiguous candidate — otherwise it declines to guess. The resolved value is always an existing dropdown option, so re-renders no longer inject duplicates or lose the selection. Thanks @josephcy95. (#6944)
- **The Artifacts tab now shows each artifact's file name first, with the directory path secondary.** Artifact rows previously rendered the full path on one line, so the file name (the part you scan for) was truncated at the end of a long path. The name now renders prominently with the directory shown smaller and muted beneath it. Thanks @webtecnica. (#6161)
- **A session row's streaming indicator now reflects that session's own activity, not a child session's.** The sidebar "streaming" state was derived from a combined flag that was also true when a *child* (subagent) session was streaming, so a row could show as busy on its child's behalf. It now keys off the session's own streaming state. Thanks @webtecnica. (#6165)
- **`prefers-reduced-motion` now also suppresses the message-row color transition.** Message rows were missing from the reduced-motion transition-disable rule, so a color change (e.g. on theme or state change) still animated for users who asked for reduced motion. `.msg-row` is now included. Thanks @webtecnica. (#6166)
- **Deliberately picking a model whose bare id also exists under the default provider no longer reverts to the default on re-render.** When a model id (e.g. `glm-5.2`) is offered by both the default provider and another provider, and the provider hint was momentarily unavailable during a dropdown rebuild, the picker's normalized-match fallback snapped to the first (default) group — silently reverting a deliberate non-default pick. The resolver now returns no match for an ambiguous bare id with no provider hint, so the caller injects the correctly provider-scoped option instead of asserting the wrong provider; hinted, uniquely-named, and exact-match lookups are unchanged. Thanks @webtecnica. (#6199, #6195)
- **Model-emitted base64 images now render as images instead of dumping raw base64 text.** A `data:image/*` URI in an assistant message — whether as `![alt](data:image/png;base64,...)`, a raw `<img src="data:...">`, or a `MEDIA:data:...` reference — previously rendered as a wall of base64 text (or was swallowed entirely), and `![alt](file:///x.png)` produced a stray `!<a>` anchor. Data-image URIs now render inline under a strict allowlist (raster formats plus base64-only SVG, a safe payload charset, and a 2 MB cap), and `![](file://)` routes through the media pipeline. Every other `data:` scheme (`data:text/html`, etc.) stays inert text and is never embedded; the sanitizer permits `data:` for `<img>` only, and the streaming renderer enforces the same predicate so an image can't render mid-stream then vanish at settle. Base64 SVG renders via `<img src>` (scripts can't execute in that context). Thanks @ai-ag2026. (#6209)
- **A single configured model no longer appears multiple times in the model picker.** A configured model was registered under its bare id, `provider/model`, AND `@provider:model` forms, so the picker treated equivalent routing ids as separate entries and showed the same model more than once. Dedup now collapses only genuinely-equivalent entries (same normalized id **and** same provider), so two different providers offering the same bare model id (e.g. both offer `gpt-4o`) stay distinct while a truly-one-model-registered-three-ways case collapses to one badge — applied to both the live and static catalog builders. A picked provider badge resolves, sends, and persists as that provider (no snap-back to a first-match provider on a bare-id collision), and a colon-bearing model id (e.g. `model-a:free`) synthesized as a missing-catalog fallback reads its authoritative `data-provider` rather than mis-parsing at the last colon. Thanks @happy5318. (#6221, #3691)
- **Reopening a WebUI tab after a large active session is fast again, even when the session has few messages but large tool outputs.** The reconnect display-path tail optimization only fired based on message *count*, so a session with a handful of messages but multi-MB tool-call JSON forced a slow full-scan merge (bootstrap could take tens of seconds). The optimization now also fires when the session sidecar file exceeds 500 KB, regardless of message count — the same messages are returned (verified: identical message set, count, offset, and truncation metadata), just via the fast path. Thanks @webtecnica. (#6260)
- **OIDC allowlist values containing spaces (e.g. a group named `Hermes Users`) are no longer split into separate entries.** The allowlist parser split comma-separated values on all whitespace, so a multi-word group/user name became two spurious entries. Allowlist parsing now splits only on commas/newlines (multi-word values stay intact and match their exact signed claim — strictly tightening, no security loosening), while OAuth **scope** parsing keeps its space-delimited behavior per RFC 6749 §3.3. A blank allowlist entry no longer bricks an OIDC-only deployment. Thanks @webtecnica. (#6244)
### Changed
- **Internal: removed a dead `rowIndex` parameter from the settled-scene row projector.** `attachLiveStream`'s `pushRow` helper carried an unused positional index left over from before final-segment eligibility moved to a row-identity `WeakSet`; dropping it is a pure no-op refactor (differential execution confirms byte-identical settled scenes — ordering, dedupe, final-prefix suppression, and sequence assignment all unchanged). Thanks @webtecnica. (#6258)
## [exp-v0.52.157] — 2026-07-29 — Experimental release (reconnect transcript ordering)
### Fixed
+44
View File
@@ -30,6 +30,50 @@ RUN apt-get update -y --fix-missing --no-install-recommends \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
# ── SQLite upgrade ──────────────────────────────────────────────────────────
# The python:3.12-slim base ships SQLite 3.46.1 (Debian Trixie), which is
# vulnerable to the WAL-reset corruption bug discovered March 2026.
# https://sqlite.org/wal.html#walresetbug
#
# Debian has not backported the fix, so we compile from the amalgamation.
# Installs to /usr/local/lib (registered in ld.so.conf.d for arm64 priority).
# Build tools are purged after compilation to keep the image lean.
# Build args are for forward version bumps only (3.54+, etc.).
# When bumping SQLITE_VERSION, recompute the SHA-256 from the official
# download and update SQLITE_SHA256 accordingly.
ARG SQLITE_VERSION=3530000
ARG SQLITE_YEAR=2026
ARG SQLITE_SHA256=851e9b38192fe2ceaa65e0baa665e7fa06230c3d9bd1a6a9662d02380d73365a
RUN apt-get update \
&& apt-get install -y --no-install-recommends gcc make libc6-dev \
&& cd /tmp \
&& curl -fsSL "https://sqlite.org/${SQLITE_YEAR}/sqlite-autoconf-${SQLITE_VERSION}.tar.gz" \
-o sqlite.tar.gz \
&& echo "${SQLITE_SHA256} sqlite.tar.gz" | sha256sum -c - \
&& tar xzf sqlite.tar.gz \
&& cd "sqlite-autoconf-${SQLITE_VERSION}" \
&& CPPFLAGS="-DSQLITE_SECURE_DELETE" ./configure --prefix=/usr/local --disable-static --disable-readline \
--enable-fts5 --enable-fts4 --enable-rtree \
&& make -j"$(nproc)" \
&& make install \
&& echo "/usr/local/lib" > /etc/ld.so.conf.d/000-usr-local-lib.conf \
&& /sbin/ldconfig \
&& cd / && rm -rf /tmp/sqlite* \
&& apt-get purge -y gcc make libc6-dev \
&& apt-get autoremove -y \
&& apt-get clean && rm -rf /var/lib/apt/lists/* \
&& python3 -c "\
import sqlite3; \
v = sqlite3.sqlite_version; \
assert tuple(int(x) for x in v.split('.')) >= (3, 51, 3), \
f'SQLite {v} still vulnerable'; \
c = sqlite3.connect(':memory:'); \
assert c.execute('PRAGMA secure_delete').fetchone()[0] == 1, \
'SQLITE_SECURE_DELETE not compiled in (deleted rows would remain recoverable)'; \
c.execute('CREATE VIRTUAL TABLE _fts5_build_check USING fts5(x)'); \
c.execute('DROP TABLE _fts5_build_check'); \
c.close()"
# Optional GPU user-space acceleration libraries for users who pass through
# host GPU devices. The default image remains CPU-only.
ARG INSTALL_GPU_LIBS=0
+1
View File
@@ -692,6 +692,7 @@ The WebUI is still coupled to Hermes Agent internals for runtime execution, prov
- [`TESTING.md`](TESTING.md) — manual browser test plan and automated coverage reference
- [`DESIGN.md`](DESIGN.md) — design tokens and the calm-console direction
- [`docs/UIUX-GUIDE.md`](docs/UIUX-GUIDE.md) — UI/UX principles sourced from the design docs and visual inventories
- [`docs/sse-streams.md`](docs/sse-streams.md) — cross-client SSE endpoint reference: session streaming, gateway SSE probe scope, heartbeats, and proxy behavior
- [`docs/CONTRACTS.md`](docs/CONTRACTS.md) — project contract/RFC/design index for contributors and agents
- [`docs/rfcs/README.md`](docs/rfcs/README.md) — RFC index for larger architecture and durability proposals
+278 -6
View File
@@ -8,19 +8,54 @@ mixed runtime and require a clean WebUI restart instead.
from __future__ import annotations
import errno
import math
import os
from pathlib import Path
import stat
import sys
import subprocess
import threading
import time
# Retain the discovered path as a diagnostic/test-visible compatibility value;
# runtime identity is deliberately captured from the loaded module below.
from api.config import _AGENT_DIR # noqa: F401
_RESTART_MESSAGE = (
"Hermes Agent was updated while Hermes WebUI was running. "
"Restart Hermes WebUI before retrying this action."
from api.config import (
PYTHON_EXE,
_AGENT_DIR, # noqa: F401
_DEFAULT_STATE_HOME,
)
from api.subprocess_utils import windows_hide_flags
_RESTART_REQUIRED_MESSAGE = (
"Hermes Agent was updated while Hermes WebUI was running. "
"WebUI cannot verify that the Agent update completed safely. "
"Check the Agent update outcome and environment first. "
"Restart Hermes WebUI manually before retrying this action."
)
_AGENT_UPDATE_MARKER = ".hermes-update-in-progress"
_AGENT_RECOVERY_MARKERS = (".update-incomplete", ".lazy-refresh-incomplete")
_AGENT_UPDATE_MAX_AGE_SECONDS = 20 * 60
# The update marker holds a PID and a start timestamp (two short numeric lines).
# Anything larger is not a legitimate marker; cap the read so a huge or growing
# regular file can never exhaust memory on the stale-runtime request path.
_AGENT_UPDATE_MARKER_MAX_BYTES = 64 * 1024
# O_NOFOLLOW is POSIX; on platforms that lack it the fast os.open() path is not
# taken at all (see _MARKER_SAFE_OPEN_AVAILABLE below).
# The marker read hardening relies on two POSIX-only open flags to stay both
# non-blocking (never hang on a FIFO/device) and symlink-safe. O_NONBLOCK is
# Unix-only and O_NOFOLLOW is absent on some platforms; accessing them
# unconditionally raises AttributeError on native Windows. Resolve them safely
# and only take the os.open() fast path when BOTH are genuinely available —
# otherwise the read cannot prove non-blocking + no-follow and must fall back to
# an lstat-only classification (see _read_live_agent_update).
_O_NOFOLLOW = getattr(os, "O_NOFOLLOW", 0)
_O_NONBLOCK = getattr(os, "O_NONBLOCK", 0)
_MARKER_SAFE_OPEN_AVAILABLE = bool(getattr(os, "O_NOFOLLOW", 0)) and bool(
getattr(os, "O_NONBLOCK", 0)
)
_HERMES_HOME = Path(_DEFAULT_STATE_HOME)
_AGENT_PYTHON = Path(PYTHON_EXE).expanduser() if PYTHON_EXE else None
def _read_agent_revision(
@@ -49,6 +84,7 @@ def _read_agent_revision(
capture_output=True,
text=True,
timeout=2,
creationflags=windows_hide_flags(),
)
if worktree_result.returncode != 0:
return None
@@ -69,6 +105,7 @@ def _read_agent_revision(
capture_output=True,
text=True,
timeout=2,
creationflags=windows_hide_flags(),
)
if tracked_result.returncode != 0:
return None
@@ -78,6 +115,7 @@ def _read_agent_revision(
capture_output=True,
text=True,
timeout=2,
creationflags=windows_hide_flags(),
)
except (OSError, subprocess.TimeoutExpired, RuntimeError, ValueError):
return None
@@ -96,6 +134,233 @@ _RUNTIME_LOCK = threading.Lock()
class AgentRuntimeChangedError(RuntimeError):
"""Raised when the loaded Agent runtime no longer matches its source tree."""
def __init__(
self,
message: str,
*,
agent_update_state: str | None = None,
) -> None:
super().__init__(message)
self.agent_update_state = agent_update_state
def agent_runtime_stale_payload(exc: AgentRuntimeChangedError) -> dict:
"""Return the shared retry response for every stale-runtime entry point."""
payload = {
"error": str(exc),
"type": "agent_runtime_stale",
"retryable": True,
"restart_scheduled": False,
}
if exc.agent_update_state is not None:
payload["agent_update_state"] = exc.agent_update_state
return payload
def _pid_is_alive(pid: int) -> bool | None:
"""Return PID liveness, or ``None`` when the platform cannot confirm it."""
if pid <= 0:
return False
if pid.bit_length() > 32:
return None
if sys.platform == "win32":
try:
import ctypes
from ctypes import wintypes
kernel32 = ctypes.WinDLL("kernel32", use_last_error=True)
kernel32.OpenProcess.argtypes = (
wintypes.DWORD,
wintypes.BOOL,
wintypes.DWORD,
)
kernel32.OpenProcess.restype = wintypes.HANDLE
kernel32.GetExitCodeProcess.argtypes = (
wintypes.HANDLE,
ctypes.POINTER(wintypes.DWORD),
)
kernel32.GetExitCodeProcess.restype = wintypes.BOOL
kernel32.CloseHandle.argtypes = (wintypes.HANDLE,)
kernel32.CloseHandle.restype = wintypes.BOOL
handle = kernel32.OpenProcess(0x1000, False, pid)
if not handle:
error = ctypes.get_last_error()
if error == 5: # ERROR_ACCESS_DENIED still proves the PID exists.
return True
if error == 87: # ERROR_INVALID_PARAMETER for a missing PID.
return False
return None
try:
exit_code = wintypes.DWORD()
if not kernel32.GetExitCodeProcess(handle, ctypes.byref(exit_code)):
return None
return exit_code.value == 259 # STILL_ACTIVE
finally:
kernel32.CloseHandle(handle)
except (AttributeError, OSError, TypeError, ValueError):
return None
try:
os.kill(pid, 0)
except ProcessLookupError:
return False
except PermissionError:
return True
except (OverflowError, ValueError):
return None
except OSError as exc:
if exc.errno == errno.ESRCH:
return False
if exc.errno == errno.EPERM:
return True
return None
return True
def _read_live_agent_update(marker: Path) -> str:
"""Classify the shared Agent update marker without changing Agent state.
The marker is attacker-adjacent shared state (any process that can write the
Agent home can create it), so the read is hardened: never follow a symlink,
never block on a FIFO/device, and never read an unbounded regular file.
Anything that is not a small regular file is classified ``unknown`` rather
than allowed to hang or exhaust memory on a stale-runtime request path.
"""
if not _MARKER_SAFE_OPEN_AVAILABLE:
# Without both O_NONBLOCK and O_NOFOLLOW we cannot prove the read is
# non-blocking and symlink-safe (e.g. native Windows), so never open the
# marker: an unverifiable marker fails closed to ``unknown``, and only a
# genuinely missing path is ``absent``.
try:
marker.lstat()
except FileNotFoundError:
return "absent"
except (OSError, ValueError, TypeError):
return "unknown"
return "unknown"
try:
fd = os.open(marker, os.O_RDONLY | _O_NONBLOCK | _O_NOFOLLOW)
except FileNotFoundError:
try:
marker.lstat()
except FileNotFoundError:
return "absent"
except OSError:
return "unknown"
# Path exists to lstat (e.g. a dangling/looping symlink) but O_NOFOLLOW
# refused to open it — treat as an unverifiable marker.
return "unknown"
except (OSError, ValueError, TypeError):
# ELOOP (symlink under O_NOFOLLOW), ENXIO/EWOULDBLOCK (FIFO with no
# writer under O_NONBLOCK), a non-path marker object, or any other open
# failure — fail closed: an unreadable marker is never proof of safety.
return "unknown"
try:
try:
st = os.fstat(fd)
except OSError:
return "unknown"
if not stat.S_ISREG(st.st_mode):
# FIFO, device, directory, socket — never a legitimate marker.
return "unknown"
if st.st_size > _AGENT_UPDATE_MARKER_MAX_BYTES:
return "unknown"
try:
# Read one byte past the cap so an oversized file that lied about
# st_size (or grew mid-read) is still rejected rather than truncated.
data = os.read(fd, _AGENT_UPDATE_MARKER_MAX_BYTES + 1)
except (OSError, BlockingIOError):
return "unknown"
finally:
try:
os.close(fd)
except OSError:
pass
if len(data) > _AGENT_UPDATE_MARKER_MAX_BYTES:
return "unknown"
try:
raw = data.decode("utf-8")
except UnicodeError:
return "unknown"
lines = raw.splitlines()
try:
pid = int(lines[0].strip())
started_at = float(lines[1].strip())
except (IndexError, TypeError, ValueError):
return "unknown"
if pid <= 0 or not math.isfinite(started_at):
return "unknown"
age_seconds = time.time() - started_at
if age_seconds < 0:
return "unknown"
if age_seconds > _AGENT_UPDATE_MAX_AGE_SECONDS:
return "stale"
alive = _pid_is_alive(pid)
if alive is None:
return "unknown"
return "active" if alive else "stale"
def _marker_presence(marker: Path) -> str:
"""Return ``present``, ``absent``, or ``unknown`` for a recovery marker."""
try:
marker.lstat()
except FileNotFoundError:
return "absent"
except OSError:
return "unknown"
return "present"
def _agent_install_roots() -> tuple[Path, ...]:
"""Return portable roots that can own the Agent's venv recovery markers."""
candidates: list[Path] = []
if _AGENT_SOURCE_DIR is not None:
candidates.append(_AGENT_SOURCE_DIR)
if _AGENT_PYTHON is not None:
# A venv Python is commonly a symlink to a shared interpreter. Keep the
# configured venv path so its installation's recovery markers are read.
python_path = _AGENT_PYTHON
if python_path.parent.name.lower() in {"bin", "scripts"}:
venv_dir = python_path.parent.parent
if venv_dir.name.lower() in {"venv", ".venv"}:
candidates.append(venv_dir.parent)
roots: list[Path] = []
seen: set[str] = set()
for candidate in candidates:
key = os.path.normcase(os.path.abspath(str(candidate)))
if key not in seen:
seen.add(key)
roots.append(candidate)
return tuple(roots)
def _agent_update_transaction_state() -> str:
"""Report marker diagnostics, never proof of successful completion.
The Agent removes its active marker on failed/interrupted exits too. Neither
its absence nor a stale PID proves the checkout or environment is healthy.
"""
live_state = _read_live_agent_update(_HERMES_HOME / _AGENT_UPDATE_MARKER)
if live_state == "unknown":
return "unknown"
recovery_present = False
for root in _agent_install_roots():
for marker_name in _AGENT_RECOVERY_MARKERS:
presence = _marker_presence(root / marker_name)
if presence == "unknown":
return "unknown"
recovery_present = recovery_present or presence == "present"
if recovery_present:
return "incomplete"
return "unverified" if live_state == "absent" else live_state
def _loaded_agent_source_identity() -> tuple[Path, Path] | None:
"""Return the source directory and file that supplied ``run_agent``."""
@@ -136,7 +401,14 @@ def ensure_agent_runtime_current() -> None:
_read_agent_revision(_AGENT_SOURCE_DIR, module_path=_AGENT_MODULE_PATH)
!= _AGENT_REVISION
):
raise AgentRuntimeChangedError(_RESTART_MESSAGE)
# Automatic restart needs an Agent-owned success receipt bound to this
# transaction, final revision and healthy environment, plus an atomic
# handoff excluding mutations across replacement. Marker polling and a
# final revision read supply neither contract. Keep this path manual.
raise AgentRuntimeChangedError(
_RESTART_REQUIRED_MESSAGE,
agent_update_state=_agent_update_transaction_state(),
)
def require_ai_agent_class():
+72 -7
View File
@@ -6,6 +6,17 @@ from pathlib import Path
logger = logging.getLogger(__name__)
# state.db paths that already produced the "no 'source' column" warning below.
# ``get_cli_sessions(all_profiles=True)`` re-reads every profile DB on every
# sidebar poll (behind a 5 s cache), so a single pre-``source`` profile DB would
# otherwise re-emit the identical WARNING line every ~15 s for the life of the
# process. The condition is a property of the DB file, not of the poll, so it
# is reported once per path. Process-lifetime only: a restart warns again,
# which is the desired behaviour (the log line is the operator's cue that the
# agent still needs upgrading). Plain ``set`` mutation under the GIL is
# sufficient here; a duplicate line from two racing first calls is harmless.
_SOURCE_COLUMN_WARNED_DB_PATHS: set[str] = set()
def open_state_db_readonly(db_path: Path, log: logging.Logger | None = None) -> sqlite3.Connection:
"""Open the live agent ``state.db`` read-only for a pure-read projection.
@@ -48,6 +59,7 @@ MESSAGING_SOURCES = {
'telegram',
'weixin',
'matrix',
'signal',
}
CLI_MIN_UNTITLED_MESSAGE_COUNT = 6
@@ -71,6 +83,7 @@ SOURCE_LABELS = {
'webui': 'WebUI',
'weixin': 'Weixin',
'matrix': 'Matrix',
'signal': 'Signal',
}
@@ -517,6 +530,19 @@ def read_importable_agent_session_rows(
``exclude_sources=None``. ``include_sources`` is an additional narrowing
filter; callers that want an include-only query should explicitly pass
``exclude_sources=None`` so the default exclusions do not also apply.
``limit`` bounds the *recency slice*, not the returned row count. Subagent
rows only render as children when their parent row is in the same payload,
so subagent ancestors of selected rows are re-added afterwards and the
result can exceed ``limit`` by the number of such anchors. Callers must
therefore iterate the result rather than assume ``len(rows) <= limit``.
That recovery is deliberately bounded by the oversampled candidate set
(``limit * 8`` newest sessions): it re-uses rows the projection already
fetched and never issues an extra query, so an ancestor older than the
oversample stays unresolved and its children render top-level, exactly as
they did before. Widening that window is a ``candidate_limit`` change, not
a change to this walk.
"""
db_path = Path(db_path)
if not db_path.exists():
@@ -549,12 +575,15 @@ def read_importable_agent_session_rows(
cur.execute("PRAGMA table_info(messages)")
message_cols = {row[1] for row in cur.fetchall()}
if 'source' not in session_cols:
log.warning(
"agent session listing skipped: state.db at %s has no 'source' column "
"(older hermes-agent?). Agent sessions unavailable. "
"Upgrade hermes-agent to fix this.",
db_path,
)
warned_key = str(db_path.resolve())
if warned_key not in _SOURCE_COLUMN_WARNED_DB_PATHS:
_SOURCE_COLUMN_WARNED_DB_PATHS.add(warned_key)
log.warning(
"agent session listing skipped: state.db at %s has no 'source' column "
"(older hermes-agent?). Agent sessions unavailable. "
"Upgrade hermes-agent to fix this.",
db_path,
)
return []
parent_expr = _optional_col('parent_session_id', session_cols)
@@ -766,7 +795,43 @@ def read_importable_agent_session_rows(
projected = [row for row in projected if is_cli_session_row_visible(row)]
if limit is None:
return projected
return projected[:max(0, int(limit))]
selected = projected[:max(0, int(limit))]
# The recency slice is per-row, but subagent rows are only renderable as
# children: the sidebar nests a child under its parent solely when that
# parent row is present in the same payload. A frozen orchestrator stops
# writing while its leaves keep streaming, so the leaves win the recency
# race and the parent falls outside the window — leaving the leaves to be
# promoted to top-level sidebar rows. Re-add subagent parents that the
# oversampled candidate set already projected (no extra query); webui
# ancestors are left out because that sidebar bucket already has them.
#
# Bounded by construction: ``by_id`` only holds the ``limit * 8``
# newest candidates, so an ancestor older than that oversample is not
# recovered and its children stay top-level — the pre-existing
# behaviour, narrowed rather than fixed. Resolving those would need an
# unbounded per-row ancestor query on the hot sidebar path; widen
# ``candidate_limit`` instead if the window proves too tight.
# NOTE: this can return more than ``limit`` rows (see docstring).
have = {row.get('id') for row in selected}
by_id = {row.get('id'): row for row in projected if row.get('id')}
pending = list(selected)
while pending:
row = pending.pop()
if str(row.get('raw_source') or row.get('source') or '').strip().lower() != 'subagent':
continue
parent_id = row.get('parent_session_id')
if not parent_id or parent_id in have:
continue
parent = by_id.get(parent_id)
if parent is None:
continue
if str(parent.get('raw_source') or parent.get('source') or '').strip().lower() != 'subagent':
continue
selected.append(parent)
have.add(parent_id)
pending.append(parent)
return selected
+20 -1
View File
@@ -105,6 +105,12 @@ _TRUSTED_AUTH_HEADER_ENV = 'HERMES_WEBUI_TRUSTED_AUTH_HEADER'
_TRUSTED_GROUPS_HEADER_ENV = 'HERMES_WEBUI_TRUSTED_GROUPS_HEADER'
_TRUSTED_GROUP_PROFILE_MAP_ENV = 'HERMES_WEBUI_GROUP_PROFILE_MAP'
_TRUSTED_AUTH_LOGOUT_URL_ENV = 'HERMES_WEBUI_TRUSTED_AUTH_LOGOUT_URL'
# Opt-in: also treat '|' as a group separator in the trusted-groups header.
# Off by default so an existing deployment whose group NAME legitimately
# contains a literal '|' is never silently re-split into two groups (which
# could change its profile binding). Set to 1/true/yes/on for identity
# providers (some Authentik outpost configs) that emit "admins|developpeur".
_TRUSTED_GROUPS_PIPE_SEPARATOR_ENV = 'HERMES_WEBUI_TRUSTED_GROUPS_PIPE_SEPARATOR'
_TRUSTED_AUTH_WARNINGS_EMITTED: set[str] = set()
@@ -722,7 +728,20 @@ def _trusted_groups_header_value(handler) -> list[str]:
if not raw:
return []
values = []
for part in str(raw).replace('\n', ',').split(','):
# Authentik's outpost typically joins multiple group names with a comma or
# newline; parse those as separators by default. Some proxy provider /
# property-mapping configs instead emit a pipe-separated list (e.g.
# "admins|developpeur"), which a comma-only split would treat as one
# unmatched group name — silently dropping the session to the unbound
# "default" profile despite a legitimate mapped membership. Pipe splitting
# is therefore available but OPT-IN (HERMES_WEBUI_TRUSTED_GROUPS_PIPE_SEPARATOR),
# because a group NAME can legitimately contain a literal '|' and must not be
# re-split by default — doing so unconditionally could change an existing
# deployment's profile binding.
normalized = str(raw).replace('\n', ',')
if str(os.getenv(_TRUSTED_GROUPS_PIPE_SEPARATOR_ENV, '')).strip().lower() in ('1', 'true', 'yes', 'on'):
normalized = normalized.replace('|', ',')
for part in normalized.split(','):
part = part.strip()
if part:
values.append(part)
+79 -28
View File
@@ -271,6 +271,77 @@ def subscribe_to_session_channel(
return ch, q
# Bounded window a cancelling worker may stay lifecycle-busy before an entry
# with no live SSE channel is treated as an orphan. Matches the unwind ceiling
# used by the chat-start successor guard in ``api.routes``.
_ACTIVE_RUN_CANCEL_UNWIND_SECONDS = 180.0
def _active_run_ids_for_session(
session_id: str,
*,
attachable_only: bool,
) -> list[str]:
"""Return this session's run stream ids under the requested semantics.
``attachable_only`` selects browser-recovery semantics and drops every
``phase="cancelling"`` row: cancellation is already terminal for the client
even while the worker unwinds. Busy checks pass ``False`` so a freshly
cancelled run still blocks a successor for its bounded unwind window.
A cancelling row past that window with no live ``STREAMS`` channel is an
orphan: it is dropped from ``ACTIVE_RUNS`` and its stream-owner entry is
released, so a wedged worker cannot suppress background wakeups forever.
"""
from api import config as _cfg
sid = str(session_id or "").strip()
if not sid:
return []
try:
with _cfg.STREAMS_LOCK:
live_stream_ids = set((_cfg.STREAMS or {}).keys())
now = time.time()
matches: list[str] = []
stale_keys: list[str] = []
with _cfg.ACTIVE_RUNS_LOCK:
for run_key, meta in list((_cfg.ACTIVE_RUNS or {}).items()):
if not isinstance(meta, dict) or meta.get("session_id") != sid:
continue
stream_id = str(meta.get("stream_id") or run_key or "").strip()
if not stream_id:
continue
cancelling = not _cfg.active_run_is_attachable(meta)
stale_cancel = (
cancelling
and _cfg.active_run_cancel_is_stale(
meta,
grace_seconds=_ACTIVE_RUN_CANCEL_UNWIND_SECONDS,
now=now,
)
and run_key not in live_stream_ids
and stream_id not in live_stream_ids
)
if stale_cancel:
stale_keys.append(run_key)
continue
if attachable_only and cancelling:
continue
matches.append(stream_id)
for run_key in stale_keys:
(_cfg.ACTIVE_RUNS or {}).pop(run_key, None)
for run_key in stale_keys:
_cfg.unregister_stream_owner(run_key)
return matches
except Exception:
logger.debug(
"ACTIVE_RUNS lookup failed for %s",
sid,
exc_info=True,
)
return []
def active_stream_id_for_session(session_id: str) -> Optional[str]:
"""Return the stream_id of the live run for *session_id*, or None.
@@ -288,25 +359,14 @@ def active_stream_id_for_session(session_id: str) -> Optional[str]:
Keys on ACTIVE_RUNS (worker-lifecycle registry) — the same source
``_session_has_active_turn`` / ``_emit_to_session_streams`` already trust
to map a stream back to its owning session. Returns the first matching
stream_id (a session has at most one live run; cancel/reconnect can
briefly hold two — either is a valid attach target, the frontend dedupes
by stream_id).
to map a stream back to its owning session — but returns only rows that
are still ATTACHABLE. A ``phase="cancelling"`` row stays lifecycle-busy
while its worker unwinds, yet the client already reached a terminal state
for that run, so replaying ``server_turn_started`` for it makes the tab
attach, receive the terminal event, resubscribe, and loop forever.
"""
from api import config as _cfg
try:
with _cfg.ACTIVE_RUNS_LOCK:
for _stream_id, meta in (_cfg.ACTIVE_RUNS or {}).items():
if isinstance(meta, dict) and meta.get("session_id") == session_id:
return str(_stream_id)
except Exception:
logger.debug(
"active_stream_id_for_session lookup failed for %s",
session_id,
exc_info=True,
)
return None
matches = _active_run_ids_for_session(session_id, attachable_only=True)
return matches[0] if matches else None
def persisted_message_count_for_session(session_id: str) -> Optional[int]:
@@ -1518,16 +1578,7 @@ def _session_has_active_turn(session_id: str) -> bool:
``_start_chat_stream_for_session``'s own active-stream guard — i.e. the
same lock /api/chat/start uses is the authoritative race backstop.
"""
from api import config as _cfg
try:
with _cfg.ACTIVE_RUNS_LOCK:
for _stream_id, meta in (_cfg.ACTIVE_RUNS or {}).items():
if isinstance(meta, dict) and meta.get("session_id") == session_id:
return True
except Exception:
logger.debug("ACTIVE_RUNS active-turn check failed", exc_info=True)
return False
return bool(_active_run_ids_for_session(session_id, attachable_only=False))
def _start_server_side_wakeup_turn(
+8 -3
View File
@@ -86,11 +86,16 @@ def clear_pending(session_key: str) -> int:
def _with_timeout_metadata(data: dict) -> dict:
item = dict(data or {})
requested_at = float(item.get("requested_at") or time.time())
timeout_seconds = int(item.get("timeout_seconds") or DEFAULT_TIMEOUT_SECONDS)
expires_at = float(item.get("expires_at") or requested_at + timeout_seconds)
raw_timeout = item.get("timeout_seconds")
timeout_seconds = int(raw_timeout) if raw_timeout is not None else DEFAULT_TIMEOUT_SECONDS
item["requested_at"] = requested_at
item["timeout_seconds"] = timeout_seconds
item["expires_at"] = expires_at
if timeout_seconds <= 0:
# Unlimited: no expiry. The frontend renders no countdown and the card
# waits until the user answers (or the run is cancelled).
item["expires_at"] = 0
return item
item["expires_at"] = float(item.get("expires_at") or requested_at + timeout_seconds)
return item
+674 -81
View File
File diff suppressed because it is too large Load Diff
+30 -2
View File
@@ -1854,7 +1854,7 @@ def install_extension(id: object, download_url: object, sha256: object) -> Dict[
ext_dir_resolved = ext_dir.resolve()
for member_name in file_members:
decoded = _fully_unquote_path(_stripped(member_name))
if not decoded or not _is_safe_relative_path(decoded):
if not decoded or not _is_safe_archive_member(decoded):
raise ExtensionInstallError("Unsafe archive member")
resolved = (ext_dir / decoded).resolve()
try:
@@ -1956,7 +1956,7 @@ def uninstall_extension(id: object) -> Dict[str, Any]:
raise ExtensionInstallError("Extension not installed", 404)
ext_dir = root / ext_id
for rel_path in entry.get("files", []):
if not _is_safe_relative_path(rel_path):
if not _is_safe_archive_member(rel_path):
continue
target = (ext_dir / rel_path).resolve()
try:
@@ -2068,6 +2068,9 @@ def inject_extension_tags(index_html: str) -> str:
def _is_safe_relative_path(rel: str) -> bool:
# Strict: reject empty, traversal, AND any dot-prefixed segment. This is shared
# by static serving, asset URLs and manifest paths, where a hidden file must
# never become reachable. Archive members use _is_safe_archive_member() below.
if not rel or "\x00" in rel or "\\" in rel:
return False
for segment in rel.split("/"):
@@ -2076,6 +2079,31 @@ def _is_safe_relative_path(rel: str) -> bool:
return True
# Benign hidden LEAF files that extension archives legitimately ship. A dot-prefixed
# name is permitted only as the final path segment and only if it is one of these;
# dot-directories (e.g. ".git/") and every other dotfile (".env", ".secret") stay rejected.
_ALLOWED_ARCHIVE_DOTFILES = frozenset({".gitkeep", ".gitignore", ".gitattributes", ".env.example"})
def _is_safe_archive_member(rel: str) -> bool:
"""Path validator for archive install/uninstall ONLY. Same containment as
_is_safe_relative_path(), but permits a narrow allowlist of benign hidden leaf
files so extensions can ship ".gitkeep"/".env.example" without loosening the
shared validator that also gates static serving."""
if not rel or "\x00" in rel or "\\" in rel:
return False
segments = rel.split("/")
for segment in segments[:-1]:
if not segment or segment in (".", "..") or segment.startswith("."):
return False
leaf = segments[-1]
if not leaf or leaf in (".", ".."):
return False
if leaf.startswith(".") and leaf not in _ALLOWED_ARCHIVE_DOTFILES:
return False
return True
def _not_found(handler) -> bool:
j(handler, {"error": "not found"}, status=404)
return True
+100 -12
View File
@@ -29,6 +29,7 @@ from api.config import (
coerce_reasoning_effort_for_model,
gateway_approval_unavailable_reason,
gateway_supports_approval,
peek_stream,
register_active_run,
unregister_active_run,
unregister_stream_owner,
@@ -327,6 +328,73 @@ def _gateway_reasoning_effort_for_request(cfg, *, model=None, model_provider=Non
return None
def _gateway_session_yolo_enabled(session_id: str) -> bool:
"""Return the WebUI-owned, in-memory YOLO state for a browser session."""
try:
from tools.approval import is_session_yolo_enabled
return bool(is_session_yolo_enabled(str(session_id or "")))
except Exception:
return False
def _settle_gateway_run_approval(
session_id: str,
approval_data: dict,
base_url: str,
api_key: str,
) -> tuple[bool, dict | None, int]:
"""Auto-approve or mirror one run approval at a session-linearized point."""
from api.route_approvals import gateway_yolo_handoff, submit_gateway_pending_mirror
run_id = str(approval_data.get("run_id") or "").strip()
identity_v1 = bool(approval_data.get("_gateway_agent_identity_v1"))
with gateway_yolo_handoff(session_id):
if _gateway_session_yolo_enabled(session_id):
try:
_auto_approve_gateway_run(
base_url,
api_key,
run_id,
approval_data["approval_id"] if identity_v1 else "",
)
return True, None, 0
except Exception:
# Fail closed: if remote approval fails, surface the real card
# before allowing a same-session toggle to pass the handoff.
logger.warning(
"WebUI YOLO could not auto-approve run %s; showing approval card",
run_id,
exc_info=True,
)
head, total = submit_gateway_pending_mirror(session_id, approval_data)
return False, head, total
def _auto_approve_gateway_run(
base_url: str,
api_key: str,
run_id: str,
approval_id: str,
) -> None:
"""Resolve one Runs API prompt using only the shipped approval contract.
This is a WebUI-owned compatibility path: the Runs API does not yet expose
session YOLO, so WebUI answers each approval request while its own session
flag is enabled. Native Agent-side YOLO would be preferable because it can
bypass gates before they pause and also covers Agent-owned computer-use
policy; https://github.com/NousResearch/hermes-agent/pull/61946 tracks that
API capability. Until then, do not send speculative fields to the Agent.
"""
from api.runner_client import HttpRunnerClient
HttpRunnerClient(base_url=base_url, api_key=api_key).respond_approval(
run_id,
approval_id,
"once",
)
def gateway_chat_config_status(config_data=None, environ: dict[str, str] | None = None) -> dict:
"""Return redacted Gateway-backed chat configuration status."""
mode = webui_chat_backend_mode(config_data, environ)
@@ -516,7 +584,7 @@ def _run_gateway_runs_api_streaming(
try:
from api.streaming import _build_native_multimodal_message
message_content = _build_native_multimodal_message("", str(msg_text or ""), attachments, str(workspace), cfg=cfg, active_provider=active_provider, active_model=(model or ""), requested_provider=active_provider)
message_content = _build_native_multimodal_message("", str(msg_text or ""), attachments, str(workspace), cfg=cfg, active_provider=active_provider, active_model=(model or ""), requested_provider=active_provider, profile=getattr(session, "profile", None))
except Exception:
logger.debug("Failed to build runs-API multimodal attachment payload", exc_info=True)
message_content = str(msg_text or "")
@@ -615,9 +683,17 @@ def _run_gateway_runs_api_streaming(
if approval_data:
approval_data["run_id"] = run_id
from api.config import gateway_supports_approval_identity_v1
approval_data["_gateway_agent_identity_v1"] = bool(approval_data.get("_gateway_raw_approval_id_present")) and gateway_supports_approval_identity_v1(base_url, api_key)
from api.route_approvals import submit_gateway_pending_mirror
head, total = submit_gateway_pending_mirror(session_id, approval_data)
identity_v1 = bool(approval_data.get("_gateway_raw_approval_id_present")) and gateway_supports_approval_identity_v1(base_url, api_key)
approval_data["_gateway_agent_identity_v1"] = identity_v1
auto_approved, head, total = _settle_gateway_run_approval(
session_id,
approval_data,
base_url,
api_key,
)
if auto_approved:
sse_event = "message"
continue
put_gateway_event("approval", {**(head or approval_data), "pending_count": total})
sse_event = "message"
continue
@@ -834,6 +910,7 @@ def _run_gateway_chat_streaming(
*,
model_provider=None,
goal_related=False,
regeneration=False,
):
"""Bridge a WebUI chat turn through Hermes Gateway's API server.
@@ -843,7 +920,7 @@ def _run_gateway_chat_streaming(
the configured Gateway API server into those local events and persists the
final user/assistant turn back into the WebUI session.
"""
q = STREAMS.get(stream_id)
q = peek_stream(stream_id)
if q is None:
_finish_gateway_run_starting(stream_id, result="fallback")
_clear_gateway_run_starting(stream_id)
@@ -1043,7 +1120,7 @@ def _run_gateway_chat_streaming(
try:
from api.streaming import _build_native_multimodal_message
message_content = _build_native_multimodal_message("", str(msg_text or ""), attachments, str(workspace), cfg=cfg, active_provider=(model_provider or ""), active_model=(model or ""), requested_provider=(model_provider or ""))
message_content = _build_native_multimodal_message("", str(msg_text or ""), attachments, str(workspace), cfg=cfg, active_provider=(model_provider or ""), active_model=(model or ""), requested_provider=(model_provider or ""), profile=getattr(s, "profile", None))
except Exception:
logger.debug("Failed to build gateway multimodal attachment payload", exc_info=True)
message_content = str(msg_text or "")
@@ -1204,18 +1281,29 @@ def _run_gateway_chat_streaming(
# same sort key; later transcript merges can then fall back to
# role/content ordering instead of turn order.
assistant_ts = now + 0.000001
user_msg = {"role": "user", "content": str(msg_text or ""), "timestamp": now}
pending_source = getattr(s, "pending_user_source", None) or "webui"
if pending_source != "webui":
user_msg["_source"] = pending_source
if attachments:
user_msg["attachments"] = list(attachments)
from api.streaming import _active_turn_authority, _materialize_active_turn_user
active_turn_identity = _active_turn_authority(s, stream_id, msg_text)
user_msg = _materialize_active_turn_user(
active_turn_identity,
str(msg_text or ""),
pending_source,
)
user_msg["timestamp"] = float(
active_turn_identity.get("timestamp") or now
)
assistant_msg = {"role": "assistant", "content": assistant_text, "timestamp": assistant_ts}
saved_reasoning = STREAM_REASONING_TEXT.get(stream_id, "")
if saved_reasoning:
assistant_msg["reasoning"] = saved_reasoning
previous_messages = list(getattr(s, "messages", None) or [])
previous_context = list(getattr(s, "context_messages", None) or getattr(s, "messages", None) or [])
stored_context = getattr(s, "context_messages", None)
previous_context = list(
stored_context
if isinstance(stored_context, list) and (regeneration or stored_context)
else getattr(s, "messages", None) or []
)
previous_process_wakeup_pause = dict(getattr(s, "process_wakeup_pause", {}) or {})
# Stamp stable ids on the two new rows (shared with the display merge
# below) so display and model-context copies share an id for the
+105 -3
View File
@@ -73,8 +73,99 @@ def _profile_db(profile_home: str | Path):
return db
def _profile_home_context_api():
"""Return the native context-local Hermes home setters when available."""
try:
from hermes_constants import reset_hermes_home_override, set_hermes_home_override
except Exception: # pragma: no cover - depends on installed hermes-agent
return None
return set_hermes_home_override, reset_hermes_home_override
def _native_profile_context_api(profile_home: str | Path):
"""Return context setters only when SessionDB resolves that context at call time."""
context_api = _profile_home_context_api()
if context_api is None:
return None
try:
from hermes_state import _default_db_path # type: ignore
except Exception: # pragma: no cover - depends on installed hermes-agent
return None
if not callable(_default_db_path):
return None
home = Path(profile_home).expanduser().resolve()
set_home, reset_home = context_api
try:
token = set_home(home)
except Exception: # pragma: no cover - depends on installed hermes-agent
return None
try:
resolved_db_path = Path(_default_db_path()).expanduser().resolve()
except Exception: # pragma: no cover - depends on installed hermes-agent
return None
finally:
reset_home(token)
if resolved_db_path != home / "state.db":
return None
return context_api
class _ProfileGoalManager:
"""Small WebUI-local GoalManager adapter with explicit profile persistence."""
"""Run the native GoalManager under a context-local profile home."""
def __init__(
self,
session_id: str,
*,
profile_home: str | Path,
default_max_turns: int = 20,
context_api=None,
):
if _NativeGoalManager is None:
raise RuntimeError("Hermes goal manager unavailable")
context_api = context_api or _profile_home_context_api()
if context_api is None:
raise RuntimeError("Hermes profile context unavailable")
self.session_id = session_id
self.profile_home = Path(profile_home).expanduser().resolve()
self._set_home, self._reset_home = context_api
self._manager = self._scoped(
_NativeGoalManager,
session_id=session_id,
default_max_turns=default_max_turns,
)
def _scoped(self, func, *args, **kwargs):
token = self._set_home(self.profile_home)
try:
return func(*args, **kwargs)
finally:
self._reset_home(token)
@property
def state(self):
return self._manager.state
def __getattr__(self, name):
value = getattr(self._manager, name)
if not callable(value):
return value
def scoped_call(*args, **kwargs):
return self._scoped(value, *args, **kwargs)
return scoped_call
def _restore_state(self, snapshot) -> None:
from hermes_cli.goals import save_goal # type: ignore
self._manager._state = snapshot
self._scoped(save_goal, self.session_id, snapshot)
class _LegacyProfileGoalManager:
"""Explicit-DB fallback for Hermes versions without profile context."""
def __init__(self, session_id: str, *, profile_home: str | Path, default_max_turns: int = 20):
if GoalState is None:
@@ -193,7 +284,7 @@ class _ProfileGoalManager:
if judge_goal is None:
verdict, reason = "continue", "goal judge unavailable"
else:
verdict, reason = judge_goal(state.goal, str(last_response or ""))
verdict, reason, *_ = judge_goal(state.goal, str(last_response or ""))
state.last_verdict = verdict
state.last_reason = reason
@@ -246,7 +337,15 @@ def _manager(session_id: str, *, profile_home: str | Path | None = None):
return None
if profile_home and GoalManager is _NativeGoalManager and GoalState is not None:
try:
return _ProfileGoalManager(
context_api = _native_profile_context_api(profile_home)
if context_api is not None:
return _ProfileGoalManager(
session_id=session_id,
profile_home=profile_home,
default_max_turns=_default_max_turns(),
context_api=context_api,
)
return _LegacyProfileGoalManager(
session_id=session_id,
profile_home=profile_home,
default_max_turns=_default_max_turns(),
@@ -407,6 +506,9 @@ def restore_goal_state(session_id: str, snapshot: Any, *, profile_home: str | Pa
pass
return
if isinstance(mgr, _ProfileGoalManager):
mgr._restore_state(snapshot)
return
if isinstance(mgr, _LegacyProfileGoalManager):
mgr._state = snapshot
mgr._save(snapshot)
return
+41 -21
View File
@@ -16,9 +16,13 @@ logger = logging.getLogger(__name__)
_PUBLIC_MESSAGE_INTERNAL_FIELDS = frozenset({
"api_content",
"_row_id",
"_state_db_row_id",
"_db_row_id",
"state_db_row_id",
"_active_turn_token",
"_active_turn_user",
"_fork_child_turn",
})
@@ -1114,8 +1118,14 @@ def scrub_internal_replay_fields(
return result
def _public_message_projection(message, *, _enabled: bool):
def _public_message_projection(message, *, _enabled: bool, _active_turn_token=None):
"""Return one public transcript message without internal replay fields."""
is_active = (
isinstance(message, dict)
and message.get("role") == "user"
and _active_turn_token is not None
and message.get("_active_turn_token") == _active_turn_token
)
message = scrub_internal_replay_fields([message], message_records=True)[0]
if not isinstance(message, dict):
return _redact_value(message, _enabled=_enabled)
@@ -1131,13 +1141,22 @@ def _public_message_projection(message, *, _enabled: bool):
]
else:
item[key] = _redact_value(value, _enabled=_enabled)
if is_active:
item["_active_turn_user"] = True
return item
def _redact_messages(messages, *, _enabled: bool):
def _redact_messages(messages, *, _enabled: bool, _active_turn_token=None):
if not isinstance(messages, list):
return _redact_value(messages, _enabled=_enabled)
return [_public_message_projection(message, _enabled=_enabled) for message in messages]
return [
_public_message_projection(
message,
_enabled=_enabled,
_active_turn_token=_active_turn_token,
)
for message in messages
]
def _redact_tool_calls(tool_calls, *, _enabled: bool):
@@ -1152,10 +1171,12 @@ def _redact_nested_message_containers(value, *, _enabled: bool):
return _redact_value(scrubbed, _enabled=_enabled)
result = {}
for key, child in scrubbed.items():
if key == "messages" and isinstance(child, list):
if key in {"messages", "context_messages"} and isinstance(child, list):
result[key] = _redact_messages(child, _enabled=_enabled)
elif key == "tool_calls" and isinstance(child, list):
result[key] = _redact_tool_calls(child, _enabled=_enabled)
elif key == "runtime_journal_snapshot" and isinstance(child, dict):
result[key] = _redact_nested_message_containers(child, _enabled=_enabled)
else:
result[key] = _redact_value(child, _enabled=_enabled)
return result
@@ -1174,7 +1195,7 @@ def strip_public_internal_fields(value, *, message_records: bool = False):
"""Deep-copy imported records through the shared schema scrubber.
JSON import uses this before constructing or saving a ``Session``. The
Four replay aliases belong to a message/content-part/tool-call/function
Five replay aliases belong to a message/content-part/tool-call/function
record itself; matching names inside user content or tool arguments are
ordinary JSON and must be preserved. This is intentionally independent of
the credential-redaction setting: caller-supplied provider sidecars must
@@ -1193,30 +1214,21 @@ def _copy_json_value(value):
def redact_session_data(session_dict: dict) -> dict:
"""Redact credentials from message content, tool data, and session sidecars.
Applies to: messages[], tool_calls[], todo_state, runtime_journal_snapshot,
and title.
The underlying session file is not modified; redaction is response-layer only.
Reads the ``api_redact_enabled`` setting ONCE for the entire response and
threads it through to avoid hundreds of settings.json reads per session
payload (a 50-message session has hundreds of nested strings). When the
setting is disabled this is also a fast path: the recursion still walks
but every string returns early.
"""
"""Redact credentials in the public session response without mutation."""
from api.config import load_settings
_enabled = bool(load_settings().get("api_redact_enabled", True))
if not isinstance(session_dict, dict):
return {}
result = {}
from api.process_event_utils import build_active_turn_token
_active_turn_token = build_active_turn_token(session_dict.get("active_stream_id"), session_dict.get("pending_started_at"))
for key, value in session_dict.items():
if key in _PUBLIC_MESSAGE_INTERNAL_FIELDS:
continue
if key == 'title' and isinstance(value, str):
result[key] = _redact_text(value, _enabled=_enabled)
elif key in {'messages', 'context_messages'}:
result[key] = _redact_messages(value, _enabled=_enabled)
result[key] = _redact_messages(value, _enabled=_enabled, _active_turn_token=_active_turn_token)
elif key == 'tool_calls' and isinstance(value, list):
result[key] = _redact_tool_calls(value, _enabled=_enabled)
elif key in {'todo_state', 'runtime_journal_snapshot'}:
@@ -1256,10 +1268,18 @@ def read_body(handler) -> dict:
pass
raise ValueError(f'Request body too large ({length} bytes, max {MAX_BODY_BYTES})')
raw = handler.rfile.read(length) if length else b'{}'
try:
return _json.loads(raw)
except Exception:
# A body-optional endpoint (e.g. DELETE /api/mcp/servers/{name}) may send a
# whitespace-only body; treat it like an empty body ({}) rather than a 400,
# matching the pre-#7336 lenient behavior for that shape (see gate finding).
if not raw.strip():
return {}
try:
parsed = _json.loads(raw)
except Exception:
raise ValueError('Invalid JSON body') from None
if not isinstance(parsed, dict):
raise ValueError('JSON body must be an object')
return parsed
# ── Profile cookie helpers (issue #798) ─────────────────────────────────────
+1165 -114
View File
File diff suppressed because it is too large Load Diff
+30 -1
View File
@@ -35,6 +35,34 @@ from api.workspace import get_last_workspace, load_workspaces
logger = logging.getLogger(__name__)
# ── OpenRouter namespace translation (#7514) ────────────────────────────────
#
# ``_FALLBACK_MODELS`` (api/config.py) is authored for the *direct* provider
# endpoints, so its Z.AI entries carry the provider's own ``zai/`` prefix.
# OpenRouter serves the same models under its canonical ``z-ai/`` slug, so
# projecting that list into the OpenRouter setup verbatim handed the wizard
# ids that 404 on openrouter.ai (``zai/glm-5.3`` instead of
# ``z-ai/glm-5.3``). Translate at this projection boundary only — the
# fallback catalog and the direct ``zai`` setup keep their own namespace.
#
# Every other namespace in the fallback list is already OpenRouter-canonical:
# openai/, anthropic/, google/, deepseek/, qwen/, x-ai/, mistralai/, minimax/,
# openrouter/, tencent/, nvidia/, arcee-ai/ — hence a single mapping entry.
_OPENROUTER_NAMESPACE_MAP = {"zai/": "z-ai/"}
def _to_openrouter_namespace(model_id: str) -> str:
"""Return *model_id* with direct-provider prefixes mapped to OpenRouter's.
Ids whose namespace already matches OpenRouter (or that carry no
namespace) are returned unchanged.
"""
for direct_prefix, openrouter_prefix in _OPENROUTER_NAMESPACE_MAP.items():
if model_id.startswith(direct_prefix):
return openrouter_prefix + model_id[len(direct_prefix):]
return model_id
_SUPPORTED_PROVIDER_SETUPS = {
# ── Easy start ──────────────────────────────────────────────────────
"openrouter": {
@@ -43,7 +71,8 @@ _SUPPORTED_PROVIDER_SETUPS = {
"default_model": "anthropic/claude-sonnet-4.6",
"requires_base_url": False,
"models": [
{"id": model["id"], "label": model["label"]} for model in _FALLBACK_MODELS
{"id": _to_openrouter_namespace(model["id"]), "label": model["label"]}
for model in _FALLBACK_MODELS
],
"category": "easy_start",
"quick": True,
+35
View File
@@ -182,6 +182,41 @@ def _host_without_port(host: str) -> str:
def rp_context(handler) -> tuple[str, str]:
# A reverse proxy that rewrites Host to its own upstream address (common
# when a separate SPA front-end proxies /api to this backend) leaves Host
# useless for deriving the RPID, while Origin still carries the origin the
# browser is actually on. Prefer Origin's hostname: WebAuthn's RPID-origin
# check compares against the page's origin, not the backend's.
#
# Origin is client-supplied, so this is deliberately NOT a trust decision:
# it only selects which name the ceremony is scoped to. The actual security
# gates are unchanged and live elsewhere — the authenticator will only
# release a credential whose RPID is a registrable suffix of the real page
# origin, `_client_data()` requires clientDataJSON.origin to equal the
# origin stored with the challenge, `_parse_auth_data()` compares the
# authenticator's rpIdHash against that same stored RPID, and the assertion
# signature is verified against the stored credential public key. A forged
# Origin therefore yields a self-consistent ceremony that still cannot
# produce a valid signature.
#
# Accept it only as a syntactically valid http(s) origin with a hostname,
# and rebuild the origin string from the parsed parts rather than echoing
# the raw header, so a malformed or non-http Origin falls through to Host.
browser_origin = handler.headers.get("Origin", "")
if browser_origin:
try:
from urllib.parse import urlparse
parsed = urlparse(browser_origin.strip())
if parsed.scheme in ("http", "https") and parsed.hostname:
# Re-bracket IPv6 literals: parsed.hostname strips the [] that the
# browser's clientDataJSON.origin carries, so an unbracketed
# "http://::1:8787" would never match the stored origin.
host_part = f"[{parsed.hostname}]" if ":" in parsed.hostname else parsed.hostname
netloc = host_part if parsed.port is None else f"{host_part}:{parsed.port}"
return parsed.hostname, f"{parsed.scheme}://{netloc}"
except Exception:
pass
# Fallback: derive from Host header (direct/internal access)
host = _host_without_port(handler.headers.get("Host", "localhost"))
proto = handler.headers.get("X-Forwarded-Proto", "").split(",", 1)[0].strip().lower()
if proto not in {"http", "https"}:
+37 -12
View File
@@ -1711,19 +1711,22 @@ def switch_profile(name: str, *, process_wide: bool = True) -> dict:
default_workspace = None
try:
from api.config import DEFAULT_WORKSPACE as _DW
from api.workspace import _resolve_path, _remote_terminal_workspace_candidate
lw_file = home / 'webui_state' / 'last_workspace.txt'
if lw_file.exists():
_p = lw_file.read_text(encoding='utf-8').strip()
if _p:
_pp = Path(_p).expanduser()
if _pp.is_dir():
default_workspace = str(_pp.resolve())
_pp = _resolve_path(_p, profile=name)
remote_cand = _remote_terminal_workspace_candidate(_p, profile=name)
if remote_cand is not None or _pp.is_dir():
default_workspace = str(_pp)
if default_workspace is None:
for _key in ('workspace', 'default_workspace'):
_v = cfg.get(_key)
if _v:
_pp = Path(str(_v)).expanduser().resolve()
if _pp.is_dir():
_pp = _resolve_path(str(_v), profile=name)
remote_cand = _remote_terminal_workspace_candidate(str(_v), profile=name)
if remote_cand is not None or _pp.is_dir():
default_workspace = str(_pp)
break
if default_workspace is None:
@@ -1731,8 +1734,9 @@ def switch_profile(name: str, *, process_wide: bool = True) -> dict:
if isinstance(_tc, dict):
_cwd = _tc.get('cwd', '')
if _cwd and str(_cwd) not in ('.', ''):
_pp = Path(str(_cwd)).expanduser().resolve()
if _pp.is_dir():
_pp = _resolve_path(str(_cwd), profile=name)
remote_cand = _remote_terminal_workspace_candidate(str(_cwd), profile=name)
if remote_cand is not None or _pp.is_dir():
default_workspace = str(_pp)
if default_workspace is None:
default_workspace = str(_DW)
@@ -2393,20 +2397,41 @@ def _clean_profile_config_value(value: Optional[str], field: str) -> Optional[st
def _split_webui_provider_model_value(default_model: Optional[str], model_provider: Optional[str]) -> tuple[Optional[str], Optional[str]]:
"""Normalize WebUI-internal @provider:model picker values for config.yaml."""
"""Normalize WebUI-internal @provider:model picker values for config.yaml.
Parsing is delegated to ``config._parse_provider_qualified_model_id()`` so
this agrees with the grammar every other call site uses (#6722, #6723). A
positional ``rsplit(":", 1)`` cannot tell a colon-tagged model
(``@ollama:qwen3.8:27b-mtp-q8_0``) from a multi-segment custom provider ID
(``@custom:backup:model-a``); it truncated the model name and persisted the
fragment into profile config, which the provider API then 404'd on (#7182).
"""
model = _clean_profile_config_value(default_model, "default_model")
provider = _clean_profile_config_value(model_provider, "model_provider")
if model and model.startswith("@") and ":" in model:
provider_part, model_part = model[1:].rsplit(":", 1)
provider = provider or _clean_profile_config_value(provider_part, "model_provider")
model = _clean_profile_config_value(model_part, "default_model")
from api.config import _parse_provider_qualified_model_id
parsed = _parse_provider_qualified_model_id(model)
if parsed:
model_part, provider_part = parsed
provider = provider or _clean_profile_config_value(provider_part, "model_provider")
model = _clean_profile_config_value(model_part, "default_model")
return model, provider
def _strip_webui_provider_prefix(model_id: object) -> str:
"""Return the bare model name from a WebUI ``@provider:model`` value.
Uses the same shared grammar as ``_split_webui_provider_model_value()`` so a
tagged model name survives the round trip (#7182).
"""
value = str(model_id or "").strip()
if value.startswith("@") and ":" in value:
return value.rsplit(":", 1)[1]
from api.config import _parse_provider_qualified_model_id
parsed = _parse_provider_qualified_model_id(value)
if parsed:
return str(parsed[0] or "").strip()
return value
+26
View File
@@ -425,6 +425,19 @@ def _entry_exhausted_ttl_seconds(error_code):
code = str(error_code or "").strip()
if code == "401":
return 5 * 60
if code == "402":
# #6626: keep WebUI's eligibility decision tied to the installed
# runtime contract. The runtime routes 402 via
# credential_pool._exhausted_ttl() (120s when the new
# EXHAUSTED_TTL_402_SECONDS is present, 1h fallback otherwise).
# Hard-coding 120s here would let display/probe code mark an entry
# usable before CredentialPool.select() is willing to lease it on
# mixed-version installations.
try:
from agent.credential_pool import _exhausted_ttl as _runtime_exhausted_ttl
return _runtime_exhausted_ttl(int(code))
except Exception:
return 60 * 60
return 60 * 60
@@ -824,6 +837,19 @@ def _entry_exhausted_ttl_seconds(error_code):
code = str(error_code or "").strip()
if code == "401":
return 5 * 60
if code == "402":
# #6626: keep WebUI's eligibility decision tied to the installed
# runtime contract. The runtime routes 402 via
# credential_pool._exhausted_ttl() (120s when the new
# EXHAUSTED_TTL_402_SECONDS is present, 1h fallback otherwise).
# Hard-coding 120s here would let display/probe code mark an entry
# usable before CredentialPool.select() is willing to lease it on
# mixed-version installations.
try:
from agent.credential_pool import _exhausted_ttl as _runtime_exhausted_ttl
return _runtime_exhausted_ttl(int(code))
except Exception:
return 60 * 60
return 60 * 60
+6
View File
@@ -17,6 +17,8 @@ from datetime import datetime
from pathlib import Path
from typing import Any
from api.subprocess_utils import windows_hide_flags
logger = logging.getLogger(__name__)
# Checkpoint identifiers are SHA-style hex hashes from the agent's
@@ -102,6 +104,7 @@ def _checkpoint_entry_modes(git: str, ckpt_dir: Path) -> dict[str, int]:
result = subprocess.run(
[git, "-C", str(ckpt_dir), "ls-files", "-s"],
capture_output=True, text=True, timeout=10,
creationflags=windows_hide_flags(),
)
if result.returncode != 0:
raise ValueError("Failed to list checkpoint files")
@@ -136,6 +139,7 @@ def _read_checkpoint_blob(git: str, ckpt_dir: Path, rel_path: str) -> bytes | No
result = subprocess.run(
[git, "-C", str(ckpt_dir), "show", f"HEAD:{rel_path}"],
capture_output=True, timeout=10,
creationflags=windows_hide_flags(),
)
if result.returncode != 0:
return None
@@ -250,6 +254,7 @@ def _inspect_checkpoint(ckpt_path: Path, git: str) -> dict[str, Any] | None:
result = subprocess.run(
[git, "-C", str(ckpt_path), "log", "--format=%H%n%s%n%aI", "-1"],
capture_output=True, text=True, timeout=5,
creationflags=windows_hide_flags(),
)
if result.returncode != 0 or not result.stdout.strip():
return None
@@ -272,6 +277,7 @@ def _inspect_checkpoint(ckpt_path: Path, git: str) -> dict[str, Any] | None:
files_result = subprocess.run(
[git, "-C", str(ckpt_path), "ls-files"],
capture_output=True, text=True, timeout=5,
creationflags=windows_hide_flags(),
)
file_count = len(files_result.stdout.strip().split("\n")) if files_result.stdout.strip() else 0
+355 -24
View File
@@ -6,6 +6,7 @@ Extracts approval state, not handlers, by design.
import queue
import threading
import uuid
from contextlib import contextmanager
from api.session_events import publish_session_list_changed
@@ -50,6 +51,95 @@ _GATEWAY_MIRROR_RETAINED = "_gateway_mirror_retained"
_GATEWAY_ENTRY_DATA_TOKEN_KEY = "_webui_mirror_token"
_GATEWAY_AGENT_IDENTITY_V1 = "_gateway_agent_identity_v1"
_gateway_relay_owners: dict[tuple[str, str], str] = {}
_yolo_transition_lock = threading.Lock()
_yolo_transitions: dict[str, dict] = {}
_gateway_yolo_handoff_guard = threading.Lock()
_gateway_yolo_handoffs: dict[str, dict] = {}
@contextmanager
def gateway_yolo_handoff(session_key: str):
"""Serialize one session's YOLO toggles with gateway approval dispatch."""
session_key = str(session_key or "").strip()
with _gateway_yolo_handoff_guard:
entry = _gateway_yolo_handoffs.get(session_key)
if entry is None:
entry = {"lock": threading.Lock(), "users": 0}
_gateway_yolo_handoffs[session_key] = entry
entry["users"] += 1
lock = entry["lock"]
lock.acquire()
try:
yield
finally:
lock.release()
with _gateway_yolo_handoff_guard:
entry["users"] -= 1
if entry["users"] == 0:
_gateway_yolo_handoffs.pop(session_key, None)
def begin_session_yolo_transition(session_key: str) -> object | None:
"""Register a pending YOLO enable until one approval relay settles.
Multiple tabs may relay approvals for different runs in the same session.
Track every in-flight enable intent so one failed relay cannot undo another
successful or explicit enable. Do not publish an unconfirmed enable to the
shared session flag: the gateway stream may only auto-approve later prompts
after a relay succeeds or an explicit enable wins.
"""
session_key = str(session_key or "").strip()
if not session_key:
return None
token = object()
with _yolo_transition_lock:
transition = _yolo_transitions.get(session_key)
if transition is None:
transition = {
"was_enabled": bool(is_session_yolo_enabled(session_key)),
"tokens": set(),
"committed": False,
}
_yolo_transitions[session_key] = transition
transition["tokens"].add(token)
return token
def finish_session_yolo_transition(session_key: str, token: object | None, *, succeeded: bool) -> None:
"""Settle one pending YOLO enable without exposing or applying stale state."""
session_key = str(session_key or "").strip()
if not session_key or token is None:
return
with _yolo_transition_lock:
transition = _yolo_transitions.get(session_key)
if transition is None or token not in transition["tokens"]:
return
transition["tokens"].remove(token)
if succeeded:
transition["committed"] = True
# The first confirmed relay commits YOLO immediately. Any remaining
# tokens may fail later but cannot revoke this successful enable.
enable_session_yolo(session_key)
if transition["tokens"]:
return
_yolo_transitions.pop(session_key, None)
if transition["committed"] or transition["was_enabled"]:
enable_session_yolo(session_key)
else:
disable_session_yolo(session_key)
def set_session_yolo_enabled(session_key: str, enabled: bool) -> None:
"""Apply an explicit YOLO choice and supersede in-flight rollbacks."""
session_key = str(session_key or "").strip()
if not session_key:
return
with _yolo_transition_lock:
_yolo_transitions.pop(session_key, None)
if enabled:
enable_session_yolo(session_key)
else:
disable_session_yolo(session_key)
def _approval_sse_subscribe(session_id: str) -> queue.Queue:
@@ -150,29 +240,51 @@ def reconcile_gateway_pending_mirror_locked(session_key: str) -> tuple[dict | No
live_head_entry = live_gateway_queue[0] if live_gateway_queue else None
live_head_data = getattr(live_head_entry, "data", None) or {}
has_no_run_mirror = any(
_is_gateway_mirror_entry(entry)
and not str(entry.get("run_id") or "").strip()
for entry in queue_list
)
live_run_id = str(live_head_data.get("run_id") or "").strip()
live_has_token = bool(live_head_data.get(_GATEWAY_ENTRY_DATA_TOKEN_KEY))
# Tokenize EVERY live no-run producer, and derive `live_token` from the
# authoritative head, whenever `_gateway_queues[session_key]` has live
# producers. This deliberately does NOT defer to a pre-existing no-run
# mirror: an unmatched/tokenless mirror A (which the fail-closed binding in
# submit_gateway_pending_mirror leaves deliberately unbound) must never
# suppress the live producer's token, or A masks the real pending approval
# B — B never surfaces as the head and can't be actioned, while responding
# to A resolves nothing. While any producer is live, a mirror survives only
# if it is bound to some live producer's own token — the head's via
# `live_token`, a non-head producer's via `live_local_tokens` (which is what
# keeps a non-head mirror resolvable) — so unmatched and tokenless copies
# are discarded instead of masking a real one. Tokenless-orphan retention is
# reserved for the genuine no-producer case (#7093), which lands here with an
# empty `live_gateway_queue` and therefore a `None` `live_token` anyway.
live_local_tokens: set[str] = set()
for live_entry in live_gateway_queue:
live_data = getattr(live_entry, "data", None) or {}
if str(live_data.get("run_id") or "").strip():
continue
live_entry_token = str(live_data.get(_GATEWAY_ENTRY_DATA_TOKEN_KEY) or "").strip()
if live_entry_token or not has_no_run_mirror:
live_entry_token = _gateway_mirror_entry_token(live_entry) or ""
if live_entry_token:
live_local_tokens.add(live_entry_token)
if not str(live_data.get("approval_id") or "").strip():
live_data["approval_id"] = f"gwlocal:{live_entry_token}"
live_entry_token = _gateway_mirror_entry_token(live_entry) or ""
if live_entry_token:
live_local_tokens.add(live_entry_token)
if not str(live_data.get("approval_id") or "").strip():
live_data["approval_id"] = f"gwlocal:{live_entry_token}"
# A raw id-less LOCAL entry in `_pending` (agent-side _pending_result drops
# `{command, pattern_key, pattern_keys, description}` verbatim when no
# gateway notifier is registered) is unactionable by contract: the frontend
# owner-capture requires an approval_id, so its card can neither be
# approved nor dismissed — the poll re-serves it forever. Mint a stable id
# on the entry itself (the same contract the producer-tokenization loop
# above applies to _gateway_queues producers). Minting on the stored dict
# keeps the id stable across polls, so frontend dismiss markers persist.
for entry in queue_list:
if not isinstance(entry, dict) or _is_gateway_mirror_entry(entry):
continue
if str(entry.get("approval_id") or "").strip():
continue
if str(entry.get("run_id") or "").strip():
continue
entry["approval_id"] = f"gwlocal-mirrorless:{uuid.uuid4().hex}"
live_token = (
_gateway_mirror_entry_token(live_head_entry)
if live_head_entry and live_head_data
and (live_run_id or live_has_token or not has_no_run_mirror)
else None
)
if live_token and live_run_id and not str(live_head_data.get("approval_id") or "").strip():
@@ -210,18 +322,28 @@ def reconcile_gateway_pending_mirror_locked(session_key: str) -> tuple[dict | No
live_mirror_present = True
continue
if live_token:
if entry.get(_GATEWAY_MIRROR_RETAINED):
rebuilt.append(entry)
continue
if entry_token:
changed = True
continue
deferred_run_entries.append(entry)
continue
if not entry_token:
if entry.get(_GATEWAY_MIRROR_RETAINED) or not entry_token:
rebuilt.append(entry)
continue
changed = True
continue
if entry.get(_GATEWAY_MIRROR_RETAINED):
# A retained mirror is one whose own producer had already vanished when
# the user responded (the missing-producer 409 kept visible until an
# explicit teardown). It survives ONLY while no producer is live: once
# `_gateway_queues[session_key]` holds a real producer again, an
# unresolvable retained mirror must not mask it, so fall through to the
# normal matching below (which keeps it if it still matches a live
# token and discards it otherwise).
if entry.get(_GATEWAY_MIRROR_RETAINED) and not live_gateway_queue:
rebuilt.append(entry)
continue
@@ -276,10 +398,16 @@ def reconcile_gateway_pending_mirror_locked(session_key: str) -> tuple[dict | No
return head, total, changed
def _gateway_pending_mirror_locked(session_key: str, approval_id: str = "", run_id: str = "") -> dict | None:
def _gateway_pending_mirror_locked(
session_key: str,
approval_id: str = "",
run_id: str = "",
mirror_token: str = "",
) -> dict | None:
"""Return the exact live run-backed mirror under `_lock`."""
approval_id = str(approval_id or "").strip()
run_id = str(run_id or "").strip()
mirror_token = str(mirror_token or "").strip()
queue = _pending.get(session_key)
entries = queue if isinstance(queue, list) else [queue] if queue else []
if approval_id:
@@ -296,6 +424,8 @@ def _gateway_pending_mirror_locked(session_key: str, approval_id: str = "", run_
continue
if run_id and entry_run_id != run_id:
continue
if mirror_token and str(entry.get(_GATEWAY_MIRROR_TOKEN) or "").strip() != mirror_token:
continue
if run_id:
return entry
if matched_entry is not None:
@@ -305,19 +435,46 @@ def _gateway_pending_mirror_locked(session_key: str, approval_id: str = "", run_
for entry in entries:
if not _is_gateway_mirror_entry(entry) or not str(entry.get("run_id") or "").strip():
continue
if run_id and entry.get("run_id") == run_id:
return entry
if mirror_token and str(entry.get(_GATEWAY_MIRROR_TOKEN) or "").strip() != mirror_token:
continue
if run_id:
if entry.get("run_id") == run_id:
return entry
continue
# With no caller-supplied identity, the queue order is authoritative:
# return the current run-backed projection and let its embedded
# `(approval_id, run_id)` identify the exact relay owner.
return entry
return None
def gateway_pending_mirror(session_key: str, approval_id: str = "", run_id: str = "") -> dict | None:
def gateway_pending_mirror(
session_key: str,
approval_id: str = "",
run_id: str = "",
mirror_token: str = "",
) -> dict | None:
"""Return an exact live run-backed mirror for this session."""
with _lock:
reconcile_gateway_pending_mirror_locked(session_key)
entry = _gateway_pending_mirror_locked(session_key, approval_id, run_id)
entry = _gateway_pending_mirror_locked(session_key, approval_id, run_id, mirror_token)
return dict(entry) if entry else None
def gateway_pending_mirrors(session_key: str) -> list[dict]:
"""Return every currently parked run-backed mirror in queue order."""
with _lock:
reconcile_gateway_pending_mirror_locked(session_key)
queue = _pending.get(session_key)
entries = queue if isinstance(queue, list) else [queue] if queue else []
return [
dict(entry)
for entry in entries
if _is_gateway_mirror_entry(entry)
and str(entry.get("run_id") or "").strip()
]
def claim_gateway_approval_relay_owner(session_key: str, run_id: str, approval_id: str) -> bool:
"""Claim the single-flight relay owner for one `(session, run)` pair."""
session_key = str(session_key or "").strip()
@@ -348,7 +505,12 @@ def release_gateway_approval_relay_owner(session_key: str, run_id: str, approval
_gateway_relay_owners.pop(key, None)
def retire_gateway_pending_mirror(session_key: str, approval_id: str = "", run_id: str = "") -> bool:
def retire_gateway_pending_mirror(
session_key: str,
approval_id: str = "",
run_id: str = "",
mirror_token: str = "",
) -> bool:
"""Retire one approval, or every mirror for a terminal run."""
with _lock:
reconcile_gateway_pending_mirror_locked(session_key)
@@ -359,7 +521,12 @@ def retire_gateway_pending_mirror(session_key: str, approval_id: str = "", run_i
retained_gateway_queue = gateway_queue
gateway_queue_changed = False
if approval_id:
match = _gateway_pending_mirror_locked(session_key, approval_id, run_id)
match = _gateway_pending_mirror_locked(
session_key,
approval_id,
run_id,
mirror_token,
)
if match is None and not normalized_run_id:
match = next((entry for entry in entries if _is_gateway_mirror_entry(entry)
and not str(entry.get("run_id") or "").strip()
@@ -419,7 +586,30 @@ def _gateway_mirrored_pending_run_id(session_key: str, approval_id: str) -> str
def submit_gateway_pending_mirror(session_key: str, approval: dict) -> tuple[dict | None, int]:
"""Mirror the live gateway head into WebUI polling state under a typed tag."""
"""Mirror the live gateway head into WebUI polling state under a typed tag.
Every mirrored entry describes one pending approval to the UI. Run-backed
mirrors carry ``run_id`` (remote gateway runs) and are bound to the parked
``_ApprovalEntry`` via ``approval_id``. No-run mirrors, which represent an
in-process (legacy) approval parked in ``_gateway_queues``, have no
``run_id`` and instead bind back to the live entry through
``_GATEWAY_MIRROR_TOKEN``: ``_resolve_approval_legacy()`` matches
``pending[_GATEWAY_MIRROR_TOKEN]`` against the live entry's
``_webui_mirror_token`` before it will call ``resolve_gateway_pending_local()``
to unblock the agent thread. A no-run mirror's token is stamped from its
OWN live producer — resolved via the mirror's ``request_id``/
``approval_id`` — and from nothing else. If no live producer's identity
matches, the mirror is left tokenless (fail closed): guessing ownership
from "first unclaimed token" or the queue head would let approving THIS
(possibly stale/foreign) mirror resolve a DIFFERENT live producer than
the one the user actually saw, which is an approval-integrity violation,
not a convenience. A tokenless mirror is only wired up automatically when
there is truly no live producer at all (#7093); ``reconcile_gateway_
pending_mirror_locked`` binds a fresh, correctly-bound mirror to the
authoritative live head on its own. Without a token, THIS specific click
returns ``ok:true`` and the card clears without unblocking any producer —
the correct producer's own card reappears on the next reconcile.
"""
with _lock:
run_id = str(approval.get("run_id") or "").strip()
approval_id = str(approval.get("approval_id") or "").strip()
@@ -439,6 +629,29 @@ def submit_gateway_pending_mirror(session_key: str, approval: dict) -> tuple[dic
),
None,
)
if exact_local_entry is None and not run_id:
# Fall back to matching on the core's per-approval `request_id`.
# The gateway core notifies WebUI with a COPY of the entry payload
# (`notify_cb(dict(entry.data))`), so the identity match above
# (`entry.data is approval`) never holds for a real gateway head,
# and a local `_ApprovalEntry` carries a `request_id` but no
# `approval_id`, so the approval_id fallback misses too. The
# `request_id` is stamped once on the source entry
# (`_ApprovalEntry.__init__` -> `data.setdefault("request_id", ...)`)
# and preserved through the copy, so it uniquely reunites the
# notified copy with its queued entry. Without this, the mirror is
# created with no token, reconcile keeps the orphan, and
# `_session_has_pending_approval` stays True after the entry is
# dropped (the stale-approval-card dead-end, #4948 local variant).
request_id = str(approval.get("request_id") or "").strip()
if request_id:
exact_local_entry = next(
(
entry for entry in live_gateway_queue
if str(((getattr(entry, "data", None) or {}).get("request_id") or "")).strip() == request_id
),
None,
)
if exact_local_entry is not None:
mirror_entries = _normalize_pending_queue_locked(session_key)
entries_to_mirror = [live_gateway_queue[0]] if live_gateway_queue else []
@@ -500,6 +713,8 @@ def submit_gateway_pending_mirror(session_key: str, approval: dict) -> tuple[dic
mirror_entry["run_id"] = run_id
mirror_entry["approval_id"] = approval_id
mirror_entry[_GATEWAY_MIRROR_FLAG] = True
mirror_entry[_GATEWAY_MIRROR_TOKEN] = uuid.uuid4().hex
mirror_entry[_GATEWAY_MIRROR_RETAINED] = True
if not _gateway_pending_mirror_locked(session_key, approval_id=approval_id, run_id=run_id):
_normalize_pending_queue_locked(session_key).append(mirror_entry)
elif not exact_local_entry:
@@ -520,9 +735,51 @@ def submit_gateway_pending_mirror(session_key: str, approval: dict) -> tuple[dic
if no_run_mirror:
approval["approval_id"] = str(no_run_mirror.get("approval_id") or approval_id).strip()
elif not _gateway_pending_mirror_locked(session_key, approval_id=approval_id):
# Stamp the mirror token from the mirror's OWN live producer so
# the first respond can link this mirror to the right
# _ApprovalEntry in _gateway_queues. Without it,
# _resolve_approval_legacy cannot match the no-run mirror to its
# gateway entry (both token fields are empty) and the agent
# thread is never unblocked on the first click (#6008 legacy).
# We MUST NOT blindly take the live head: for a non-head mirror
# (multiple parked producers, #7093 exact-producer isolation)
# the head belongs to a sibling, and stamping its token would
# bind the mirror to the wrong entry. Only an explicit
# request_id/approval_id match may bind a token.
#
# FAIL CLOSED when no producer's identity matches: never infer
# ownership from "first unclaimed token" or the queue head. A
# stale/foreign approval (mismatched request_id, or no
# identity at all) that borrows another live producer's token
# would let approving THIS mirror resolve a DIFFERENT producer
# than the one the user actually saw — an approval-integrity
# violation, not a convenience (found in review of 818fd2fd).
# A tokenless orphan is only legitimate when there is no live
# producer at all (#7093); reconcile_gateway_pending_mirror_locked
# binds a real mirror to the authoritative live head on its own.
live_queue_for_mirror = _gateway_queues.get(session_key) or []
request_id_for_mirror = str(approval.get("request_id") or "").strip()
mirror_producer = None
if request_id_for_mirror or approval_id:
for cand in live_queue_for_mirror:
cand_data = getattr(cand, "data", None) or {}
if (request_id_for_mirror and
str(cand_data.get("request_id") or "").strip() == request_id_for_mirror):
mirror_producer = cand
break
if (approval_id and not request_id_for_mirror and
str(cand_data.get("approval_id") or "").strip() == approval_id):
mirror_producer = cand
break
mirror_token = (
_gateway_mirror_entry_token(mirror_producer)
if mirror_producer is not None else None
)
mirror_entry = dict(approval)
mirror_entry["approval_id"] = approval_id
mirror_entry[_GATEWAY_MIRROR_FLAG] = True
if mirror_token:
mirror_entry[_GATEWAY_MIRROR_TOKEN] = mirror_token
_normalize_pending_queue_locked(session_key).append(mirror_entry)
head, total, _changed = reconcile_gateway_pending_mirror_locked(session_key)
_approval_sse_notify_locked(session_key, head, total)
@@ -609,6 +866,80 @@ def resolve_gateway_pending_local_no_run_mirror(
return True, 1, head, total
def resolve_gateway_pending_local_all(
session_key: str,
choice: str,
reason: str | None = None,
) -> tuple[int, dict | None, int]:
"""Resolve every parked local/no-run approval without touching remote runs."""
targets = []
removed_pending = False
with _lock:
reconcile_gateway_pending_mirror_locked(session_key)
gateway_queue = _gateway_queues.get(session_key) or []
retained_gateway_queue = []
for entry in gateway_queue:
data = getattr(entry, "data", None) or {}
if str(data.get("run_id") or "").strip():
retained_gateway_queue.append(entry)
else:
targets.append(entry)
if retained_gateway_queue:
_gateway_queues[session_key] = retained_gateway_queue
else:
_gateway_queues.pop(session_key, None)
queue = _pending.get(session_key)
entries = queue if isinstance(queue, list) else [queue] if queue else []
retained_pending = [
entry
for entry in entries
if _is_gateway_mirror_entry(entry)
and str(entry.get("run_id") or "").strip()
]
removed_pending = len(retained_pending) != len(entries)
if retained_pending:
_pending[session_key] = retained_pending
else:
_pending.pop(session_key, None)
head, total, _changed = reconcile_gateway_pending_mirror_locked(session_key)
_approval_sse_notify_locked(session_key, head, total)
for entry in targets:
entry.result = choice
if reason:
entry.reason = reason
entry.event.set()
if targets or removed_pending:
publish_session_list_changed("attention_resolved")
return len(targets), head, total
def settle_gateway_pending_local_notification(
session_key: str,
approval: dict,
) -> tuple[bool, dict | None, int]:
"""Auto-resolve or publish one local approval at the YOLO handoff boundary.
The Agent adds its blocking entry before invoking WebUI's notify callback.
Serialize that callback with session YOLO commit/disable so a waiter arriving
after a drain snapshot cannot be parked behind an already-committed enable.
Run-backed approvals stay on the Runs API path and are never resolved here.
"""
with gateway_yolo_handoff(session_key):
run_id = str((approval or {}).get("run_id") or "").strip()
if not run_id and is_session_yolo_enabled(session_key):
_resolved, head, total = resolve_gateway_pending_local_all(
session_key,
"once",
)
return True, head, total
head, total = submit_gateway_pending_mirror(session_key, approval)
return False, head, total
def submit_pending(session_key: str, approval: dict) -> None:
"""Append a pending approval to the per-session queue.
+161 -7
View File
@@ -1,7 +1,7 @@
"""Session-list cache helpers extracted from api.routes."""
import os
import copy
import os
import re
import threading
import time
@@ -44,6 +44,143 @@ _SESSIONS_CACHE_ALL_PROFILES_INVALIDATION_VERSION = 0
_SESSIONS_CACHE_PROFILE_INVALIDATION_VERSION: dict[str, int] = {}
_SIDEBAR_SESSION_RESPONSE_FIELDS = {
"session_id",
"title",
"display_title",
"_state_db_title",
"workspace",
"model",
"model_provider",
"message_count",
"user_message_count",
"created_at",
"updated_at",
"last_message_at",
"pinned",
"archived",
"project_id",
"profile",
"input_tokens",
"output_tokens",
"estimated_cost",
"cache_read_tokens",
"cache_write_tokens",
"cache_hit_percent",
"personality",
"context_length",
"config_context_length",
"window_usage_percent",
"source_tag",
"raw_source",
"session_source",
"source_label",
"is_cli_session",
"is_messaging_session",
"is_streaming",
"cron_running",
"active_stream_id",
"has_pending_user_message",
"pending_started_at",
"default_hidden",
"worktree_path",
"worktree_branch",
"parent_session_id",
"parent_title",
"parent_source",
"relationship_type",
"pre_compression_snapshot",
"_lineage_root_id",
"_lineage_tip_id",
"_compression_segment_count",
"_lineage_collapsed_count",
"_parent_lineage_root_id",
"_parent_lineage_tip_id",
"_cross_surface_child_session",
"match_type",
"match_preview",
# Preserved so the sidebar can suppress rename / action-menu / swipe on
# read-only sessions and render the detailed gateway model label. Only the
# latest bounded routing object is included; routing history stays excluded.
"read_only",
"is_read_only",
"gateway_routing",
}
def _session_list_cache_sidebar_fields() -> set[str]:
"""Return the canonical bounded field set used by the list response."""
return _SIDEBAR_SESSION_RESPONSE_FIELDS
def _session_list_cache_copy_value(value):
"""Copy only mutable values after the transcript-bearing projection."""
if isinstance(value, (dict, list, set)):
return copy.deepcopy(value)
return value
def _session_list_cache_copy_row(row: dict) -> dict:
return {
key: _session_list_cache_copy_value(value)
for key, value in row.items()
}
def _session_list_cache_bounded_payload(payload: dict) -> dict:
"""Project a builder payload to data the sidebar can actually consume.
Builders operate on full session snapshots because they also perform
reconciliation and lineage decisions. The cache must not retain those
snapshots: a single long transcript can otherwise make every cache set/get
allocate hundreds of megabytes via deepcopy().
"""
fields = _session_list_cache_sidebar_fields()
def project_rows(rows):
projected = []
for row in rows or []:
if not isinstance(row, dict):
projected.append({})
continue
item = {key: row[key] for key in fields if key in row}
projected.append(_session_list_cache_copy_row(item))
return projected
bounded = {
key: _session_list_cache_copy_value(value)
for key, value in payload.items()
if key not in {"sessions", "sidebar_reference_sessions"}
}
if "sessions" in payload:
bounded["sessions"] = project_rows(payload.get("sessions"))
if "sidebar_reference_sessions" in payload:
bounded["sidebar_reference_sessions"] = project_rows(
payload.get("sidebar_reference_sessions")
)
return bounded
def _session_list_cache_copy_payload(payload: dict) -> dict:
"""Copy only the already-bounded cache shape for request-local mutation."""
copied = {
key: _session_list_cache_copy_value(value)
for key, value in payload.items()
if key not in {"sessions", "sidebar_reference_sessions"}
}
if "sessions" in payload:
copied["sessions"] = [
_session_list_cache_copy_row(row)
for row in payload.get("sessions", [])
]
if "sidebar_reference_sessions" in payload:
copied["sidebar_reference_sessions"] = [
_session_list_cache_copy_row(row)
for row in payload.get("sidebar_reference_sessions", [])
]
return copied
def get_session_list_cache_snapshot() -> dict[str, object]:
"""Return scalar cache occupancy without waiting or changing LRU state.
@@ -240,7 +377,7 @@ def _session_list_cache_get(
if stamp != current_stamp:
if allow_stale:
_SESSIONS_CACHE.move_to_end(key)
return copy.deepcopy(payload), False
return _session_list_cache_copy_payload(payload), False
_SESSIONS_CACHE.pop(key, None)
return None, False
# #4808: widen the freshness window while a turn is streaming so the fixed
@@ -251,10 +388,10 @@ def _session_list_cache_get(
fresh = (now - ts) < ttl
if fresh:
_SESSIONS_CACHE.move_to_end(key)
return copy.deepcopy(payload), True
return _session_list_cache_copy_payload(payload), True
if allow_stale:
_SESSIONS_CACHE.move_to_end(key)
return copy.deepcopy(payload), False
return _session_list_cache_copy_payload(payload), False
_SESSIONS_CACHE.pop(key, None)
return None, False
@@ -278,15 +415,32 @@ def _session_list_cache_stale_reason(key: tuple) -> str | None:
return None
def _session_list_cache_set(key: tuple, payload: dict) -> None:
def _session_list_cache_set(
key: tuple,
payload: dict,
*,
expected_invalidation_stamp: tuple[int, int] | None = None,
) -> bool:
if not isinstance(payload, dict):
return
return False
stamp = _session_list_cache_resolved_source_stamp(key)
bounded = _session_list_cache_bounded_payload(payload)
with _SESSIONS_CACHE_LOCK:
_SESSIONS_CACHE[key] = (time.monotonic(), stamp, copy.deepcopy(payload))
# Projection intentionally happens outside the lock so large source rows
# cannot block cache hits. Re-check the caller's pre-build generation
# atomically before insertion, otherwise a rename/archive/delete clear
# that lands during projection can be undone by this stale write.
if (
expected_invalidation_stamp is not None
and _session_list_cache_invalidation_stamp(key)
!= expected_invalidation_stamp
):
return False
_SESSIONS_CACHE[key] = (time.monotonic(), stamp, bounded)
_SESSIONS_CACHE.move_to_end(key)
while len(_SESSIONS_CACHE) > _SESSIONS_CACHE_MAX_ENTRIES:
_SESSIONS_CACHE.popitem(last=False)
return True
def _session_list_cache_clear(profile: str | None = None) -> None:
+2762 -872
View File
File diff suppressed because it is too large Load Diff
+536 -16
View File
@@ -9,17 +9,513 @@ from __future__ import annotations
import json
import logging
import uuid
import copy
import hashlib
import math
from dataclasses import dataclass
from contextlib import nullcontext
from bisect import bisect_left
from typing import Any
from api.config import LOCK, _get_session_agent_lock
from api.models import get_session, SESSIONS
from api.agent_sessions import normalize_agent_session_source
logger = logging.getLogger(__name__)
AUTO_TITLE_LABELS = {'untitled', 'new chat'}
class RegenerationUnavailable(Exception):
def __init__(self, code: str, status: int = 409, message: str | None = None):
super().__init__(message or code)
self.code = code
self.status = status
def _regeneration_source_class(value):
raw = str(value or "").strip().lower()
if not raw:
return ""
if raw == "fork":
return "fork"
normalized = normalize_agent_session_source(raw).get("session_source")
return str(normalized or raw).strip().lower()
def _regeneration_source_allowed(value):
return _regeneration_source_class(value) in {"webui", "fork"}
def _selected_regeneration_turn_owned(session, row) -> bool:
"""Accept only a final row whose provenance proves WebUI ownership."""
if getattr(session, "read_only", False) or not isinstance(row, dict):
return False
session_source = _regeneration_source_class(getattr(session, "session_source", None))
imported_session = bool(
getattr(session, "is_cli_session", False)
or session_source not in {"", "webui", "fork"}
)
raw_sources = (
getattr(session, "raw_source", None),
getattr(session, "source_tag", None),
)
row_source = row.get("_source") or row.get("source")
if session_source == "fork":
if any(source and not _regeneration_source_allowed(source) for source in raw_sources):
return False
if row_source and not _regeneration_source_allowed(row_source):
return False
return bool(
getattr(session, "parent_session_id", None)
and row.get("_fork_child_turn") == getattr(session, "session_id", None)
)
if imported_session:
token = row.get("_active_turn_token")
if not isinstance(token, str) or not token.strip() or ":" not in token:
return False
stream_id, started_at = token.rsplit(":", 1)
try:
started = float(started_at)
except (TypeError, ValueError):
return False
from api.process_event_utils import build_active_turn_token
if not math.isfinite(started) or started <= 0:
return False
if build_active_turn_token(stream_id, started) != token:
return False
else:
if any(source and not _regeneration_source_allowed(source) for source in raw_sources):
return False
if row_source and not _regeneration_source_allowed(row_source):
return False
return True
@dataclass(frozen=True)
class RegenerationTurn:
user_index: int
assistant_index: int
message: dict
message_text: str
attachments: list
source: str
message_count: int
revision: str
row_digest: str
@dataclass(frozen=True)
class RegenerationPlan:
canonical_rows: list
canonical_context: list
turn: RegenerationTurn
revision: str
row_digest: str
message_count: int
truncation_boundary: int
def plan_regeneration(session, *, expected_revision=None, lock_held=False):
"""Prepare one canonical display/context pair for a locked regeneration."""
lock_context = nullcontext() if lock_held else _get_session_agent_lock(session.session_id)
with lock_context:
rows, context = regeneration_state(session, use_sidecar=True)
revision = regeneration_revision_for(rows, session=session, context=context)
if expected_revision is not None and expected_revision != revision:
raise RegenerationUnavailable("stale_regeneration_revision")
turn = resolve_regeneration_turn(
rows, session=session, expected_revision=revision,
lock_held=True, context=context,
)
return RegenerationPlan(
canonical_rows=copy.deepcopy(rows),
canonical_context=copy.deepcopy(context),
turn=turn,
revision=revision,
row_digest=turn.row_digest,
message_count=len(rows),
truncation_boundary=turn.user_index + 1,
)
def apply_regeneration_plan(
session,
plan: RegenerationPlan,
*,
return_context_user: bool = False,
):
"""Install the prepared pair and truncate it without a second authority read."""
def _result(success, context_user=None):
return (success, context_user) if return_context_user else success
if not isinstance(plan, RegenerationPlan):
return _result(False)
rows = copy.deepcopy(plan.canonical_rows)
context = copy.deepcopy(plan.canonical_context)
if len(rows) != plan.message_count or plan.truncation_boundary != plan.turn.user_index + 1:
return _result(False)
if regeneration_revision_for(rows, session=session, context=context) != plan.revision:
return _result(False)
session.messages = rows
session.context_messages = context
current = session.messages[plan.turn.user_index]
if not isinstance(current, dict) or current.get("role") != "user":
return _result(False)
truncate_session_at_keep(session, plan.truncation_boundary)
prepared_context, context_boundary_index = truncate_context_for_display_keep(
context,
rows,
plan.truncation_boundary,
return_boundary_index=True,
)
session.context_messages = prepared_context if prepared_context or not context else context[: plan.truncation_boundary]
retained_context_user = None
if context_boundary_index is not None:
for context_row in reversed(session.context_messages[: context_boundary_index + 1]):
if isinstance(context_row, dict) and context_row.get("role") == "user":
retained_context_user = context_row
break
return _result(True, retained_context_user)
def snapshot_regeneration_state(session):
return copy.deepcopy(session.__dict__)
def restore_regeneration_state(session, snapshot):
session.__dict__.clear()
session.__dict__.update(copy.deepcopy(snapshot))
def regeneration_revision_for(rows, *, session=None, context=None) -> str:
"""Hash the canonical writable transcript and its aligned context."""
payload = json.dumps(
{
"session_id": str(getattr(session, "session_id", "") or "") if session is not None else "",
"messages": list(rows or []),
"context_messages": list(context or []),
"truncation_watermark": getattr(session, "truncation_watermark", None) if session is not None else None,
"truncation_boundary": getattr(session, "truncation_boundary", None) if session is not None else None,
},
sort_keys=True,
separators=(",", ":"),
default=str,
)
return hashlib.sha256(payload.encode("utf-8")).hexdigest()
def regeneration_transcript(session, *, state_messages=None):
"""Return the state.db-reconciled transcript used by every authority consumer."""
if state_messages is None:
return regeneration_state(session)[0]
from api.models import reconciled_state_db_messages_for_session
return reconciled_state_db_messages_for_session(session, state_messages=state_messages)
def regeneration_context(session):
return regeneration_state(session)[1]
_REGENERATION_SIDECAR_ANCHOR_BUDGET = 200
def _sidecar_regeneration_read_floor(session):
"""Return a state.db tail-read floor anchored by the already-loaded sidecar.
#6826: regenerating a large session must not re-materialize the full
state.db transcript (a >1min stall on big sessions). When the in-memory
sidecar is a usable reconciliation base — append-only session with no
active truncation markers and timestamped rows — return the timestamp
floor for a bounded ``since_timestamp`` tail read. Rows at/after the
floor (including any gateway/server-applied tail the sidecar has not seen
yet) are re-read and merged, and rows the sidecar already carries are
deduplicated by the append-only merge, so the #6611 reconciliation
authority is preserved on the fast path.
Returns ``None`` when the sidecar cannot anchor a tail read; callers then
fall back to the full state.db read (unchanged behavior).
"""
if getattr(session, "truncation_watermark", None) not in (None, ""):
return None
if getattr(session, "truncation_boundary", None) not in (None, ""):
return None
messages = getattr(session, "messages", None)
if not isinstance(messages, list) or not messages:
return None
from api.models import _message_timestamp_as_float
timestamps = [_message_timestamp_as_float(message) for message in messages]
if any(timestamp is None for timestamp in timestamps):
return None
# Conservative anchor: re-read a bounded tip window so sub-second/clock
# drift near the sidecar tip cannot hide a concurrently appended state.db
# row, while the raw read stays tiny for huge sessions.
return min(timestamps[-_REGENERATION_SIDECAR_ANCHOR_BUDGET:])
def _bounded_tail_snapshot_if_safe(session, read_floor):
"""Return the bounded tail rows ONLY when it is provably identical to the
full read; otherwise None (caller must fall back to the full read).
#6826 r3: the skipped state.db prefix (rows older than the floor) must be
represented identically in the sidecar — same count AND same ordered
visible identity — and the bounded tail must not repeat any skipped key
(occurrence-count collision: a new tail turn repeating an older prompt
would be mistaken for the old sidecar duplicate and dropped). The prefix
proof and the tail data come from ONE read transaction (no TOCTOU).
Any mismatch, missing database, or uncertainty returns None, so the #6611
regeneration authority never operates on an unreconciled view.
"""
sid = getattr(session, "session_id", None)
if not sid:
return None
profile = getattr(session, "profile", None)
from api.models import (
_session_message_visible_key,
get_state_db_regeneration_tail_snapshot,
)
snap = get_state_db_regeneration_tail_snapshot(sid, read_floor, profile=profile)
if snap is None:
return None # cannot obtain a stable single-snapshot → full read
# Compression-anchor coverage: if the anchor predates the floor the bounded
# read can drop compacted-tail context rows (display may still match).
anchor = getattr(session, "compression_anchor_message_key", None)
if isinstance(anchor, dict):
try:
anchor_ts = float(anchor.get("ts"))
except (TypeError, ValueError):
anchor_ts = None
if anchor_ts is None or anchor_ts < read_floor:
return None
prefix = snap["prefix"]
if prefix.get("count") == 0 and prefix.get("null_timestamp_count") == 0:
# Empty skipped prefix: the bounded read already covers every row.
return snap["tail"]
# Non-empty skipped prefix: prove identical ordered visible identity.
sidecar_keys = []
for message in getattr(session, "messages", None) or []:
if not isinstance(message, dict):
continue
try:
ts = float(message.get("timestamp"))
except (TypeError, ValueError):
ts = None
if ts is not None and ts < read_floor:
key = _session_message_visible_key(message)
if key is None:
return None
sidecar_keys.append(key)
if list(snap["prefix_keys"]) != sidecar_keys:
return None # mismatch → full read
# Occurrence-count collision (#6826 r3 #1): if any bounded-tail key ALSO
# occurs in the skipped prefix, the reconciler may drop the repeated tail
# row — fall back conservatively.
prefix_key_set = set(snap["prefix_keys"])
for key in snap["tail_keys"]:
if key in prefix_key_set:
return None
# In-tail duplicates (#6826 r5): a repeated message wholly inside the
# bounded tail makes the reconciler's context dedup diverge from the full
# read (display may still match) → refuse the bounded path.
if len(snap["tail_keys"]) != len(set(snap["tail_keys"])):
return None
return snap["tail"]
def regeneration_state(session, *, use_sidecar=False):
"""Read one immutable state.db snapshot and reconcile both transcript views.
``use_sidecar=True`` (#6826) anchors the state.db read to the already
loaded in-memory sidecar: only a bounded tail (``since_timestamp`` floor)
is re-read instead of the full transcript, and both views still route
through :func:`reconciled_state_db_messages_for_session`, so the #6611
reconciliation authority (recovered display/context pair survives local
and gateway apply) is preserved on the fast path.
The bounded tail is only trusted when
:func:`_bounded_tail_snapshot_if_safe` proves the skipped state.db prefix
is identical in the sidecar (count + ordered visible identity + no
occurrence collision + compression anchor coverage), and the tail rows
come from the SAME single read transaction as the proof (no TOCTOU);
otherwise the read falls back to the full transcript.
"""
from api.models import (
get_state_db_session_messages,
reconciled_state_db_messages_for_session,
)
bounded_tail = None
if use_sidecar:
read_floor = _sidecar_regeneration_read_floor(session)
if read_floor is not None:
bounded_tail = _bounded_tail_snapshot_if_safe(session, read_floor)
if bounded_tail is not None:
state_messages = bounded_tail
else:
state_messages = get_state_db_session_messages(
getattr(session, "session_id", None),
profile=getattr(session, "profile", None),
)
return (
reconciled_state_db_messages_for_session(
session,
state_messages=state_messages,
),
reconciled_state_db_messages_for_session(
session,
prefer_context=True,
state_messages=state_messages,
),
)
def regeneration_revision(session) -> str:
rows, context = regeneration_state(session, use_sidecar=True)
return regeneration_revision_for(
rows,
session=session,
context=context,
)
def regeneration_authority(
session,
rows=None,
*,
context=None,
full_transcript=True,
canonical_state=None,
):
"""Mint a revision only for a complete, writable, canonical transcript."""
if not full_transcript:
return None
if getattr(session, "active_stream_id", None) or getattr(session, "pending_user_message", None):
return None
canonical_rows, canonical_context = canonical_state or regeneration_state(session)
rows = list(canonical_rows if rows is None else rows)
if not rows:
return None
if rows != canonical_rows:
return None
if context is not None and list(context or []) != canonical_context:
return None
try:
resolve_regeneration_turn(
canonical_rows,
session=session,
context=canonical_context,
)
except RegenerationUnavailable:
return None
return regeneration_revision_for(
canonical_rows,
session=session,
context=canonical_context,
)
def resolve_regeneration_turn(
rows,
*,
session=None,
expected_revision=None,
lock_held=False,
context=None,
):
"""Select the current session's final complete local exchange under its lock."""
legacy_session_call = session is None and not isinstance(rows, (list, tuple))
legacy_context = None
if legacy_session_call:
session = rows
rows, legacy_context = regeneration_state(session)
lock_context = (
_get_session_agent_lock(session.session_id)
if legacy_session_call and not lock_held
else nullcontext()
)
with lock_context:
rows = list(rows or [])
if context is None:
context = legacy_context
if context is None:
_, context = regeneration_state(session)
context = list(context)
revision = regeneration_revision_for(rows, session=session, context=context)
if expected_revision is not None and expected_revision != revision:
raise RegenerationUnavailable("stale_regeneration_revision")
if getattr(session, "active_stream_id", None):
raise RegenerationUnavailable("session_active")
if getattr(session, "pending_user_message", None):
raise RegenerationUnavailable("session_active")
assistant_index = next(
(
index
for index in range(len(rows) - 1, -1, -1)
if isinstance(rows[index], dict)
and rows[index].get("role") == "assistant"
and _assistant_message_has_final_visible_text(rows[index])
),
None,
)
if assistant_index is not None:
index = next(
(
candidate
for candidate in range(assistant_index - 1, -1, -1)
if isinstance(rows[candidate], dict)
and rows[candidate].get("role") == "user"
),
None,
)
else:
index = None
if index is not None:
if any(
isinstance(row, dict) and row.get("role") == "user"
for row in rows[assistant_index + 1:]
):
raise RegenerationUnavailable("no_regenerable_turn", 400)
if any(
isinstance(row, dict) and row.get("role") in {"assistant", "tool"}
for row in rows[assistant_index + 1:]
):
raise RegenerationUnavailable("no_regenerable_turn", 400)
row = rows[index]
if not _selected_regeneration_turn_owned(session, row):
raise RegenerationUnavailable("regeneration_read_only", 403)
content = _extract_text(row.get("content", ""))
if content:
row_digest = hashlib.sha256(
json.dumps(
row,
sort_keys=True,
separators=(",", ":"),
default=str,
).encode("utf-8")
).hexdigest()
return RegenerationTurn(
index,
assistant_index,
copy.deepcopy(row),
content,
copy.deepcopy(row.get("attachments") or []),
str(row.get("_source") or "webui"),
len(rows),
revision,
row_digest,
)
raise RegenerationUnavailable("no_regenerable_turn", 400)
def _assistant_message_has_final_visible_text(message) -> bool:
from api.streaming import _assistant_message_has_final_visible_text as _has_final_text
return _has_final_text(message)
def _live_active_stream_id(session) -> str | None:
"""Return session.active_stream_id ONLY if that stream is live in THIS
process; else None.
@@ -30,19 +526,26 @@ def _live_active_stream_id(session) -> str | None:
the hidden-tab poller) would make a client attach its renderer to a stream
that never emits — a permanent fake "thinking" state. Liveness test mirrors
routes._clear_stale_stream_state: live iff present in STREAMS (open SSE
channel) or ACTIVE_RUNS (worker bookkeeping).
channel) or ACTIVE_RUNS (worker bookkeeping) — except that a
``phase="cancelling"`` run is excluded on both paths, because the worker may
still be unwinding while the client already reached a terminal state for
that stream.
"""
stream_id = getattr(session, 'active_stream_id', None)
if not stream_id:
return None
try:
from api import config as _cfg
with _cfg.ACTIVE_RUNS_LOCK:
_active_run_present = stream_id in (_cfg.ACTIVE_RUNS or {})
_active_run_entry = (_cfg.ACTIVE_RUNS or {}).get(stream_id)
if _active_run_present and not _cfg.active_run_is_attachable(_active_run_entry):
return None
with _cfg.STREAMS_LOCK:
if stream_id in _cfg.STREAMS:
return stream_id
with _cfg.ACTIVE_RUNS_LOCK:
if stream_id in (_cfg.ACTIVE_RUNS or {}):
return stream_id
if _active_run_present:
return stream_id
except Exception:
# On any introspection failure, fail SAFE (report no live stream) rather
# than surfacing a possibly-stale id.
@@ -112,16 +615,21 @@ def truncate_context_for_display_keep(
context_messages: list | None,
full_messages: list | None,
keep: int,
*,
return_boundary_index: bool = False,
) -> list:
"""Align model context with display prefix ``full_messages[:keep]``."""
def _result(rows, boundary_index=None):
return (rows, boundary_index) if return_boundary_index else rows
if keep <= 0:
return []
return _result([])
ctx = context_messages if isinstance(context_messages, list) else []
msgs = full_messages if isinstance(full_messages, list) else []
if not ctx:
return []
return _result([])
if len(msgs) == 0:
return []
return _result([])
# Only the perfectly-parallel case (display and context row-for-row) can be
# sliced at the raw display index. When the two arrays differ in length —
# in EITHER direction — they have diverged and need alignment:
@@ -138,7 +646,7 @@ def truncate_context_for_display_keep(
# strips unanswered tool_calls; gateway: it forwards no tool_calls/tool rows
# at all), so we do not re-do that trimming here.
if len(ctx) == len(msgs):
return ctx[:keep]
return _result(ctx[:keep], min(keep, len(ctx)) - 1)
def _row_signature(row: Any) -> tuple[str, ...] | None:
if not isinstance(row, dict):
@@ -367,8 +875,8 @@ def truncate_context_for_display_keep(
and isinstance(msgs[keep - 1], dict)
and msgs[keep - 1].get('role') == 'user'
):
return ctx[:last_kept + 1]
return ctx[:first_unkept]
return _result(ctx[:last_kept + 1], last_kept)
return _result(ctx[:first_unkept], first_unkept - 1)
if last_kept is not None:
ambiguous_first_unkept = ambiguous_matches[keep]
if (
@@ -376,8 +884,8 @@ def truncate_context_for_display_keep(
and isinstance(msgs[keep - 1], dict)
and msgs[keep - 1].get('role') != 'user'
):
return ctx[:ambiguous_first_unkept]
return ctx[:last_kept + 1]
return _result(ctx[:ambiguous_first_unkept], ambiguous_first_unkept - 1)
return _result(ctx[:last_kept + 1], last_kept)
# Both boundary rows were ambiguous/unmatched (common in large sessions
# where context rows have lost their id/timestamp so the matcher can't
@@ -396,14 +904,15 @@ def truncate_context_for_display_keep(
for i in range(keep - 1, -1, -1):
resolved = matches[i] if matches[i] is not None else ambiguous_matches[i]
if resolved is not None:
return ctx[:resolved + 1]
return _result(ctx[:resolved + 1], resolved)
# Final fallback preserves #5096 behavior when alignment is unreliable
# (no display row resolved to a context index, or keep >= len(msgs)).
prefix_len = max(0, len(ctx) - len(msgs))
prefix = ctx[:prefix_len]
suffix = ctx[prefix_len:]
return prefix + suffix[:keep]
result = prefix + suffix[:keep]
return _result(result, len(result) - 1 if result else None)
def truncate_session_at_keep(session, keep: int) -> tuple[int, int]:
@@ -611,7 +1120,18 @@ def _extract_text(content: Any) -> str:
if isinstance(content, list):
parts = []
for p in content:
if isinstance(p, dict) and p.get('type') == 'text':
parts.append(p.get('text', ''))
if not isinstance(p, dict):
continue
part_type = str(p.get('type') or '').lower()
if part_type not in ('', 'text', 'input_text', 'output_text'):
continue
part_text = (
p.get('text')
or p.get('content')
or p.get('input_text')
or p.get('output_text')
or ''
)
parts.append(str(part_text))
return ' '.join(parts)
return str(content)
+14 -5
View File
@@ -205,9 +205,11 @@ def sync_session_title(session_id: str, title: str, profile: Optional[str] = Non
titles for WebUI sessions. This function bridges that gap and is called
from the background title update/refresh paths after a title is persisted.
Uses ``set_auto_title_if_empty`` so it will only populate a NULL title and
never overwrite a manual rename made via CLI/Gateway/TUI. This means
title refreshes (where state.db already holds the initial auto-title) are
Uses ``set_auto_title`` (LLM provenance) so it will only populate a row that
is NULL or holds a lower-authority auto-title, and never overwrites a manual
rename made via CLI/Gateway/TUI (``set_auto_title`` returns ``False``,
untouched, when a higher-authority title holds the row). This means title
refreshes (where state.db already holds the initial auto-title) are
effectively no-ops at the state.db layer -- acceptable because the primary
goal is ensuring ``hermes sessions list`` is not blank.
@@ -223,8 +225,15 @@ def sync_session_title(session_id: str, title: str, profile: Optional[str] = Non
try:
# Ensure the session row exists (idempotent) so the UPDATE has a target.
db.ensure_session(session_id=session_id, source='webui')
# hermes-agent's SessionDB.set_auto_title_if_empty was renamed to
# set_auto_title(session_id, title, *, source) in the state-module
# split (agent commit 53db597201, released v2026.9.7). set_auto_title
# preserves the same "only populate NULL / never clobber a manual
# rename" semantics (returns False, untouched, when a higher-authority
# title holds the row) and requires an explicit auto source.
_llm_source = getattr(db, "TITLE_SOURCE_LLM", "llm")
try:
db.set_auto_title_if_empty(session_id, title)
db.set_auto_title(session_id, title, source=_llm_source)
except ValueError:
# state.db enforces uniqueness on sessions.title, so a byte-identical
# auto-title generated for two sessions raises ValueError here. Derive
@@ -232,7 +241,7 @@ def sync_session_title(session_id: str, title: str, profile: Optional[str] = Non
# retry instead of leaving the second row blank (#6964).
alt = db.get_next_title_in_lineage(title)
if alt and alt != title:
db.set_auto_title_if_empty(session_id, alt)
db.set_auto_title(session_id, alt, source=_llm_source)
except Exception:
logger.debug("Failed to sync session title to state.db for %s", session_id)
finally:
+1454 -277
View File
File diff suppressed because it is too large Load Diff
+18
View File
@@ -0,0 +1,18 @@
"""Dependency-light helpers for launching child processes consistently."""
from __future__ import annotations
import subprocess
import sys
def windows_hide_flags() -> int:
"""Hide a short-lived console child on Win32 and remain a POSIX no-op.
``CREATE_NO_WINDOW`` keeps captured stdout and stderr connected, unlike
detaching the process. Passing ``0`` elsewhere preserves the subprocess
default. See #5692.
"""
if sys.platform == "win32":
return getattr(subprocess, "CREATE_NO_WINDOW", 0)
return 0
+59 -70
View File
@@ -28,6 +28,7 @@ from api.agent_health import get_active_profile_gateway_running_pid
from api.gateway_restart import restart_active_profile_gateway
from api.profiles import get_active_profile_name
from api.config import REPO_ROOT, STREAMS, STREAMS_LOCK
from api.subprocess_utils import windows_hide_flags
logger = logging.getLogger(__name__)
@@ -75,6 +76,29 @@ _GIT_LOCK_SIGNATURES = (
'another git process seems to be running',
'unable to create .git/index.lock',
)
def _windows_restart_spawn(args, **kwargs):
"""Spawn the replacement process for a Windows self-restart."""
return subprocess.Popen(args, **kwargs)
def _windows_restart_exit(code):
"""Exit the old process after a Windows replacement is running."""
os._exit(code)
def _windows_restart_command():
"""Return the canonical replacement command for the current packaging mode."""
if getattr(sys, "frozen", False):
return list(sys.argv)
executable = sys.executable
if executable.lower().endswith("python.exe"):
windowless_executable = executable[:-4] + "w.exe"
if os.path.isfile(windowless_executable):
executable = windowless_executable
return [executable, str(REPO_ROOT / "server.py")]
# Lock files we previously enumerated for auto-removal in v2. v2.2 no longer
# removes anything on the server, so the enumerable list is no longer needed;
# ``_inventory_locks`` reports whatever ``.git/**/*.lock`` files currently exist
@@ -215,9 +239,14 @@ def _run_git(args, cwd, timeout=10):
return 'git executable not found', False
try:
r = subprocess.run(
[git_executable] + args, cwd=str(cwd), capture_output=True,
text=True, timeout=timeout,
encoding='utf-8', errors='replace',
[git_executable] + args,
cwd=str(cwd),
capture_output=True,
text=True,
timeout=timeout,
encoding='utf-8',
errors='replace',
creationflags=windows_hide_flags(),
)
# On non-UTF-8 locales (e.g. Chinese Windows GBK), a binary git
# output that fails to decode used to leave r.stdout = None and crash
@@ -1711,9 +1740,9 @@ def _schedule_restart(delay: float = 2.0) -> None:
loaded on the next request, rather than running with a mix of old and
new Python modules in sys.modules.
os.execv() replaces the current process image with a fresh interpreter
running the same argv — sessions are preserved on disk, the HTTP port
is reclaimed within the delay window, and the client's own
The restart replaces the current process image or starts the canonical
server entrypoint, depending on platform and packaging mode. Sessions are
preserved on disk, the HTTP port is reclaimed within the delay window, and the client's own
``setTimeout(() => location.reload(), 2500)`` lands after the restart.
Coordinates with ``_apply_lock``: when the user updates both webui
@@ -1751,84 +1780,47 @@ def _schedule_restart(delay: float = 2.0) -> None:
try:
# Re-exec into the just-pulled image.
#
# sys.argv[0]'s meaning depends on how the server was launched:
#
# * Source checkout (`python server.py` via bootstrap.py /
# ctl.sh / start.sh): sys.argv[0] is the SCRIPT path
# (e.g. "/root/hermes-webui/server.py"), sys.executable is
# the interpreter. CPython treats argv[1] as the script to
# run, so we must pass [sys.executable] + sys.argv.
#
# * Frozen/packaged build (PyInstaller, embedded zipapp,
# etc.): sys.argv[0] == sys.executable == <binary>. Passing
# [sys.executable] + sys.argv would re-insert the binary as
# argv[1] — the kernel launches it, the interpreter treats
# the binary itself as the "script" to run, and execv
# effectively becomes a recursive no-op that never reaches
# bind(), leaving the WebUI stuck "offline" after every
# self-update. Pass argv as-is instead.
#
# Distinguish the two cases with sys.frozen (set by
# PyInstaller / zipapp / similar). For source checkouts the
# `[sys.executable] + sys.argv` form is the canonical CPython
# re-exec idiom (same shape Flask/Django reloaders use) and
# is the correct path.
#
# IMPORTANT: On Windows, os.execv() does NOT replace the
# current process — it spawns a new process while the old
# one keeps running. This causes "address already in use"
# because the old process still holds the port. On Windows
# we use subprocess.Popen() + os._exit() instead.
# we use a detached spawn + exit instead.
if sys.platform == 'win32':
import subprocess
if getattr(sys, "frozen", False):
args = sys.argv
else:
args = [sys.executable] + sys.argv
# Prefer pythonw.exe over python.exe so the restarted
# server does not create a visible console window.
# sys.executable may point at python.exe (console
# subsystem); substitute pythonw.exe if it exists
# next to python.exe.
_exe = sys.executable
if _exe.lower().endswith('python.exe'):
_w_exe = _exe[:-4] + 'w.exe' # python.exe -> pythonw.exe
if os.path.isfile(_w_exe):
if getattr(sys, "frozen", False):
args = sys.argv
else:
args = [_w_exe] + sys.argv
args = _windows_restart_command()
# Start new process fully detached with NO console
# window. DETACHED_PROCESS alone is not sufficient
# on modern Windows — without CREATE_NO_WINDOW a
# python.exe (console-subsystem) child still flashes
# an empty terminal window, which the user then
# manually kills (taking the WebUI with it).
subprocess.Popen(
args,
cwd=os.getcwd(),
creationflags=(
subprocess.DETACHED_PROCESS
| subprocess.CREATE_NEW_PROCESS_GROUP
| subprocess.CREATE_NO_WINDOW
),
close_fds=True,
stdin=subprocess.DEVNULL,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
)
try:
_windows_restart_spawn(
args,
cwd=os.getcwd(),
creationflags=(
subprocess.DETACHED_PROCESS
| subprocess.CREATE_NEW_PROCESS_GROUP
| subprocess.CREATE_NO_WINDOW
),
close_fds=True,
stdin=subprocess.DEVNULL,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
)
except Exception:
logger.exception("Windows WebUI restart spawn failed")
return
# Exit immediately — the port is released as soon as
# this process dies, allowing the new process to bind.
os._exit(0)
_windows_restart_exit(0)
else:
if getattr(sys, "frozen", False):
os.execv(sys.executable, sys.argv)
else:
os.execv(sys.executable, [sys.executable] + sys.argv)
except Exception:
# Last-resort: if execv fails for any reason, just exit so the
# process supervisor (start.sh / Docker) restarts us.
os._exit(0)
# Last-resort: let the process supervisor restart us.
_windows_restart_exit(0)
threading.Thread(target=_do, daemon=True).start()
@@ -2400,11 +2392,8 @@ def _apply_update_inner(target, channel=DEFAULT_UPDATE_CHANNEL):
}
# Schedule a self-restart so the updated code is loaded fresh. A plain
# git pull leaves stale Python modules in sys.modules — agent imports that
# reference new symbols (functions, classes) added in the update will fail
# on the next request with AttributeError / ImportError. os.execv() re-
# execs the same interpreter with the same argv, picking up the new code
# cleanly without requiring the user to restart manually.
# git pull leaves stale Python modules in sys.modules. Replacing the process
# loads the updated code cleanly without requiring a manual restart.
#
# The 2 s delay gives the HTTP response time to flush to the client before
# the process replaces itself. The client already does
+23 -6
View File
@@ -122,8 +122,8 @@ def _attachment_root() -> Path:
return (STATE_DIR / 'attachments').resolve()
def _upload_destination(session_id: str, safe_name: str) -> Path:
dest_dir = _session_attachment_dir(session_id)
def _upload_destination(session_id: str, safe_name: str, dest_dir: Path | None = None) -> Path:
dest_dir = dest_dir if dest_dir is not None else _session_attachment_dir(session_id)
dest_dir.mkdir(parents=True, exist_ok=True)
dest = (dest_dir / safe_name).resolve()
if not dest.is_relative_to(dest_dir):
@@ -224,8 +224,20 @@ def handle_upload(handler):
if _reject_invisible_session(handler, s):
return True
safe_name = _sanitize_upload_name(filename)
dest = _upload_destination(session_id, safe_name)
dest.write_bytes(file_bytes)
dest_dir = _session_attachment_dir(session_id)
dest = _upload_destination(session_id, safe_name, dest_dir)
# #3398-style TOCTOU hardening, mirrored from the workspace upload
# path: O_CREAT|O_EXCL|O_NOFOLLOW anchored open so a raced duplicate
# cannot be silently overwritten and a raced symlink subpath cannot
# redirect the write outside the attachment dir.
try:
_wfd = open_anchored_create_fd(dest_dir, dest)
except FileExistsError:
return j(handler, {'error': f'Upload destination already exists: {safe_name}'}, status=409)
except (ValueError, OSError):
return j(handler, {'error': 'Upload destination rejected'}, status=403)
with os.fdopen(_wfd, 'wb', closefd=True) as _wfh:
_wfh.write(file_bytes)
mime = mimetypes.guess_type(safe_name)[0] or 'application/octet-stream'
return j(handler, {
'filename': dest.name,
@@ -623,8 +635,13 @@ def handle_workspace_upload(handler):
if _reject_invisible_session(handler, session):
return True
# Resolve workspace root from session
workspace = resolve_trusted_workspace(session.workspace)
# Resolve workspace root using the session profile, not the ambient request profile.
try:
workspace = resolve_trusted_workspace(
session.workspace, profile=getattr(session, "profile", None)
)
except TypeError:
workspace = resolve_trusted_workspace(session.workspace)
# Resolve target subdirectory within workspace
target_dir = safe_resolve_ws(workspace, subpath) if subpath else workspace
+320 -74
View File
@@ -12,6 +12,7 @@ import json
import logging
import os
import posixpath
import re
import secrets
import shutil
import stat
@@ -35,12 +36,19 @@ from api.config import (
DEFAULT_WORKSPACE as _BOOT_DEFAULT_WORKSPACE,
MAX_FILE_BYTES, IMAGE_EXTS, MD_EXTS
)
from api.subprocess_utils import windows_hide_flags
# ── Profile-aware path resolution ───────────────────────────────────────────
def _profile_state_dir() -> Path:
"""Return the webui_state directory for the active profile.
# Logical profile-name grammar — mirrors api.profiles._PROFILE_ID_RE. Kept as
# a local copy so workspace-layer validation does not import profiles at module
# load time (profiles may not be importable in every embedding context).
_PROFILE_NAME_RE = re.compile(r'^[a-z0-9][a-z0-9_-]{0,63}$')
def _profile_state_dir(profile: str | Path | None = None) -> Path:
"""Return the webui_state directory for the active or given profile.
For the default profile, returns the global STATE_DIR (respects
HERMES_WEBUI_STATE_DIR env var for test isolation).
@@ -48,6 +56,28 @@ def _profile_state_dir() -> Path:
"""
try:
from api.profiles import get_active_profile_name, get_active_hermes_home
if profile is not None:
# Literal-"default" STATE routing (#7168 re-gate round 7): the
# default profile's workspace state always lives in the global
# state files, even when isolated mode pins the default home at
# <base>/profiles/default. The round-6 resolver change made
# _resolve_profile_home_param("default") return that pinned home,
# so the canonical-home check below sent explicit
# profile="default" state reads/writes to {pinned}/webui_state/
# while ambient calls kept using the global dir — splitting saved
# workspaces between two authorities. Config and workspace PATH
# resolution still use the pinned home via
# _resolve_profile_home_param; only this state-file tier stays
# global for the logical string "default".
if isinstance(profile, str) and profile.strip() == 'default':
return _GLOBAL_WS_FILE.parent
profile_home = _resolve_profile_home_param(profile)
if not _is_default_profile_home(profile_home):
d = profile_home / 'webui_state'
d.mkdir(parents=True, exist_ok=True)
return d
return _GLOBAL_WS_FILE.parent
name = get_active_profile_name()
if name and name != 'default':
d = get_active_hermes_home() / 'webui_state'
@@ -58,14 +88,41 @@ def _profile_state_dir() -> Path:
return _GLOBAL_WS_FILE.parent
def _workspaces_file() -> Path:
"""Return the workspaces.json path for the active profile."""
return _profile_state_dir() / 'workspaces.json'
def _workspaces_file(profile: str | Path | None = None) -> Path:
"""Return the workspaces.json path for the active or given profile."""
return _profile_state_dir(profile=profile) / 'workspaces.json'
def _last_workspace_file() -> Path:
"""Return the last_workspace.txt path for the active profile."""
return _profile_state_dir() / 'last_workspace.txt'
def _last_workspace_file(profile: str | Path | None = None) -> Path:
"""Return the last_workspace.txt path for the active or given profile."""
return _profile_state_dir(profile=profile) / 'last_workspace.txt'
def _workspaces_file_for_profile(profile: str | Path | None = None) -> Path | None:
"""Profile-scoped workspaces.json path, or None for an INVALID profile.
``None`` is the fail-closed contract (#7168 re-gate round 4): a malformed
profile name must not be clamped onto the default/global state files, so
callers treat it as "no readable/writable profile-local state".
"""
try:
return _workspaces_file(profile=profile) if profile is not None else _workspaces_file()
except TypeError:
return _workspaces_file()
except ValueError:
logger.debug("Ignoring invalid profile name %r for workspaces file", profile)
return None
def _last_workspace_file_for_profile(profile: str | Path | None = None) -> Path | None:
"""Profile-scoped last_workspace.txt path, or None for an INVALID profile."""
try:
return _last_workspace_file(profile=profile) if profile is not None else _last_workspace_file()
except TypeError:
return _last_workspace_file()
except ValueError:
logger.debug("Ignoring invalid profile name %r for last-workspace file", profile)
return None
def _expanduser_path(path: str | Path) -> Path:
@@ -100,8 +157,11 @@ def _expanduser_path(path: str | Path) -> Path:
return Path(raw)
def _resolve_path(path: str | Path) -> Path:
"""Resolve *path* after env-aware home expansion, without raising."""
def _resolve_path(path: str | Path, profile: str | Path | None = None) -> Path:
"""Resolve *path* after env-aware home expansion, preserving remote POSIX paths without raising."""
remote_candidate = _remote_terminal_workspace_candidate(path, profile=profile)
if remote_candidate is not None:
return remote_candidate
return _safe_resolve(_expanduser_path(path))
@@ -147,12 +207,75 @@ def _is_remote_terminal_backend(terminal_cfg: dict | None) -> bool:
return backend not in ('', 'local')
def _remote_terminal_cwd() -> str | None:
"""Return target-side terminal cwd for remote profiles, without local stat()."""
def _resolve_profile_home_param(profile: str | Path | None) -> Path:
"""Resolve a profile parameter (name string, directory path string, or Path) to a profile home Path.
Logical profile names are validated strictly (#7168 re-gate round 4): a
name that fails the profile-id grammar raises ValueError instead of being
silently clamped onto the default home — the old clamp let a malformed
name such as ``"bad name"`` read/write the DEFAULT profile's state.
Round 5 tightens the grammar gate: a STRING profile value is strictly a
logical profile id and is NEVER treated as a path-shaped home — the old
``"/" in raw`` branch resolved any slash-bearing string directly, so a
malformed value like ``"../evil"`` bypassed validation entirely and could
read/overwrite an arbitrary ``webui_state/last_workspace.txt``
(#7168 re-gate round 5, path traversal on the profile-isolation boundary).
Round 6 closes the last isolation hole in this resolver: the logical
string ``"default"`` used to short-circuit to ``_DEFAULT_HERMES_HOME``
before reaching ``get_hermes_home_for_profile()``, bypassing the
isolated-mode clamp in ``api.profiles._resolve_profile_home_for_name``
(#7168 re-gate round 6). In an isolated deployment pinned at
``<base>/profiles/default``, a session created with ``profile="default"``
therefore resolved workspace/config from the BASE root home instead of
the pinned one. The name now flows through the same delegated path as
every other logical id; literal-default routing to the global state
files is retained in ``_profile_state_dir``/``get_last_workspace`` via
canonical ``_is_default_profile_home`` identity.
An explicit home directory is expressed as a ``Path`` object (callers such
as streaming's legacy ``_profile_home`` fallback wrap their home strings
in ``Path``); Path values are honored and canonicalized so identity
comparisons never depend on lexical spelling (e.g. a symlink alias of the
default home must compare equal to it).
"""
if profile is None or str(profile).strip() == "":
from api.profiles import get_active_hermes_home
return get_active_hermes_home()
raw = str(profile).strip()
if isinstance(profile, Path):
return _safe_resolve(profile.expanduser())
# Strings are LOGICAL PROFILE IDS ONLY — no path-shaped strings, ever.
if not _PROFILE_NAME_RE.fullmatch(raw):
raise ValueError(f"invalid profile name: {raw!r}")
from api.profiles import get_hermes_home_for_profile
return _safe_resolve(get_hermes_home_for_profile(raw))
def _is_default_profile_home(profile_home: Path) -> bool:
"""Canonical identity check against the root/default Hermes home.
Compares resolved paths so a symlink alias of _DEFAULT_HERMES_HOME is
recognized as the default profile rather than treated as a foreign,
lexically-different directory (#7168 re-gate round 4).
"""
try:
from api.config import get_config
from api.profiles import _DEFAULT_HERMES_HOME
return _safe_resolve(profile_home) == _safe_resolve(_DEFAULT_HERMES_HOME)
except Exception:
return False
def _remote_terminal_cwd(profile: str | Path | None = None) -> str | None:
"""Return target-side terminal cwd for a remote profile, without local stat()."""
try:
from api.config import get_config_for_profile_home
profile_home = _resolve_profile_home_param(profile)
terminal_cfg = get_config_for_profile_home(profile_home).get('terminal', {})
terminal_cfg = get_config().get('terminal', {})
if not _is_remote_terminal_backend(terminal_cfg):
return None
cwd = str(terminal_cfg.get('cwd') or '').strip()
@@ -164,9 +287,19 @@ def _remote_terminal_cwd() -> str | None:
return None
def _remote_terminal_workspace_candidate(path: str | Path) -> Path | None:
"""Return a non-stat'ed target-side Path when it is under terminal.cwd."""
cwd = _remote_terminal_cwd()
def _remote_terminal_workspace_candidate(path: str | Path, profile: str | Path | None = None) -> Path | None:
"""Return a non-stat'ed target-side Path when it is under terminal.cwd for the given profile.
Remote workspace paths live on the target host (e.g., remote SSH/Docker
backend). For valid target-side POSIX paths under ``terminal.cwd``, the
normalized target path is preserved as a ``Path`` object without invoking
local host-filesystem resolution (avoiding host-specific firmlink rewriting
such as macOS synthetic ``/home`` -> ``/System/Volumes/Data/home``).
"""
try:
cwd = _remote_terminal_cwd(profile=profile) if profile is not None else _remote_terminal_cwd()
except TypeError:
cwd = _remote_terminal_cwd()
if not cwd:
return None
raw = _strip_surrounding_quotes(str(path)).strip()
@@ -182,10 +315,10 @@ def _remote_terminal_workspace_candidate(path: str | Path) -> Path | None:
if _is_blocked_workspace_path(Path(normalized_raw), normalized_raw) or _is_blocked_workspace_path(Path(normalized_cwd), normalized_cwd):
return None
if posix_candidate == posix_base or _posix_is_within(posix_candidate, posix_base):
return _resolve_path(normalized_raw)
return Path(normalized_raw)
return None
candidate = _resolve_path(raw)
base = _resolve_path(cwd)
candidate = _safe_resolve(_expanduser_path(raw))
base = _safe_resolve(_expanduser_path(cwd))
if _is_blocked_workspace_path(candidate, raw) or _is_blocked_workspace_path(base, cwd):
return None
if candidate == base or _is_within(candidate, base):
@@ -193,7 +326,7 @@ def _remote_terminal_workspace_candidate(path: str | Path) -> Path | None:
return None
def _profile_default_workspace() -> str:
def _profile_default_workspace(profile: str | Path | None = None) -> str:
"""Read the profile's default workspace from its config.yaml.
Checks keys in priority order:
@@ -209,8 +342,9 @@ def _profile_default_workspace() -> str:
Falls back to the live DEFAULT_WORKSPACE from api.config.
"""
try:
from api.config import get_config
cfg = get_config()
from api.config import get_config_for_profile_home
profile_home = _resolve_profile_home_param(profile)
cfg = get_config_for_profile_home(profile_home)
terminal_cfg = cfg.get('terminal', {})
remote_terminal = _is_remote_terminal_backend(terminal_cfg)
# Explicit webui workspace keys first
@@ -219,7 +353,7 @@ def _profile_default_workspace() -> str:
if ws:
if remote_terminal:
return str(ws).strip()
p = _resolve_path(str(ws))
p = _resolve_path(str(ws), profile=profile)
if remote_terminal or p.is_dir():
return str(p)
# Fall through to terminal.cwd — the agent's configured working directory
@@ -228,7 +362,7 @@ def _profile_default_workspace() -> str:
if cwd and str(cwd) not in ('.', ''):
if remote_terminal:
return str(cwd).strip()
p = _resolve_path(str(cwd))
p = _resolve_path(str(cwd), profile=profile)
if remote_terminal or p.is_dir():
return str(p)
except (ImportError, Exception):
@@ -236,15 +370,17 @@ def _profile_default_workspace() -> str:
try:
from api.config import DEFAULT_WORKSPACE as _LIVE_DEFAULT_WORKSPACE
return str(_resolve_path(_LIVE_DEFAULT_WORKSPACE))
return str(_resolve_path(_LIVE_DEFAULT_WORKSPACE, profile=profile))
except Exception:
return str(_resolve_path(_BOOT_DEFAULT_WORKSPACE))
return str(_resolve_path(_BOOT_DEFAULT_WORKSPACE, profile=profile))
# ── Public API ──────────────────────────────────────────────────────────────
def _clean_workspace_list(workspaces: list) -> list:
def _clean_workspace_list(workspaces: list, profile: str | Path | None = None) -> list:
"""Sanitize a workspace list:
- Preserve target-side remote terminal workspace paths (SSH/Docker) without
resolving them against the local WebUI host filesystem.
- Preserve saved paths even when they are currently missing or inaccessible;
picker state must not be destroyed by a transient stat/permission failure.
- Remove entries whose paths live inside another profile's directory
@@ -260,7 +396,11 @@ def _clean_workspace_list(workspaces: list) -> list:
name = w.get('name', '')
if not path:
continue
p = _safe_resolve(_expanduser_path(path))
remote_cand = _remote_terminal_workspace_candidate(path, profile=profile)
if remote_cand is not None:
p = remote_cand
else:
p = _safe_resolve(_expanduser_path(path))
# Skip paths inside a DIFFERENT profile's directory (cross-profile leak).
# Allow paths inside the CURRENT profile's own directory (e.g. test workspaces
# created under ~/.hermes/profiles/webui/webui-mvp-test/).
@@ -269,7 +409,14 @@ def _clean_workspace_list(workspaces: list) -> list:
# p is under ~/.hermes/profiles/ — only skip if it's under a DIFFERENT profile
try:
from api.profiles import get_active_hermes_home
own_profile_dir = get_active_hermes_home().resolve()
if profile is not None:
# Explicit profile wins: the list belongs to that profile,
# so "own" is defined by the profile parameter, never by the
# ambient home (loading profile A's list under ambient B must
# not silently drop A's own workspaces).
own_profile_dir = _resolve_profile_home_param(profile).resolve()
else:
own_profile_dir = get_active_hermes_home().resolve()
p.relative_to(own_profile_dir)
# p is under our own profile dir — keep it
except (ValueError, Exception):
@@ -336,12 +483,12 @@ def _migrate_global_workspaces() -> list:
return []
def load_workspaces() -> list:
ws_file = _workspaces_file()
if ws_file.exists():
def load_workspaces(profile: str | Path | None = None) -> list:
ws_file = _workspaces_file_for_profile(profile)
if ws_file is not None and ws_file.exists():
try:
raw = json.loads(ws_file.read_text(encoding='utf-8'))
cleaned = _clean_workspace_list(raw)
cleaned = _clean_workspace_list(raw, profile=profile)
if len(cleaned) != len(raw):
# Persist the cleaned version so stale entries don't keep reappearing
try:
@@ -350,7 +497,7 @@ def load_workspaces() -> list:
)
except Exception:
logger.debug("Failed to persist cleaned workspace list")
return cleaned or [{'path': _profile_default_workspace(), 'name': 'Home'}]
return cleaned or [{'path': _profile_default_workspace(profile=profile), 'name': 'Home'}]
except Exception:
logger.debug("Failed to load workspaces from %s", ws_file)
# No profile-local file yet.
@@ -358,7 +505,11 @@ def load_workspaces() -> list:
# For NAMED profiles: always start clean with just their own workspace.
try:
from api.profiles import get_active_profile_name
is_default = get_active_profile_name() in ('default', None)
if profile is not None:
profile_home = _resolve_profile_home_param(profile)
is_default = _is_default_profile_home(profile_home)
else:
is_default = get_active_profile_name() in ('default', None)
except ImportError:
is_default = True
if is_default:
@@ -366,16 +517,20 @@ def load_workspaces() -> list:
if migrated:
return migrated
# Fresh start: single entry from the profile's configured workspace, labeled "Home"
return [{'path': _profile_default_workspace(), 'name': 'Home'}]
return [{'path': _profile_default_workspace(profile=profile), 'name': 'Home'}]
def save_workspaces(workspaces: list) -> None:
ws_file = _workspaces_file()
def save_workspaces(workspaces: list, profile: str | Path | None = None) -> None:
ws_file = _workspaces_file_for_profile(profile)
if ws_file is None:
# Fail-closed: an invalid profile name must not write any state file
# (it would land in the default profile's directory via the old clamp).
raise ValueError(f"cannot save workspaces for invalid profile {profile!r}")
ws_file.parent.mkdir(parents=True, exist_ok=True)
ws_file.write_text(json.dumps(workspaces, ensure_ascii=False, indent=2), encoding='utf-8')
def get_profile_default_workspace() -> str:
def get_profile_default_workspace(profile: str | Path | None = None) -> str:
"""Resolve the ACTIVE PROFILE's default workspace, never the global file.
Like get_last_workspace() but WITHOUT the global ``_GLOBAL_LW_FILE``
@@ -390,32 +545,38 @@ def get_profile_default_workspace() -> str:
Priority: profile-scoped ``last_workspace.txt`` -> profile ``config.yaml``
``workspace``/``default_workspace`` -> ``terminal.cwd`` -> process default.
"""
remote_cwd = _remote_terminal_cwd()
try:
remote_cwd = _remote_terminal_cwd(profile=profile) if profile is not None else _remote_terminal_cwd()
except TypeError:
remote_cwd = _remote_terminal_cwd()
def _valid(raw: str) -> str | None:
if not raw:
return None
if remote_cwd:
if _remote_terminal_workspace_candidate(raw) is not None:
if _remote_terminal_workspace_candidate(raw, profile=profile) is not None:
return raw
return None
if Path(raw).is_dir():
return raw
return None
lw_file = _last_workspace_file()
if lw_file.exists():
lw_file = _last_workspace_file_for_profile(profile)
if lw_file is not None and lw_file.exists():
try:
p = _valid(lw_file.read_text(encoding='utf-8').strip())
if p:
return p
except Exception:
logger.debug("Failed to read profile last workspace from %s", lw_file)
return _profile_default_workspace()
return _profile_default_workspace(profile=profile)
def get_last_workspace() -> str:
remote_cwd = _remote_terminal_cwd()
def get_last_workspace(profile: str | Path | None = None) -> str:
try:
remote_cwd = _remote_terminal_cwd(profile=profile) if profile is not None else _remote_terminal_cwd()
except TypeError:
remote_cwd = _remote_terminal_cwd()
def valid_last_workspace(raw: str) -> str | None:
if not raw:
@@ -424,35 +585,58 @@ def get_last_workspace() -> str:
# For remote/SSH profiles, last_workspace is target-side state. Do
# not accept stale server-local paths merely because they exist on
# the WebUI host; require the value to stay under terminal.cwd.
if _remote_terminal_workspace_candidate(raw) is not None:
if _remote_terminal_workspace_candidate(raw, profile=profile) is not None:
return raw
return None
if Path(raw).is_dir():
return raw
return None
lw_file = _last_workspace_file()
if lw_file.exists():
lw_file = _last_workspace_file_for_profile(profile)
if lw_file is not None and lw_file.exists():
try:
p = valid_last_workspace(lw_file.read_text(encoding='utf-8').strip())
if p:
return p
except Exception:
logger.debug("Failed to read last workspace from %s", lw_file)
# Fallback: try global file
if _GLOBAL_LW_FILE.exists():
# Fallback: try global file — but ONLY for the root/default profile. A named
# profile must never inherit another profile's last-workspace binding through
# the legacy global state (#7168 re-gate round 3). Identity is canonical:
# a symlink alias of the default home still counts as default, while a
# malformed profile name fails validation and is denied the fallback
# rather than being clamped onto the global file (#7168 re-gate round 4).
_global_fallback_allowed = True # ambient / no explicit profile: historical behavior
if profile is not None:
if str(profile).strip() == 'default':
_global_fallback_allowed = True
else:
try:
_global_fallback_allowed = _is_default_profile_home(
_resolve_profile_home_param(profile)
)
except Exception:
# Conservative default: an unresolvable explicit profile is NOT the
# root profile — deny the global fallback rather than leak.
_global_fallback_allowed = False
if _global_fallback_allowed and _GLOBAL_LW_FILE.exists():
try:
p = valid_last_workspace(_GLOBAL_LW_FILE.read_text(encoding='utf-8').strip())
if p:
return p
except Exception:
logger.debug("Failed to read global last workspace")
return _profile_default_workspace()
return _profile_default_workspace(profile=profile)
def set_last_workspace(path: str) -> None:
def set_last_workspace(path: str, profile: str | Path | None = None) -> None:
try:
lw_file = _last_workspace_file()
lw_file = _last_workspace_file_for_profile(profile)
if lw_file is None:
# Fail-closed: an invalid profile name must not write the default
# profile's last_workspace.txt (#7168 re-gate round 4).
logger.debug("Refusing to set last workspace for invalid profile %r", profile)
return
lw_file.parent.mkdir(parents=True, exist_ok=True)
lw_file.write_text(str(path), encoding='utf-8')
except Exception:
@@ -669,7 +853,15 @@ def _is_within(path: Path, root: Path) -> bool:
return False
def _trusted_workspace_roots() -> list[Path]:
def _trusted_workspace_roots(profile: str | Path | None = None) -> list[Path]:
"""Return the host directories workspace suggestions may traverse.
Saved-workspace roots follow the same trust rule as
:func:`resolve_trusted_workspace`: with an explicit *profile*, only
workspaces saved under THAT profile widen the boundary (plus the ambient
home / boot-default carve-outs); ``None`` keeps the historical ambient /
global saved-list behaviour.
"""
roots: list[Path] = []
def add(candidate: str | Path | None) -> None:
@@ -688,24 +880,26 @@ def _trusted_workspace_roots() -> list[Path]:
add(_home_path())
add(_BOOT_DEFAULT_WORKSPACE)
for w in load_workspaces():
for w in load_workspaces(profile=profile):
add(w.get("path"))
roots.sort(key=lambda p: len(str(p)))
return roots
def list_workspace_suggestions(prefix: str = "", limit: int = 12) -> list[str]:
def list_workspace_suggestions(
prefix: str = "", limit: int = 12, profile: str | Path | None = None
) -> list[str]:
"""Return workspace path suggestions under trusted roots only.
Suggestions are limited to directories under one of:
- Path.home()
- the boot default workspace
- already-saved workspace roots
- already-saved workspace roots (scoped to *profile* when given)
Arbitrary system prefixes return an empty list rather than an error so the
UI can safely autocomplete while the user types.
"""
roots = _trusted_workspace_roots()
roots = _trusted_workspace_roots(profile=profile)
if not roots:
return []
@@ -801,7 +995,7 @@ def list_workspace_suggestions(prefix: str = "", limit: int = 12) -> list[str]:
return suggestions[:limit]
def resolve_trusted_workspace(path: str | Path | None = None) -> Path:
def resolve_trusted_workspace(path: str | Path | None = None, profile: str | Path | None = None) -> Path:
"""Resolve and validate a workspace path.
A path is trusted if it satisfies at least one of:
@@ -823,12 +1017,12 @@ def resolve_trusted_workspace(path: str | Path | None = None) -> Path:
trusted (it was validated at server startup).
"""
if path in (None, ""):
return _resolve_path(_BOOT_DEFAULT_WORKSPACE)
return _resolve_path(_BOOT_DEFAULT_WORKSPACE, profile) if profile is not None else _resolve_path(_BOOT_DEFAULT_WORKSPACE)
candidate = _resolve_path(path)
candidate = _resolve_path(path, profile) if profile is not None else _resolve_path(path)
access_error = _workspace_access_error(candidate)
remote_candidate = _remote_terminal_workspace_candidate(path)
remote_candidate = _remote_terminal_workspace_candidate(path, profile=profile)
if access_error:
# For remote terminal profiles, workspace paths belong to the target
# machine. Allow paths under terminal.cwd so session switching can
@@ -855,6 +1049,11 @@ def resolve_trusted_workspace(path: str | Path | None = None) -> Path:
# (B) Trusted if already in the saved workspace list — covers non-home installs
try:
saved = load_workspaces(profile=profile)
saved_paths = {_resolve_path(w["path"], profile) for w in saved if w.get("path")}
if candidate in saved_paths:
return candidate
except TypeError:
saved = load_workspaces()
saved_paths = {_resolve_path(w["path"]) for w in saved if w.get("path")}
if candidate in saved_paths:
@@ -884,15 +1083,22 @@ def resolve_trusted_workspace(path: str | Path | None = None) -> Path:
def resolve_implicit_workspace_with_recovery(
candidate: str | Path | None,
fallback: str | Path | None | Callable[[], str | Path | None],
profile: str | Path | None = None,
) -> tuple[Path, bool]:
"""Resolve an implicit workspace, recovering only a genuinely missing path.
The fallback still passes through :func:`resolve_trusted_workspace`. Existing
but untrusted, inaccessible, or non-directory candidates are not recovery
cases: their original validation error is preserved so fallback cannot widen
the workspace trust boundary.
the workspace trust boundary. When *profile* is given, both the trust
resolution and the recovery fallback are scoped to that profile.
"""
try:
if profile is not None:
try:
return resolve_trusted_workspace(candidate, profile=profile), False
except TypeError:
pass
return resolve_trusted_workspace(candidate), False
except ValueError as original_error:
if candidate in (None, ""):
@@ -903,19 +1109,49 @@ def resolve_implicit_workspace_with_recovery(
# never prove target-side deletion. Config-read uncertainty also fails
# closed by preserving the original validation error.
try:
from api.config import get_config
from api.config import get_config, get_config_for_profile_home
terminal_cfg = get_config().get("terminal", {})
if profile is not None:
# Classify the backend from THIS profile's own config, never the
# ambient one: a remote profile without terminal.cwd must still be
# recognized as remote when loaded under a different ambient home.
terminal_cfg = get_config_for_profile_home(
_resolve_profile_home_param(profile)
).get("terminal", {})
else:
terminal_cfg = get_config().get("terminal", {})
except Exception:
logger.debug("Failed to classify terminal backend for workspace recovery", exc_info=True)
raise original_error from None
if _is_remote_terminal_backend(terminal_cfg):
raise original_error from None
try:
local_candidate = _resolve_path(candidate)
local_candidate = (
_resolve_path(candidate, profile=profile)
if profile is not None
else _resolve_path(candidate)
)
local_candidate.stat()
except FileNotFoundError:
fallback_value = fallback() if callable(fallback) else fallback
def _profile_bound_fallback():
if profile is not None and callable(fallback):
# Profile-bound getter FIRST (#7168 re-gate round 3): the
# production call sites pass profile-aware getters such as
# get_last_workspace, so binding the explicit profile must
# take precedence over any zero-argument compatibility call,
# which would read ambient/global state.
try:
return fallback(profile)
except TypeError:
pass
return fallback() if callable(fallback) else fallback
fallback_value = _profile_bound_fallback()
if profile is not None:
try:
return resolve_trusted_workspace(fallback_value, profile=profile), True
except TypeError:
pass
return resolve_trusted_workspace(fallback_value), True
except (OSError, RuntimeError, ValueError):
raise original_error from None
@@ -943,7 +1179,7 @@ def _strip_surrounding_quotes(path: str) -> str:
return s
def validate_workspace_to_add(path: str) -> Path:
def validate_workspace_to_add(path: str, profile: str | Path | None = None) -> Path:
"""Validate a path for *adding* to the workspace list (less restrictive than resolve_trusted_workspace).
When a user explicitly adds a new workspace path, we trust their intent — they
@@ -958,10 +1194,10 @@ def validate_workspace_to_add(path: str) -> Path:
and users routinely paste those into the Add Space input.
"""
path = _strip_surrounding_quotes(path)
candidate = _resolve_path(path)
candidate = _resolve_path(path, profile) if profile is not None else _resolve_path(path)
access_error = _workspace_access_error(candidate)
remote_candidate = _remote_terminal_workspace_candidate(path)
remote_candidate = _remote_terminal_workspace_candidate(path, profile=profile)
if access_error:
# Remote terminal profiles validate workspace existence on the target
# machine, not on the WebUI server. Permit target-side paths under
@@ -1016,6 +1252,7 @@ def safe_resolve_ws(root: Path, requested: str) -> Path:
_DIR_FD_OK = os.open in getattr(os, "supports_dir_fd", set())
_O_NOFOLLOW = getattr(os, "O_NOFOLLOW", 0)
_O_DIRECTORY = getattr(os, "O_DIRECTORY", 0)
_O_BINARY = getattr(os, "O_BINARY", 0)
def open_anchored_fd(workspace: Path, target: Path, *, want_dir: bool) -> int:
@@ -1037,7 +1274,11 @@ def open_anchored_fd(workspace: Path, target: Path, *, want_dir: bool) -> int:
# Windows / no openat: fall back to a plain pathname open. No new race
# protection, but no regression vs the prior path-based behaviour, and
# symlink creation needs admin on Windows anyway.
flags = os.O_RDONLY | (_O_DIRECTORY if want_dir else 0) | _O_NOFOLLOW
flags = (
os.O_RDONLY
| (_O_DIRECTORY if want_dir else _O_BINARY)
| _O_NOFOLLOW
)
try:
return os.open(str(target), flags)
except OSError:
@@ -1052,7 +1293,11 @@ def open_anchored_fd(workspace: Path, target: Path, *, want_dir: bool) -> int:
for i, part in enumerate(rel_parts):
is_last = i == len(rel_parts) - 1
want_directory = (not is_last) or want_dir
flags = os.O_RDONLY | _O_NOFOLLOW | (_O_DIRECTORY if want_directory else 0)
flags = (
os.O_RDONLY
| _O_NOFOLLOW
| (_O_DIRECTORY if want_directory else _O_BINARY)
)
try:
nfd = os.open(part, flags, dir_fd=fd)
except OSError:
@@ -1770,6 +2015,7 @@ def _run_git(args, cwd, timeout=3):
r = subprocess.run(
['git'] + args, cwd=str(cwd), capture_output=True,
text=True, timeout=timeout,
creationflags=windows_hide_flags(),
)
return r.stdout.strip() if r.returncode == 0 else None
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
+5 -19
View File
@@ -12,7 +12,6 @@ import logging
import os
import shutil
import subprocess
import sys
import tempfile
import threading
import re
@@ -21,24 +20,11 @@ from dataclasses import dataclass
from pathlib import Path
from typing import Iterable
logger = logging.getLogger(__name__)
def _windows_hide_flags() -> int:
"""Win32 ``creationflags`` that hide a short-lived console child's window
(``CREATE_NO_WINDOW``) without detaching it, so ``capture_output`` still
works. Returns ``0`` on non-Windows — the ``subprocess`` default, a genuine
no-op. Mirrors the ``api/updates.py`` pattern; kept local so workspace-git
never takes a hard dependency on the optional ``hermes_cli`` package (a
standalone/agent-less WebUI must keep full git functionality). See #5692.
"""
if sys.platform == "win32":
return getattr(subprocess, "CREATE_NO_WINDOW", 0)
return 0
from api.subprocess_utils import windows_hide_flags
from api.workspace import rmtree_anchored, safe_resolve_ws, unlink_anchored
logger = logging.getLogger(__name__)
GIT_TIMEOUT = 5
GIT_REMOTE_TIMEOUT = 60
@@ -263,7 +249,7 @@ def _run_git(
text=True,
timeout=timeout,
env=run_env,
creationflags=_windows_hide_flags(),
creationflags=windows_hide_flags(),
)
except subprocess.TimeoutExpired as exc:
raise GitWorkspaceError("Git command timed out", "timeout") from exc
@@ -306,7 +292,7 @@ def _config_names_for_scope(
capture_output=True,
timeout=GIT_TIMEOUT,
env=env,
creationflags=_windows_hide_flags(),
creationflags=windows_hide_flags(),
)
if result.returncode not in {0, 1}:
if ignore_unsupported:
+4
View File
@@ -10,6 +10,8 @@ from pathlib import Path
import logging
from api.subprocess_utils import windows_hide_flags
logger = logging.getLogger(__name__)
@@ -21,6 +23,7 @@ def _run_git(args: list[str], cwd: str | Path, timeout: float = 2) -> subprocess
capture_output=True,
timeout=timeout,
check=False,
creationflags=windows_hide_flags(),
)
@@ -326,6 +329,7 @@ def find_git_repo_root(workspace: str | Path) -> Path:
capture_output=True,
timeout=5,
check=False,
creationflags=windows_hide_flags(),
)
except (OSError, subprocess.TimeoutExpired) as exc:
raise ValueError("Workspace is not inside a git repository") from exc
+13 -2
View File
@@ -3,6 +3,10 @@
# QUICK START:
# docker compose -f docker-compose.three-container.yml up -d
# Open http://localhost:8787 (chat) and http://localhost:9119 (dashboard)
# Remote access (issue #6876): echo "HERMES_DASHBOARD_BIND=0.0.0.0" >> .env
# and echo "API_SERVER_KEY=<16+ chars>" >> .env, then recreate the stack.
# The WebUI reports "Hermes agent is not responding" when the gateway
# listener is down — API_SERVER_KEY must be set for port 8642 to open.
#
# This extends the two-container setup with the Hermes Dashboard for
# monitoring agent activity, sessions, and resource usage.
@@ -39,7 +43,10 @@ services:
container_name: hermes-agent
command: gateway run
ports:
- "127.0.0.1:8642:8642"
# Bind to loopback by default. For remote access (issue #6876), set
# HERMES_AGENT_BIND=0.0.0.0 in .env — the gateway requires
# API_SERVER_KEY (>=16 chars) to open the listener at all.
- "${HERMES_AGENT_BIND:-127.0.0.1}:8642:8642"
volumes:
# Persist config, state, sessions, skills, memory across restarts
- hermes-home:/home/hermes/.hermes
@@ -83,7 +90,11 @@ services:
container_name: hermes-dashboard
command: dashboard --host 0.0.0.0 --insecure
ports:
- "127.0.0.1:9119:9119"
# Loopback-only by default (the dashboard runs --insecure). To reach the
# dashboard from another machine (issue #6876), set
# HERMES_DASHBOARD_BIND=0.0.0.0 in .env and put it behind auth / a
# reverse proxy.
- "${HERMES_DASHBOARD_BIND:-127.0.0.1}:9119:9119"
volumes:
- hermes-home:/home/hermes/.hermes
environment:
+59 -17
View File
@@ -57,17 +57,31 @@ if [ ! -d "$itdir" ]; then error_exit "Failed to create $itdir"; fi
# logic: if not set and file exists, use file value, else use default. Create file for persistence when the container is re-run
# reasoning: needed when using docker compose as the file will exist in the stopped container, and changing the value from environment variables or configuration file must be propagated from the root init phase to the hermeswebui runtime phase
it=$itdir/hermeswebui_user_uid
if [ -z "${WANTED_UID+x}" ]; then
if [ -f $it ]; then WANTED_UID=$(cat $it); fi
# Remember where the value came from. Auto-detection must never overwrite an
# identity the operator chose explicitly — including the legitimate value 1024,
# which is also the built-in default and was therefore indistinguishable from
# "unset" (#7027). The source is persisted next to the value because `su` drops
# the environment when this script re-enters as the runtime user, so an explicit
# choice would otherwise look like a detected one on the second pass.
it_source=$itdir/hermeswebui_user_uid_source
_wanted_uid_source=""
if [ -n "${WANTED_UID+x}" ]; then
_wanted_uid_source=explicit
elif [ -f $it ]; then
WANTED_UID=$(cat $it)
if [ -f $it_source ]; then _wanted_uid_source=$(cat $it_source); fi
fi
# Auto-detect from mounted volumes if still unset (#569, #668).
# Auto-detect from mounted volumes if still unset (#569, #668, #7027).
# On macOS, host UIDs start at 501. Using the wrong UID means the container
# user cannot read the bind-mounted files, making the workspace appear empty.
# In two-container setups (hermes-agent + hermes-webui), the shared hermes-home
# volume may be owned by the agent container's UID — detect from there first.
if [ -z "${WANTED_UID+x}" ] || [ "${WANTED_UID}" = "1024" ]; then
# Priority 1: hermes-home shared volume — covers two-container Zeabur/Compose setups (#668)
for _probe_dir in "/home/hermeswebui/.hermes" "$HERMES_HOME" "/opt/data"; do
if [ "$_wanted_uid_source" != "explicit" ] && { [ -z "${WANTED_UID+x}" ] || [ "${WANTED_UID}" = "1024" ]; }; then
# Priority 1: the configured state dir first (#7027) — in a single-container
# deploy it is a bind mount by definition, so its owner is the host identity
# the container has to match — then the hermes-home shared volume, which
# covers two-container Zeabur/Compose setups (#668)
for _probe_dir in "${HERMES_WEBUI_STATE_DIR:-/app/data}" "/home/hermeswebui/.hermes" "$HERMES_HOME" "/opt/data"; do
if [ -d "$_probe_dir" ]; then
_detected_uid=$(stat -c '%u' "$_probe_dir" 2>/dev/null || echo "")
if [ -n "$_detected_uid" ] && [ "$_detected_uid" != "0" ]; then
@@ -78,8 +92,10 @@ if [ -z "${WANTED_UID+x}" ] || [ "${WANTED_UID}" = "1024" ]; then
fi
done
fi
if [ -z "${WANTED_UID+x}" ] || [ "${WANTED_UID}" = "1024" ]; then
# Priority 2: /workspace bind-mount — the standard single-container mount point
if [ "$_wanted_uid_source" != "explicit" ] && { [ -z "${WANTED_UID+x}" ] || [ "${WANTED_UID}" = "1024" ]; }; then
# Priority 2: /workspace bind-mount — a useful signal when it is actually
# bind-mounted, but a stock image ships /workspace owned by the build-time
# 1024, which says nothing about the host, so it stays below the state dir (#7027)
if [ -d "/workspace" ]; then
_detected_uid=$(stat -c '%u' "/workspace" 2>/dev/null || echo "")
if [ -n "$_detected_uid" ] && [ "$_detected_uid" != "0" ]; then
@@ -90,16 +106,22 @@ if [ -z "${WANTED_UID+x}" ] || [ "${WANTED_UID}" = "1024" ]; then
fi
WANTED_UID=${WANTED_UID:-1024}
write_privtmpfile $it "$WANTED_UID"
write_privtmpfile $it_source "${_wanted_uid_source:-detected}"
echo "-- WANTED_UID: \"${WANTED_UID}\""
it=$itdir/hermeswebui_user_gid
if [ -z "${WANTED_GID+x}" ]; then
if [ -f $it ]; then WANTED_GID=$(cat $it); fi
it_source=$itdir/hermeswebui_user_gid_source
_wanted_gid_source=""
if [ -n "${WANTED_GID+x}" ]; then
_wanted_gid_source=explicit
elif [ -f $it ]; then
WANTED_GID=$(cat $it)
if [ -f $it_source ]; then _wanted_gid_source=$(cat $it_source); fi
fi
# Auto-detect GID from mounted volumes to match (#569, #668)
if [ -z "${WANTED_GID+x}" ] || [ "${WANTED_GID}" = "1024" ]; then
# Priority 1: hermes-home shared volume
for _probe_dir in "/home/hermeswebui/.hermes" "$HERMES_HOME" "/opt/data"; do
# Auto-detect GID from mounted volumes to match (#569, #668, #7027)
if [ "$_wanted_gid_source" != "explicit" ] && { [ -z "${WANTED_GID+x}" ] || [ "${WANTED_GID}" = "1024" ]; }; then
# Priority 1: configured state dir first, then the hermes-home shared volume
for _probe_dir in "${HERMES_WEBUI_STATE_DIR:-/app/data}" "/home/hermeswebui/.hermes" "$HERMES_HOME" "/opt/data"; do
if [ -d "$_probe_dir" ]; then
_detected_gid=$(stat -c '%g' "$_probe_dir" 2>/dev/null || echo "")
if [ -n "$_detected_gid" ] && [ "$_detected_gid" != "0" ]; then
@@ -110,8 +132,8 @@ if [ -z "${WANTED_GID+x}" ] || [ "${WANTED_GID}" = "1024" ]; then
fi
done
fi
if [ -z "${WANTED_GID+x}" ] || [ "${WANTED_GID}" = "1024" ]; then
# Priority 2: /workspace bind-mount
if [ "$_wanted_gid_source" != "explicit" ] && { [ -z "${WANTED_GID+x}" ] || [ "${WANTED_GID}" = "1024" ]; }; then
# Priority 2: /workspace bind-mount — lower priority for the same reason as UID (#7027)
if [ -d "/workspace" ]; then
_detected_gid=$(stat -c '%g' "/workspace" 2>/dev/null || echo "")
if [ -n "$_detected_gid" ] && [ "$_detected_gid" != "0" ]; then
@@ -122,6 +144,7 @@ if [ -z "${WANTED_GID+x}" ] || [ "${WANTED_GID}" = "1024" ]; then
fi
WANTED_GID=${WANTED_GID:-1024}
write_privtmpfile $it "$WANTED_GID"
write_privtmpfile $it_source "${_wanted_gid_source:-detected}"
echo "-- WANTED_GID: \"${WANTED_GID}\""
echo "== Most Environment variables set"
@@ -146,7 +169,26 @@ load_env() {
obfuscate_part="${ENV_OBFUSCATE_PART}"
if [ -f "$tocheck" ]; then
echo "-- Loading environment variables from $tocheck (overwrite existing: $overwrite_if_different) (ignorelist: $ignore_list) (obfuscate: $obfuscate_part)"
while IFS='=' read -r key value; do
# Read whole lines, then split at the FIRST '=' only. Splitting on every
# '=' (IFS='=' read -r key value) makes read drop a single trailing '='
# from values, corrupting secrets such as openssl rand -base64 32 output
# (44 chars ending in '=').
while IFS= read -r key; do
# Skip empty lines and comment lines.
case "$key" in
''|\#*) continue ;;
esac
if [[ "$key" == *=* ]]; then
# Split at the FIRST '=' only. Derive value before shortening key,
# since ${key%%=*} mutates key. Reusing the existing 'key' read
# target (rather than a new 'line' var) avoids silently clobbering a
# user env var named 'line' on the post-su reload.
value=${key#*=}
key=${key%%=*}
else
# Line without '=': whole line is the key with an empty value.
value=
fi
doit=false
# checking if the key is in the ignorelist
for i in $ignore_list; do
+4
View File
@@ -23,6 +23,10 @@ contributor guidance; it does not change runtime behavior or CI gates.
## Runtime, durability, and state contracts
- [`docs/remote-workspaces.md`](remote-workspaces.md):
architecture contract for remote terminal workspaces (SSH/Docker), target-side
POSIX path preservation against macOS synthetic firmlink expansion, and
per-profile isolation boundaries.
- [`docs/rfcs/webui-run-state-consistency-contract.md`](rfcs/webui-run-state-consistency-contract.md):
proposed consistency rules for current WebUI streaming, recovery, replay,
model-context reconstruction, compression, UI scene/cache, and sidebar metadata
+42
View File
@@ -300,6 +300,48 @@ sandbox, isolation mechanism, capability or permission system, or security
boundary. Extensions still execute with the full WebUI session authority
described above.
### Custom Configure editors
Extensions with structured configuration that does not fit scalar
`settings_schema` fields can register one custom editor entry point on their
scoped E0 settings handle:
```js
const ext = window.hermesExt?.register?.("dictionary-manager");
const unregister = ext?.settings?.registerConfigure?.(({ opener, restoreFocus }) => {
openDictionaryManager({ opener, restoreFocus });
});
```
`registerConfigure(handler)` is available only through a valid boot-trusted E0
handle. It does not require `settings_schema` or extension-owned storage. The
first handler registered for an extension wins; duplicates return `null` and do
not replace it. A successful registration returns an idempotent unregister
function.
Core shows one **Configure** button only for the extension's current
effective-enabled row under **Settings → Extensions → Installed**. Diagnostics
never shows the button. Late registration updates an already-mounted Installed
row. Disable or uninstall status removes the entry point immediately; uninstall
followed by a same-ID reinstall cannot revive the old page-local handler until a
full WebUI reload injects and registers the new extension script.
Core enters a pending state before invoking the handler and passes the activated
button as `opener` plus a one-shot `restoreFocus()` callback. The extension owns
its dialog, validation, persistence, and close lifecycle. It must either call
`restoreFocus()` when its UI closes or return a thenable whose settlement means
the Configure UI is closed. Both paths end pending and converge on Core's
one-shot focus restoration. A synchronous non-thenable handler remains disabled
while unsettled. If it never calls `restoreFocus()`, it remains disabled until reload.
Core does not guess dialog lifetime with a timeout.
Synchronous throws, thenable-access failures, and asynchronous rejections are
isolated, logged with the extension ID, and surfaced as a generic failure. Focus
returns to the connected opener when possible, otherwise to the current visible
Configure button or the Installed tab without forcing navigation. The hook does
not let an extension return DOM for Core to render, add a backend settings route,
or change the trusted same-origin extension model.
### Turn lifecycle events
A registered extension can react when a session turn observed by the current page
+32
View File
@@ -55,6 +55,23 @@ Hermes WebUI derives a provisional session title from the first user message
and, after the first response, may call an LLM to generate a better title
(and periodically refresh it for long sessions).
For structured messages containing text and native images, title generation
uses the user text without flattening or modifying the stored message. Title
comparison and title-model inputs remove the internal `[Workspace::v1: ...]`
prefix and one terminal `[Attached files: ...]` or
`[Attached files for this steer: ...]` suffix separated by a blank line.
For structured content, this cleanup applies to the first text part that
provides title content. Literal legacy `[Workspace: ...]` text and later text
parts remain unchanged. Initial generation, explicit regeneration, and adaptive
refresh use this title-specific cleanup.
Background generation requires user text and a substantive assistant response.
It recognizes the sanitized provisional title as well as the existing raw
placeholder, so internal metadata does not make an image-containing turn look
manually titled. Image-only or metadata-only content does not provide title
text. Existing manual-title protection and the title-generation setting still
apply; this cleanup does not rewrite the transcript or native image parts.
Automatic title-generation LLM calls honor the active Hermes profile's
`auxiliary.title_generation.enabled` setting (default: `true`):
@@ -102,6 +119,21 @@ HERMES_WEBUI_GATEWAY_USE_RUNS_API=true \
Use this when the connected gateway advertises approval support and you want tool approval cards to appear in WebUI. Without `HERMES_WEBUI_GATEWAY_USE_RUNS_API=true`, gateway chat stays on the legacy chat-completions transport and approval-capable commands can remain pending in the agent without a WebUI approval card.
When YOLO is enabled for a gateway-backed browser session, WebUI approves every
approval already parked for that session: Runs API prompts are relayed by their
exact `run_id` and mirror token, and local/no-run waiters are all released. It
then automatically answers later Runs API approval requests while the WebUI
session flag remains active. The flag is committed only after every currently
parked remote relay succeeds; a later prompt that races that unconfirmed drain
remains visible instead of being speculatively auto-approved. The handoff is
also shared with local approval admission: a local waiter arriving after
the current drain snapshot waits for the same session handoff and is released
immediately if YOLO has committed, rather than being parked behind an enabled
session. This is client-managed compatibility behavior: the current Runs API has
no session-YOLO toggle, so a request briefly reaches the approval boundary before WebUI answers
it, and Agent-owned policy such as unrestricted computer-use mode is unchanged.
Native API session YOLO is tracked in [Hermes Agent PR #61946](https://github.com/NousResearch/hermes-agent/pull/61946).
`HERMES_WEBUI_CHAT_BACKEND` is intentionally strict: only `gateway`,
`api_server`, or `api-server` enable the bridge. Generic truthy values such as
`1` or `true` are ignored so existing deployments do not change execution
+100 -1
View File
@@ -41,10 +41,109 @@ helpers are genuinely shared code.
| `docker_agent_source_volume` | Compose files and Docker docs expose `hermes-agent-src` and `/opt/hermes` to make the agent checkout visible to WebUI. | Remove the WebUI source mount only after startup install and runtime imports have migrated. This needs Docker/compose follow-up work, not a runtime behavior change in this audit PR. |
| `startup_dependency_install` | `api/startup.py` discovers `HERMES_WEBUI_AGENT_DIR` or `$HERMES_HOME/hermes-agent`; `server.py` calls `auto_install_agent_deps()` after import verification fails; `docker_init.bash` installs from the staged agent source. | Replace source-tree pip installs with a packaged hermes-agent WebUI client plus an agent health/version capability contract. Keep `HERMES_WEBUI_AGENT_DIR` during migration as an override/debug path, but it should stop being required in normal multi-container startup. |
| `runtime_auxiliary_model_metadata` | `api/streaming.py`, `api/routes.py`, `api/config.py`, and `api/providers.py` import `agent.auxiliary_client`, `agent.model_metadata`, `agent.models_dev`, `hermes_cli.models`, and `agent.account_usage`. | Existing provider/model WebUI endpoints can keep serving UI data where they already wrap agent helpers. Missing surfaces need hermes-agent endpoints or a client package for auxiliary task config, text auxiliary calls, context length, token estimate, provider catalog, and account usage. |
| `runtime_session_state` | `api/streaming.py`, `api/goals.py`, and `api/state_sync.py` import `hermes_state.SessionDB` directly. `api/models.py` also opens the active profile's canonical `state.db` for scoped session deletion because the current canonical helper does not preserve branch/compression evidence ahead of inherited delegate metadata or expose retryable artifact-cleanup semantics. | Move cross-container state reads and writes, including destructive session deletion, behind hermes-agent session/state endpoints once the agent API provides equivalent lineage precedence, transaction, and retry-manifest guarantees. WebUI-only presentation state can remain local, but agent session storage should not be opened from the WebUI container. |
| `runtime_session_state` | `api/streaming.py`, `api/goals.py`, and `api/state_sync.py` import `hermes_state.SessionDB` directly. `api/models.py` also reads the `messages` table directly (see [state.db message content encoding](#statedb-message-content-encoding) for the storage-format coupling that creates) and opens the active profile's canonical `state.db` for scoped session deletion because the current canonical helper does not preserve branch/compression evidence ahead of inherited delegate metadata or expose retryable artifact-cleanup semantics. | Move cross-container state reads and writes, including destructive session deletion, behind hermes-agent session/state endpoints once the agent API provides equivalent lineage precedence, transaction, and retry-manifest guarantees. WebUI-only presentation state can remain local, but agent session storage should not be opened from the WebUI container. |
| `runtime_gateway_provider` | `api/streaming.py` and `api/routes.py` import `hermes_cli.runtime_provider`; adapter helpers such as `agent.anthropic_adapter` are also imported for gateway normalization. | Provider resolution, runtime routing, and gateway invocation should be hermes-agent API calls. WebUI can keep request validation and display formatting, but it should not import runtime provider internals from the agent checkout. |
| `webui_local_or_client_package` | WebUI imports `hermes_cli.auth`, `hermes_cli.config`, `hermes_cli.plugins`, `hermes_cli.profiles`, `hermes_cli.goals`, `agent.skill_utils`, `agent.credential_pool`, and `hermes_constants`. | Pure schemas, constants, and parsing helpers can move into a small versioned client/shared package. Privileged data such as credential pools, auth status, profile mutation, plugin discovery, and goal persistence need hermes-agent endpoints. UI-only formatting can remain in WebUI. |
## state.db message content encoding
`api/models.py` reads the agent's `messages` table with its own SQL, so it also
depends on how hermes-agent *encodes* that table, not only on its schema. This
is a storage-format coupling and belongs with the `runtime_session_state`
dependency class above.
`hermes_state` stores list/dict message content (multimodal parts) as a
sentinel-prefixed JSON string, because sqlite3 binds only scalars:
```
_CONTENT_JSON_PREFIX = "\x00json:" # hermes_state.py
```
It provides `_decode_content()` to reverse this. Any WebUI read path that
projects that column must apply an equivalent decode; a raw read hands the
frontend an encoded string that no reader recognises, and an image part's
base64 data URI then renders as literal transcript text.
### WebUI decoding contract
`_decode_state_db_content()` in `api/models.py` is the single decode point. It
is deliberately narrower than the agent's own decoder, because the WebUI can
only accept shapes the rest of its pipeline already renders:
| Input | Result | Why |
| --- | --- | --- |
| Sentinel + list with non-whitespace text and only valid image parts | decoded `list` | `msgContent()` joins the text parts, so the row renders its text |
| Sentinel + image-only list, or text that is empty/whitespace | unchanged string | `msgContent()` discards image parts, so `_messageIsRenderable()` would hide the row with no error |
| Sentinel + list containing a malformed image part | unchanged string | a part must carry a valid per-type payload, not just a matching `type` |
| Sentinel + dict or scalar root | unchanged string | a dict reaches `_getCachedRender()`, and `_renderCacheKey()` calls `text.slice()` on it, blanking the turn |
| Sentinel + `NaN`/`Infinity`/overflowed float | unchanged string | Python emits them, browser `JSON.parse()` rejects the whole `/api/session` payload |
| Sentinel + unsupported part shapes | unchanged string | `input_text`, `output_text`, scalar and unknown parts are dropped by the JS readers, so decoding them would silently lose content that is visible today |
| Anything without the sentinel | unchanged | non-sentinel content is not this contract's concern |
Supported parts are `{"type": "text", "text": <str>}` plus image parts whose
payload validates for their type: `image_url` with a non-empty URL (string or
`{"url": ...}`), `input_image` with a URL or `file_id`, and `image` with a
`base64` source carrying `data` and `media_type` or a `url` source. At least one
text part must contain non-whitespace text.
**Image parts do not render from this projection.** The shared JS readers drop
them, and the state.db projection supplies no `attachments`. Decoding a
text-and-image row shows its text and keeps the base64 payload out of the DOM;
it does not display the image. Rendering images from state.db rows would need a
shared inline-image projection first, at which point image-only lists could be
accepted too.
Widening the accepted schema requires teaching every shared content reader
through one extractor first; until then unsupported shapes must keep falling
back to the raw string.
### Consequences for identity and bounded reads
Decoding changes the runtime type of `content`, so every consumer that derives
an identity from it must agree on one representation:
- Every key -- merge, dedup, content, visible and the fuzzy fallback -- derives
content identity through `_content_identity_for_key()`. Non-list values key
exactly as on master, `str(content or "")`. Non-empty lists get an
**out-of-band** tuple identity, so no message body can compare equal to one:
an in-band string marker would be forgeable by a scalar that contains it.
Two rich turns sharing visible text and timestamp stay distinct when their
images differ.
- Fuzzy duplicate matching is text-only. Structured identities match by exact
identity or not at all, so a rich row can never fuzzy-match a scalar.
- The merge key cache never writes a key component back into message content.
To avoid re-serialising large payloads for every key it instead memoises the
canonical serialisation per content object, scoped to one
`merge_session_messages_append_only()` call.
- The multimodal mirror bridge pairs one rich image-bearing row with one scalar
mirror only. `require_image_parts` and `require_scalar_mirror` are mutually
exclusive so rich-to-rich pairing cannot occur.
- Every read path that projects the `content` column applies the decoder, so
keys derived on one path cannot disagree with keys derived on another. There
are exactly three such call sites:
| Call site | Role |
| --- | --- |
| `_project_state_db_message()` | canonical row projection, shared by the transcript read and the regeneration tail |
| `get_state_db_session_message_keys_before_timestamp()` | bounded prefix keys |
| `get_state_db_regeneration_tail_snapshot()` | regeneration prefix keys |
If prefix keys stayed encoded while the projected tail was decoded, the
prefix/tail collision proof could miss a genuine repeated recovered turn and
`_bounded_tail_snapshot_if_safe` would reject the bounded path, reading the
entire transcript during regeneration.
Two nearby paths deliberately need no decode.
`get_state_db_session_message_prefix_summary()` projects only timestamp
counts and never selects `content`. State-db sidecar reconstruction
(`_sync_sidecar_from_state_db_if_newer()`) sources its rows through
`get_state_db_session_messages()`, so it inherits the canonical decoded
projection rather than reading the column itself.
When session state moves behind hermes-agent endpoints, this decode should move
with it: the agent should return structured content over the API and the WebUI
should stop depending on the sentinel format at all.
## Replacement contract
### Existing endpoint candidates
+39
View File
@@ -468,6 +468,44 @@ services:
Then configure the URL as `http://host.docker.internal:<port>`. Also ensure the host service binds to an address reachable from containers (not only a loopback interface the Docker bridge cannot reach) and that your host firewall allows the connection.
### 9. "Failed to verify state directory" / restart loop on a bind-mounted state dir (#7027)
**Symptom**: Single-container deploy with the state directory bind-mounted from a
host directory, no `WANTED_UID` set. The container exits 1 and restart-loops:
```
-- Auto-detected workspace UID: 1024 (from /workspace)
touch: cannot touch '/app/data/.testfile': Permission denied
!! ERROR: Failed to verify state directory at /app/data
```
**Cause**: UID auto-detection used to read `/workspace` before the configured
state directory. In a stock image `/workspace` exists and is owned by the
image's own build-time `1024:1024`, so detection returned a value that carries
no information about the host — and because `1024` is also the fallback default,
the log read as if detection had found nothing.
**Fix**: Fixed in the init script — the configured `HERMES_WEBUI_STATE_DIR` is
now probed first. The full order is:
1. `$HERMES_WEBUI_STATE_DIR` (default `/app/data`) — a bind mount by definition
in a single-container deploy, so its owner is the host identity to match
2. `/home/hermeswebui/.hermes`, `$HERMES_HOME`, `/opt/data` — the hermes-home
shared volume in two-container setups (#668)
3. `/workspace` — used only when nothing above resolves
4. `1024` — fallback default
Root-owned candidates (UID 0, e.g. a freshly created named volume) are skipped
at every step. An explicitly supplied `WANTED_UID`/`WANTED_GID` always wins and
is never overwritten by detection — including the value `1024`, which earlier
versions treated as "unset".
If you are on an older image, the workaround is to set the IDs explicitly:
```bash
docker run -e WANTED_UID=$(id -u) -e WANTED_GID=$(id -g) ...
```
## Multi-container architecture
The two- and three-container setups use **named Docker volumes** (not bind mounts) by default for a reason: named volumes solve the UID/GID problem by construction. Docker creates the volume's root directory with the correct ownership, all containers reading/writing to it see the same files, no host-side permission setup required.
@@ -588,6 +626,7 @@ volumes:
- #681 — tools running in WebUI container, not agent container (architectural)
- #668 — auto-detect UID/GID from mounted volume
- #569 — UID/GID detection priority order
- #7027 — state dir probed before `/workspace` in UID/GID detection (see [#9 above](#9-failed-to-verify-state-directory--restart-loop-on-a-bind-mounted-state-dir-7027))
If you hit a new failure mode not covered here, please [open an issue](https://github.com/nesquena/hermes-webui/issues/new) with:
+39 -5
View File
@@ -34,19 +34,53 @@ The Hermes Web UI is fully responsive with a mobile-optimized layout
(hamburger sidebar, sidebar top tabs in the drawer, touch-friendly controls),
so it works well as a daily-driver agent interface from your phone.
**Setup:**
**Preferred setup: Tailscale Serve**
1. Install [Tailscale](https://tailscale.com/download) on your server and
your iPhone/Android.
2. Start the WebUI listening on all interfaces with password auth enabled:
2. Keep the WebUI bound to localhost and enable password auth:
```bash
HERMES_WEBUI_PASSWORD=your-secret ./start.sh
```
3. Publish the local WebUI port through Tailscale Serve:
```bash
tailscale serve --bg 8787
```
4. Open the HTTPS MagicDNS URL that Tailscale prints in your phone's browser.
Tailscale Serve keeps WebUI on loopback while giving your tailnet an HTTPS
MagicDNS hostname. On Linux, changing Serve configuration may require elevated
permissions. If `tailscale serve --bg 8787` reports `Access denied: serve
config denied`, either run it with sudo:
```bash
sudo -S -p '' tailscale serve --bg 8787
```
Or allow the supervised non-root WebUI/Hermes user to manage Tailscale:
```bash
sudo -S -p '' tailscale set --operator=$USER
tailscale serve --bg 8787
```
**Fallback: direct tailnet IP**
Use direct tailnet access when Tailscale Serve is unavailable, disabled, or not
permitted. Because this binds WebUI beyond loopback, always enable password
auth:
```bash
HERMES_WEBUI_HOST=0.0.0.0 HERMES_WEBUI_PASSWORD=your-secret ./start.sh
```
3. Open `http://<server-tailscale-ip>:8787` in your phone's browser
(find your server's Tailscale IP in the Tailscale app or with
`tailscale ip -4` on the server).
Then open `http://<server-tailscale-ip>:8787` in your phone's browser (find
your server's Tailscale IP in the Tailscale app or with `tailscale ip -4` on
the server).
That's it. Traffic is encrypted end-to-end by WireGuard, and password auth
protects the UI at the application level. You can add it to your home screen
+29
View File
@@ -0,0 +1,29 @@
# Remote Terminal Workspaces
Architecture contract and path-resolution semantics for remote terminal profiles (SSH, Docker) in Hermes WebUI.
---
## 1. Overview
When a Hermes profile is configured with a remote terminal backend (e.g. `terminal.backend: "ssh"` or `"docker"`), its working directory (`terminal.cwd`) lives on the remote target host rather than the local WebUI host filesystem.
On hosts such as macOS, local path resolution via `Path.resolve()` or `os.path.realpath()` expands synthetic firmlinks (e.g. rewriting `/home/<user>` to `/System/Volumes/Data/home/<user>`). Because the target-side path does not exist on the local macOS server, unconstrained local resolution causes runtime validation failures and session corruption.
---
## 2. Path Resolution Contract
1. **Target-side POSIX Path Preservation:**
- Any valid POSIX path at or beneath a remote profile's configured `terminal.cwd` is preserved verbatim as a `Path` object without calling host-local `Path.resolve()`.
- Traversal escapes (`..`), null bytes (`\0`), and blocked system roots (`/etc`, `/usr`, `/var`, `/sys`, `/proc`, `/dev`) continue to be strictly rejected.
2. **Profile Isolation & Boundary Enforcement:**
- Remote path recognition is **strictly scoped to the target/active profile**.
- If an active profile has a local backend, target-side remote paths belonging to inactive named profiles are **not** treated as remote for the active local profile.
- For local profiles, workspace addition (`/api/workspaces/add`) enforces host-local directory existence and permissions.
- For remote profiles, workspace addition validates against the profile's remote `terminal.cwd` and skips host-local directory creation (`mkdir`).
3. **Session & Streaming Lifecycle:**
- `Session.__init__` and `Session.load` normalize session workspace paths through `_resolve_path(..., profile=profile)`, preserving target-side paths for remote profiles and preventing corruption of `session.workspace` or `session.created_workspace`.
- Streaming execution (`_run_agent_streaming`) and multimodal asset root resolution preserve the remote session workspace when updating runtime run state.
@@ -3,7 +3,7 @@
- **Status:** Proposed
- **Author:** @franksong2702
- **Created:** 2026-05-16
- **Updated:** 2026-07-16
- **Updated:** 2026-08-22
- **Tracking issue:** [#2361](https://github.com/nesquena/hermes-webui/issues/2361)
- **Related architecture:** [#1925](https://github.com/nesquena/hermes-webui/issues/1925), [`hermes-run-adapter-contract.md`](hermes-run-adapter-contract.md), [`stable-assistant-turn-anchors.md`](stable-assistant-turn-anchors.md)
@@ -70,6 +70,7 @@ and 5; it does not mark every run-state boundary implemented.
| Model context / `context_messages` | Supplies conversation state to the agent | Must include the current visible user turn unless deliberately excluded with a user-visible reason | Let the agent resume from context that contradicts what the user can see |
| Pending turn metadata | Bridges submitted-but-not-yet-finalized user input | Must identify the user turn and stream that own active work | Become a permanent duplicate transcript row after recovery |
| Live stream / SSE | Delivers active runtime events to the browser | Must remain an observation path, not the only durable truth for already-emitted events | Lose the visible scene on refresh, reconnect, or session switch |
| Worker lifecycle registry (`ACTIVE_RUNS`) | Tracks whether a worker still occupies the session, so a successor turn cannot start on top of it | Broader than "attachable UI work": a cancelled worker stays registered while it unwinds | Be read directly as the set of runs a browser may attach to |
| Run journal / replay | Rebuilds emitted runtime events after reconnect or restart | Must be cursor-safe and idempotent | Duplicate assistant text, thinking text, tool cards, or compression cards |
| Compression summary / handoff | Gives the agent recovery context after automatic compression | Must remain agent-facing recovery material unless explicitly rendered as history | Pollute the active turn or become implicit current user intent |
| Live UI scene/cache | Preserves expanded rows, in-progress cards, local scroll, and transient grouping | May optimize presentation but must be rebuildable or degradable from transcript/replay | Become the only place where chronological ordering exists |
@@ -98,6 +99,16 @@ and 5; it does not mark every run-state boundary implemented.
browser-facing timeline renderer as live SSE events so recovery does not
downgrade a structured Thinking / progress / tool / compression turn into a
separate flattened presentation.
When session loading combines a WebUI sidecar with Hermes Agent `state.db`, a
native-image user turn may appear as both rich multipart content and scalar
text that replaces each image part with `[screenshot]`. Reconciliation may
treat those rows as one turn only when the multipart value contains text and
recognized native-image parts, its exact scalar projection matches, role and
tool shape match, timestamps match exactly, stable IDs and provider metadata
do not conflict, and the pairing is unambiguous. Keep the rich sidecar row;
if any requirement is missing or contradictory, preserve both rows rather
than deduplicating. Literal scalar `[screenshot]` text alone is not identity
evidence.
Visible interim assistant progress must remain visible timeline content; a
compact Activity disclosure may summarize adjacent tool/debug detail, but it
must not be the only place where the user can see emitted progress text.
@@ -121,6 +132,26 @@ and 5; it does not mark every run-state boundary implemented.
8. **Every mutation names its layer.** A PR touching streaming, recovery,
context reconstruction, compression, replay, or sidebar metadata should state
which layer it changes and what regression proves the invariant still holds.
9. **Lifecycle-busy is not client-attachable.** `ACTIVE_RUNS` answers "may a new
turn start?", not "may a browser attach a renderer?". Cancellation splits the
two: `cancel_stream()` keeps the row as `phase="cancelling"` so a successor
cannot overlap the unwinding worker, but the client has already reached a
terminal state for that stream because its run journal ends in a terminal
event. Recovery paths that hand a stream id to a renderer — session SSE
recovery and hidden-tab status polling — must therefore exclude cancelling
rows, while busy/admission checks must keep counting them. Reading the
registry with a single meaning resurrects a cancelled run on every fresh
subscription: the client attaches, consumes the terminal event, tears the
renderer down, resubscribes, and the loop repeats indefinitely.
Because a cancelling row can otherwise persist forever, cancellation unwind is
bounded: a cancelling row older than that window **and** owning no live
`STREAMS` channel is reclaimed from `ACTIVE_RUNS` along with its stream-owner
entry, so a wedged worker cannot suppress background wakeups permanently.
Reclamation requires both conditions — age alone must not evict a row that
still owns a live channel. Staleness is measured from the cancellation
timestamp (falling back to run start), so a long-running turn cancelled
moments ago is never mistaken for an orphan.
## Review Checklist
@@ -138,6 +169,11 @@ context reconstruction, or session metadata:
interim assistant text, tool cards, compression cards, and terminal states?
- Can this change move a session in the sidebar without meaningful user or
assistant activity?
- Does this change read `ACTIVE_RUNS` for admission ("may a turn start?") or for
attachment ("may a browser render this?"), and does it use the matching
predicate for that question?
- If it introduces or changes a reclamation window, what proves an in-flight
cancellation is not evicted early, and that a wedged one is eventually freed?
- Can automatic compression or recovery text become visible active-turn content?
- What test or manual evidence proves the invariant?
+69
View File
@@ -0,0 +1,69 @@
# SSE streams and capability signaling
Cross-client reference for the server-sent events (SSE) endpoints Hermes WebUI
exposes. Browser and non-browser clients (Android wrapper, CLI observers)
should integrate against this page so every client describes the same
behavior.
All endpoints below are served by the WebUI origin and sit behind the same
authentication as every other `/api/*` route: when a WebUI password or OIDC
is configured, clients must authenticate before opening any stream.
## Endpoint inventory
| Endpoint | Availability | Purpose |
|---|---|---|
| `GET /api/chat/stream?stream_id=<id>` | Always on | Live agent-turn relay (tokens, tool calls, approvals, `done`, `stream_end`). Falls back to run-journal replay when the in-memory stream is gone. Resume cursors: `after_event_id` / `after_seq` query params, with the standard `Last-Event-ID` header as fallback (events carry `id: <stream_id>:<seq>`). An invalid/foreign/ahead-of-stream cursor is honored as replay-from-start rather than silently skipping events. |
| `GET /api/session/stream?session_id=<id>` | Always on | Persistent per-session channel that survives across agent turns (`initial`, `server_turn_started`, `session-updated`, `bg_task_complete`). This is the stream the WebUI frontend keeps open per session and the stream non-browser clients should prefer for background session updates. |
| `GET /api/sessions/events` | Always on | Global session-list invalidation (`sessions_changed` + keepalives). A signal to re-read `/api/sessions`, not a per-session lifecycle feed. |
| `GET /api/sessions/{session_id}/events` | Always on | Per-session run-journal relay with `Last-Event-ID` / `after_event_id` resume and snapshot fallback. See `docs/rfcs/session-sse-contract-v1.md` for the contract and its proof gates. |
| `GET /api/sessions/gateway/stream` | Optional | Real-time updates for CLI/TUI/messaging (agent) sessions merged into the sidebar. Only streams when the **Agent sessions** setting (`show_cli_sessions`) is enabled and the gateway watcher thread is running. |
The authoritative `event:` names on `/api/chat/stream` are listed in the
**Authoritative emitted events** table of
[`docs/rfcs/session-sse-contract-v1.md`](rfcs/session-sse-contract-v1.md).
## Gateway probe scope (important for non-browser clients)
`GET /api/sessions/gateway/stream?probe=1` returns a JSON capability payload
for the **optional gateway stream only** instead of holding an SSE connection:
```json
{
"enabled": false,
"ok": false,
"watcher_running": false,
"fallback_poll_ms": 30000,
"error": "agent sessions not enabled",
"scope": "gateway_sessions",
"session_stream_available": true,
"session_stream_path": "/api/session/stream"
}
```
- `404` + `error: "agent sessions not enabled"` means only that the optional
gateway/agent-sessions stream is disabled on this server.
- `503` + `error: "watcher not started"` means the setting is on but the
gateway watcher thread is not running.
- `200` with `ok: true` means gateway SSE is usable.
**A negative gateway probe result must not be treated as "SSE unavailable".**
The persistent per-session stream (`/api/session/stream`) and the chat-turn
relay (`/api/chat/stream`) are always on and are not gated by the Agent
sessions setting. The `scope`, `session_stream_available`, and
`session_stream_path` fields make this explicit so clients do not need to
infer it from the status code. Clients that only need session updates should
use `/api/session/stream` directly; the gateway probe is only relevant for
clients that display CLI/TUI/messaging sessions.
## Heartbeats and proxy behavior
- All long-lived streams emit SSE keepalive comment lines on the
`_SSE_HEARTBEAT_INTERVAL_SECONDS` cadence (currently 5 seconds), which is
short enough to survive typical reverse-proxy idle timeouts.
- Handlers send `X-Accel-Buffering: no` so nginx-style proxies pass events
through unbuffered.
- Deployments behind buffering proxies that read-until-close (notably
Tornado-based `jupyter-server-proxy`) can set `HERMES_WEBUI_SSE_CHUNKED=1`
to frame each event as an HTTP/1.1 chunk. The default wire format is
unchanged when the flag is unset.
+18 -6
View File
@@ -179,13 +179,23 @@ turn in the exhausted session instead of being blocked with recovery guidance.
## "Hermes Agent was updated while Hermes WebUI was running"
**Symptom.** An action that uses the in-process Agent runtime stops with a message telling you to restart Hermes WebUI. This can happen after `hermes update`, a Git checkout/pull in the Agent source tree, or another tool updates Hermes Agent without restarting the already-running WebUI backend.
**Symptom.** An action that uses the in-process Agent runtime stops with a message telling you to restart Hermes WebUI manually. This can happen after `hermes update`, a Git checkout/pull in the Agent source tree, or another tool updates Hermes Agent without restarting the already-running WebUI backend.
**Why.** WebUI currently imports `run_agent.AIAgent` into its long-lived Python process. Python keeps imported modules in memory. Continuing after a known Agent Git revision changes could combine cached modules from the old revision with source read from the new revision, producing misleading `ImportError`s or inconsistent runtime state. For local Agent-backed chat, WebUI therefore returns a retryable `409 agent_runtime_stale` before claiming or mutating session state instead of attempting a partial in-process reload. Gateway-backed chat runs in the gateway process and is not blocked by this WebUI-local check. Non-Git Agent installs preserve their existing behavior because there is no revision identity to compare.
**Why.** WebUI imports `run_agent.AIAgent` into its long-lived Python process. Continuing after a known Agent Git revision changes could combine cached modules from the old revision with source read from the new revision. Local Agent-backed actions return a retryable `409 agent_runtime_stale` with `restart_scheduled: false` before accepting a new turn. Gateway- and runner-owned chat keep their existing runtime ownership. Non-Git Agent installs preserve their existing behavior because there is no revision identity to compare; losing a previously known revision remains fail-closed.
**Diagnostic.** Compare the running WebUI process start time with the Agent checkout revision and recent update history. If the Agent was updated after WebUI started, restart WebUI before investigating individual missing-symbol errors.
**Diagnostic.** The stale-runtime response includes `agent_update_state`, also preserved in asynchronous compression error status:
**Fix.** Restart using the same launch method that started WebUI:
| Value | Observation |
| --- | --- |
| `active` | A recent Agent update marker names a live PID. |
| `incomplete` | An Agent recovery marker exists in the loaded checkout or configured venv installation. |
| `stale` | The update marker names a dead PID or is older than the diagnostic age limit. |
| `unknown` | Marker contents, PID liveness, or recovery-marker presence cannot be read or classified. |
| `unverified` | No active or recovery marker was found. Update completion and environment health remain unverified. |
These are observations, not success receipts. Hermes Agent removes `.hermes-update-in-progress` on failed and interrupted exits too. A missing or stale marker, or a readable Git revision, does not prove a completed update or a healthy environment. WebUI only reads these markers; it does not remove or repair them.
**Fix.** Check the Agent updater's outcome and resolve any failed or incomplete Agent update first. Once the Agent checkout and environment are healthy and no updater is running, restart WebUI using the same launch method that started it:
```bash
./ctl.sh restart
@@ -193,9 +203,11 @@ turn in the exhausted session instead of being blocked with recovery guidance.
systemctl --user restart hermes-webui.service
```
If you launched `python3 bootstrap.py` in the foreground, stop it with Ctrl-C and start it again. Restarting the whole computer or WSL is not required when restarting the WebUI backend succeeds.
For a foreground `python3 bootstrap.py`, stop it with Ctrl-C and start it again. Restarting the whole computer or WSL is not required when restarting the WebUI backend succeeds. Retry the action after restarting the backend; refreshing the browser alone does not replace its imported Agent modules.
**When to file a bug.** File a WebUI bug if the restart-required message appears even though the Agent revision did not change, or if a clean WebUI restart still produces the same import error. Include the WebUI launch method, WebUI revision, Agent revision, and the sanitized error text.
**Automatic restart prerequisite.** Revision mismatch does not schedule a WebUI restart. Safe automation requires an Agent-owned terminal success receipt bound to the exact update transaction, final revision, and healthy environment, plus an Agent-owned atomic handoff or lease that excludes new mutations across process replacement (or an Agent updater that performs the restart itself). No such public contract is verified for this integration. Repeated readiness checks followed by `os.execv()` leave a race; WebUI's own update lock does not exclude an external Agent updater. Explicit updates initiated through WebUI retain their existing behavior and are outside this revision-mismatch guard.
**When to file a bug.** File a WebUI bug if the restart-required message appears even though the Agent revision did not change or become unreadable, or if a clean WebUI restart still produces the same import error. Include the launch method, WebUI and Agent revisions, the marker diagnostic, and sanitized error text.
---
+18 -1
View File
@@ -7,7 +7,7 @@ Option A rewrite (2026-05-08): imports api.models and api.profiles
directly from the webui codebase, using canonical helpers for
locking, profile scoping, index consistency, and validation.
pip install "mcp<2" # one-time setup
pip install "mcp>=1.28,<2" # one-time setup
python3 mcp_server.py # start via stdio
MCP config for Hermes Agent (add to config.yaml):
@@ -447,6 +447,23 @@ async def handle_move_session(arguments: dict) -> list[TextContent]:
}, ensure_ascii=False, indent=2))]
# ═══════════════════════════════════════════════════════════════════════════
# MCP SDK version guard
# ═════════════════════════════════════════════════════════════════════════==
#
# mcp>=2.0.0 removed the Server.list_tools / Server.call_tool decorator API
# used by this server. Fail fast instead of producing secondary errors.
if not hasattr(Server, "list_tools"): # pragma: no cover
import sys
print(
"ERROR: mcp SDK incompatible. This server requires mcp>=1.28,<2 "
"(the 1.x Python SDK with the decorator API). "
"Run: pip install 'mcp>=1.28,<2'",
file=sys.stderr,
)
sys.exit(1)
# ═══════════════════════════════════════════════════════════════════════════
# MCP Server wiring
# ═══════════════════════════════════════════════════════════════════════════
+8
View File
@@ -21,6 +21,14 @@ dynamic = ["version", "dependencies"]
Repository = "https://github.com/nesquena/hermes-webui"
Issues = "https://github.com/nesquena/hermes-webui/issues"
[project.scripts]
# Packaged CLI entry point (#6739): `hermes-webui` starts the WebUI server, the
# pip-installed equivalent of `python server.py` from a source checkout. The
# generated console script calls bootstrap.main(), which runs launcher-python
# discovery and dep-ensure before launching the server — matching the effective
# behavior of the `python server.py` / `python -m server` launch surface.
hermes-webui = "bootstrap:main"
[tool.setuptools]
include-package-data = true
packages = ["api", "static"]
+1 -1
View File
@@ -8,7 +8,7 @@ pytest-timeout
pytest-asyncio
pytest-shard
ruff
mcp<2
mcp>=1.28,<2
# Optional Office parsers — runtime-optional (commented in requirements.txt), but
# installed here so the workspace Office-preview tests (tests/test_issue540_*.py)
# run under the documented ./scripts/test.sh entry point. These files
+2 -2
View File
@@ -716,9 +716,9 @@ def main() -> None:
try:
signal.signal(signal.SIGTERM, _request_shutdown)
signal.signal(signal.SIGINT, _request_shutdown) # Ctrl-C / ctl.sh daemons (#7078)
except (ValueError, OSError):
# Not on the main thread (e.g. embedded/test harness); skip handler.
logger.debug("Could not install SIGTERM handler", exc_info=True)
logger.debug("Could not install shutdown signal handlers", exc_info=True)
try:
httpd.serve_forever()
+11 -3
View File
@@ -2104,9 +2104,17 @@ $('btnDownload').onclick=()=>{
const a=document.createElement('a');a.href=URL.createObjectURL(blob);
a.download=`hermes-${S.session.session_id}.md`;a.click();URL.revokeObjectURL(a.href);
};
function _buildSessionExportUrl(sessionId,params){
const url=new URL('api/session/export',document.baseURI||location.href);
url.searchParams.set('session_id',String(sessionId||''));
Object.entries(params||{}).forEach(([key,value])=>{
if(value!==undefined&&value!==null)url.searchParams.set(key,String(value));
});
return url.href;
}
$('btnExportJSON').onclick=()=>{
if(!S.session)return;
const url=`/api/session/export?session_id=${encodeURIComponent(S.session.session_id)}`;
const url=_buildSessionExportUrl(S.session.session_id);
const a=document.createElement('a');a.href=url;
a.download=`hermes-${S.session.session_id}.json`;a.click();
};
@@ -2183,7 +2191,7 @@ function exportSessionHTML(session){
// Drop empties so the inlined fallback keeps working for anything we couldn't read.
const clean={};for(const k in palette){if(palette[k])clean[k]=palette[k];}
const paletteB64=btoa(unescape(encodeURIComponent(JSON.stringify(clean))));
const url=`/api/session/export?session_id=${encodeURIComponent(sid)}&format=html&theme=${theme}&palette=${encodeURIComponent(paletteB64)}`;
const url=_buildSessionExportUrl(sid,{format:'html',theme,palette:paletteB64});
const a=document.createElement('a');a.href=url;
a.download=`hermes-${sid}.html`;a.click();
}
@@ -3513,7 +3521,7 @@ window._mirrorSpeechSettingsFromServer=_mirrorSpeechSettingsFromServer;
const _testUpdates=new URLSearchParams(location.search).get('test_updates')==='1';
if(_testUpdates||(_bootSettings.check_for_updates!==false&&!sessionStorage.getItem('hermes-update-checked')&&!sessionStorage.getItem('hermes-update-dismissed'))){
const _checkUrl='api/updates/check'+(_testUpdates?'?simulate=1':'');
api(_checkUrl,{method:_testUpdates?'GET':'POST',body:_testUpdates?undefined:JSON.stringify({force:false})}).then(d=>{if(!_testUpdates)sessionStorage.setItem('hermes-update-checked','1');if((d.webui&&d.webui.behind>0)||(d.agent&&d.agent.behind>0))_showUpdateBanner(d);}).catch(()=>{});
api(_checkUrl,{method:_testUpdates?'GET':'POST',body:_testUpdates?undefined:JSON.stringify({force:false}),timeoutMs:300000}).then(d=>{if(!_testUpdates)sessionStorage.setItem('hermes-update-checked','1');if((d.webui&&d.webui.behind>0)||(d.agent&&d.agent.behind>0))_showUpdateBanner(d);}).catch(()=>{});
}
const _bootActiveProfileUnauthRedirectBudget=(()=>{
const markerKey='hermes-webui-active-profile-bootstrap-401';
+78 -10
View File
@@ -1232,12 +1232,39 @@ async function cmdGoal(args){
if(!S.session||!S.session.session_id){showToast(t('no_active_session'));return;}
const activeSid=S.session.session_id;
try{
// #6703: re-assert the explicit-pick marker on /api/goal the same way
// /api/chat/start does. Without it the server's model resolver treats a
// persisted cross-provider pick as stale and silently reverts the session
// to the profile default mid-session (e.g. while /goal is running).
const _goalModel=S.session.model||($('modelSelect')&&$('modelSelect').value)||'';
const _goalProvider=S.session.model_provider||null;
const _pendingPick=(typeof _readPendingSessionModel==='function')
? _readPendingSessionModel(activeSid)
: null;
const _pendingPickMatch=_pendingPick
&& _pendingPick.model===_goalModel
&& String(_pendingPick.model_provider||'')===String(_goalProvider||'');
const _defaultModel=(typeof window!=='undefined' && window._defaultModel)||'';
const _activeProvider=(typeof window!=='undefined' && window._activeProvider)||null;
const _isCrossProviderPick=_goalModel
&& _goalProvider
&& _defaultModel
&& _activeProvider
&& _goalModel !== _defaultModel
&& String(_goalProvider||'') !== String(_activeProvider||'');
const _explicitPick=(_pendingPickMatch||_isCrossProviderPick)||undefined;
// Do NOT consume the pending explicit-pick marker here: a control-only
// invocation (e.g. /goal status) skips server-side model resolution, so a
// pre-request clear would drop the pick without using it. Consume it below,
// only after a successful kickoff (r.stream_id), re-checking that the stored
// marker still matches the model/provider captured for this kickoff (#6705).
const r=await api('/api/goal',{method:'POST',body:JSON.stringify({
session_id:activeSid,
args:args||'',
workspace:S.session.workspace,
model:S.session.model||($('modelSelect')&&$('modelSelect').value)||'',
model_provider:S.session.model_provider||null,
model:_goalModel,
model_provider:_goalProvider,
explicit_model_pick:_explicitPick,
profile:S.activeProfile||S.session.profile||'default',
})});
const msg = (() => {
@@ -1257,6 +1284,19 @@ async function cmdGoal(args){
showToast(msg.split('\n')[0],2600);
}
if(!r||!r.stream_id)return;
// #6705: consume the one-shot pending explicit-pick marker only after a
// successful kickoff. Re-read the stored marker and clear it only if it
// still matches the model/provider captured above — a control command (no
// stream_id) must leave the marker intact for the next real send, and a
// marker re-recorded mid-flight (newer onchange) must not be clobbered.
if(_pendingPickMatch && typeof _readPendingSessionModel==='function' && typeof _clearPendingSessionModel==='function'){
const _stillPending=_readPendingSessionModel(activeSid);
if(_stillPending
&& _stillPending.model===_goalModel
&& String(_stillPending.model_provider||'')===String(_goalProvider||'')){
_clearPendingSessionModel(activeSid);
}
}
S.toolCalls=[];
if(typeof clearLiveToolCards==='function')clearLiveToolCards();
appendThinking();setBusy(true);
@@ -1901,22 +1941,50 @@ function cmdVoice(){
async function cmdYolo(){
const sid=S.session&&S.session.session_id;
if(!sid){showToast(t('yolo_no_session'));return;}
const generation=_loadSessionGeneration;
const viewIsCurrent=()=>!!(
S.session&&S.session.session_id===sid&&_loadSessionGeneration===generation
);
let approvalOwner=null;
try{
// Check current state first to toggle
// Check current state first to toggle.
const status=await api('/api/session/yolo?session_id='+encodeURIComponent(sid));
if(!viewIsCurrent())return;
const enable=!status.yolo_enabled;
await api('/api/session/yolo',{
// A visible approval must belong to this exact session load before any
// command handler may POST through it. Otherwise fail closed.
const card=$('approvalCard');
if(card&&card.classList.contains('visible')){
approvalOwner=typeof _captureApprovalResponseOwner==='function'
?_captureApprovalResponseOwner()
:null;
if(!approvalOwner)return;
if(enable&&typeof toggleYoloFromApproval==='function'){
await toggleYoloFromApproval();
return;
}
}
const result=await api('/api/session/yolo',{
method:'POST',
body:JSON.stringify({session_id:sid,enabled:enable}),
});
_yoloEnabled=enable;
if(!viewIsCurrent()||(approvalOwner&&!_approvalResponseOwnerIsCurrent(approvalOwner)))return;
const settled=(result&&typeof result.yolo_enabled==='boolean')?result.yolo_enabled:enable;
_yoloEnabled=settled;
_updateYoloPill();
showToast(enable?t('yolo_enabled'):t('yolo_disabled'));
if(enable){
// Dismiss any visible approval card
hideApprovalCard(true);
showToast(settled?t('yolo_enabled'):t('yolo_disabled'));
}catch(e){
if(!viewIsCurrent()||(approvalOwner&&!_approvalResponseOwnerIsCurrent(approvalOwner)))return;
let errorPayload=null;
if(e&&typeof e.body==='string'){
try{errorPayload=JSON.parse(e.body);}catch(_){}
}
}catch(e){showToast('YOLO: '+e.message);}
if(errorPayload&&typeof errorPayload.yolo_enabled==='boolean'){
_yoloEnabled=errorPayload.yolo_enabled;
_updateYoloPill();
}
showToast('YOLO: '+((errorPayload&&(errorPayload.error||errorPayload.message))||e.message));
}
}
// ── Branch / fork command ──
+197 -3
View File
@@ -11,7 +11,12 @@
const registrations=new Map();
const turnLifecycleListeners=new Map();
const turnLifecycleStates=new Map();
const configureRegistrations=new Map();
const configureQuarantinedIds=new Set();
const configureChangeListeners=new Set();
const currentExtensionStatus=new Map();
let trustedSeeded=false;
let extensionStatusSeeded=false;
function extensionId(value){
return String(value||'').trim();
@@ -116,6 +121,7 @@
name:text(entry&&entry.name,id),
storage_owned:storageOwned,
settings_schema:storageOwned?normalizeSchema(entry&&entry.settings_schema):[],
effective_enabled:!(entry&&entry.effective_enabled===false),
});
}
return entries;
@@ -135,6 +141,20 @@
}
trustedSeeded=true;
}
const previousStatus=new Map(currentExtensionStatus);
const nextIds=new Set(entries.map(entry=>entry.id));
if(extensionStatusSeeded){
for(const id of previousStatus.keys()){
if(nextIds.has(id)) continue;
configureQuarantinedIds.add(id);
notifyConfigureChange(id,'quarantine');
}
}
currentExtensionStatus.clear();
for(const entry of entries){
currentExtensionStatus.set(entry.id,{effective_enabled:entry.effective_enabled===true});
}
extensionStatusSeeded=true;
schemas.clear();
for(const entry of entries){
const trusted=trustedExtensions.get(entry.id);
@@ -146,6 +166,14 @@
settings_schema:trusted.storage_owned===true?trusted.settings_schema:[],
});
}
for(const [id,record] of configureRegistrations){
if(!record||!record.active) continue;
const before=previousStatus.get(id);
const after=currentExtensionStatus.get(id);
if(!before||!after||before.effective_enabled!==after.effective_enabled){
notifyConfigureChange(id,'status');
}
}
}
function safeReadState(key){
@@ -233,7 +261,7 @@
return !!(meta&&meta.storage_owned&&Array.isArray(meta.settings_schema)&&meta.settings_schema.length);
}
function settingsAccessor(clean,meta,isTrusted){
function settingsAccessor(clean,meta,isTrusted,allowConfigureRegistration){
const schema=supportsSettings(meta)?meta.settings_schema:[];
const key=settingsKey(clean);
function current(){
@@ -249,7 +277,7 @@
const saved=safeWrite(key,overridesFromValues(schema,checked.values));
return {ok:saved,values:checked.values,errors:saved?{}:{storage:'unavailable'}};
}
return {
const accessor={
extensionId:clean,
trusted:isTrusted,
storageOwned:!!meta.storage_owned,
@@ -278,6 +306,10 @@
return true;
},
};
if(allowConfigureRegistration===true){
accessor.registerConfigure=handler=>registerConfigureHandler(clean,handler);
}
return accessor;
}
function settingsForExtension(id){
@@ -350,6 +382,165 @@
});
}
function notifyConfigureChange(id,reason){
const change=Object.freeze({id,reason});
for(const listener of [...configureChangeListeners]){
try{
listener(change);
}catch(error){
if(typeof console!=='undefined'&&typeof console.error==='function'){
try{console.error('[Hermes extensions] Configure change listener failed:',error);}catch(_loggingError){}
}
}
}
}
function onConfigureChange(listener){
if(typeof listener!=='function') return null;
configureChangeListeners.add(listener);
let active=true;
return function unsubscribe(){
if(!active) return false;
active=false;
configureChangeListeners.delete(listener);
return true;
};
}
function registerConfigureHandler(clean,handler){
if(typeof handler!=='function'||!trustedExtensions.has(clean)||configureQuarantinedIds.has(clean)) return null;
if(configureRegistrations.has(clean)) return null;
const record={handler,pending:false,active:true};
configureRegistrations.set(clean,record);
notifyConfigureChange(clean,'registration');
let active=true;
return function unregister(){
if(!active) return false;
active=false;
record.active=false;
if(configureRegistrations.get(clean)===record) configureRegistrations.delete(clean);
notifyConfigureChange(clean,'registration');
return true;
};
}
function configureStateForExtension(id){
const clean=extensionId(id);
const record=configureRegistrations.get(clean);
const status=currentExtensionStatus.get(clean);
const available=!!(
clean&&record&&record.active&&!configureQuarantinedIds.has(clean)
&&status&&status.effective_enabled===true
);
return {available,pending:available&&record.pending===true};
}
function focusableConfigureTarget(node){
if(!node||typeof node.focus!=='function'||node.isConnected===false||node.hidden===true||node.disabled===true) return false;
if(typeof node.getAttribute==='function'&&node.getAttribute('aria-disabled')==='true') return false;
if(typeof node.closest==='function'&&node.closest('[hidden]')) return false;
return true;
}
function focusConfigureTarget(clean,opener){
const candidates=[];
if(focusableConfigureTarget(opener)) candidates.push(opener);
if(typeof document!=='undefined'&&document){
if(typeof document.querySelectorAll==='function'){
const buttons=document.querySelectorAll('[data-extension-configure-id]');
for(const button of buttons){
if(button&&button.dataset&&button.dataset.extensionConfigureId===clean&&focusableConfigureTarget(button)){
candidates.push(button);
break;
}
}
}
if(typeof document.querySelector==='function'){
const installedTab=document.querySelector('[data-extensions-tab="installed"]');
if(focusableConfigureTarget(installedTab)) candidates.push(installedTab);
}
}
for(const candidate of candidates){
try{
candidate.focus({preventScroll:true});
return true;
}catch(_focusOptionsError){
try{
candidate.focus();
return true;
}catch(_focusError){}
}
}
return false;
}
function reportConfigureFailure(clean,error,onError){
if(typeof console!=='undefined'&&typeof console.error==='function'){
try{console.error(`[Hermes extensions] ${clean} Configure handler failed:`,error);}catch(_loggingError){}
}
if(typeof onError==='function'){
try{onError(error);}catch(callbackError){
if(typeof console!=='undefined'&&typeof console.error==='function'){
try{console.error(`[Hermes extensions] ${clean} Configure failure reporter failed:`,callbackError);}catch(_loggingError){}
}
}
}
}
function invokeConfigure(id,options){
const clean=extensionId(id);
const state=configureStateForExtension(clean);
const record=configureRegistrations.get(clean);
if(!state.available||state.pending||!record) return false;
const opener=options&&options.opener;
const onError=options&&options.onError;
let settled=false;
let failureReported=false;
record.pending=true;
notifyConfigureChange(clean,'pending');
function settle(){
if(settled) return false;
settled=true;
record.pending=false;
notifyConfigureChange(clean,'pending');
focusConfigureTarget(clean,opener);
return true;
}
function fail(error){
if(!failureReported){
failureReported=true;
reportConfigureFailure(clean,error,onError);
}
settle();
}
let result;
try{
result=record.handler(Object.freeze({opener,restoreFocus:settle}));
}catch(error){
fail(error);
return true;
}
let then;
try{
then=result!==null&&(typeof result==='object'||typeof result==='function')?result.then:null;
}catch(error){
fail(error);
return true;
}
if(typeof then==='function'){
try{
then.call(result,settle,fail);
}catch(error){
fail(error);
}
}
return true;
}
function turnLifecycleKey(sessionId,streamId){
return `${sessionId}\u0000${streamId}`;
}
@@ -416,7 +607,7 @@
if(!trusted) return null;
const handle=Object.freeze({
id:clean,
settings:settingsAccessor(clean,trusted,true),
settings:settingsAccessor(clean,trusted,true,true),
storage:storageAccessor(clean,trusted),
events:eventAccessor(clean),
});
@@ -431,6 +622,9 @@
settingsForExtension,
storageForExtension,
_dispatchTurnLifecycle:dispatchTurnLifecycle,
_configureStateForExtension:configureStateForExtension,
_invokeConfigure:invokeConfigure,
_onConfigureChange:onConfigureChange,
resetSettingsForExtension(id){return settingsForExtension(id).reset();},
clearStorageForExtension(id){return storageForExtension(id).clear();},
};
+84 -84
View File
@@ -2235,12 +2235,12 @@ const LOCALES = {
workspace_empty_no_path: 'Nessun workspace selezionato. Imposta un workspace in Impostazioni \u2192 Workspace per esplorare i file.',
workspace_empty_dir: 'Questo workspace è vuoto.',
workspace_show_hidden_files: 'Mostra file nascosti',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Ordina per',
workspace_sort_name_asc: 'Nome (A → Z)',
workspace_sort_name_desc: 'Nome (Z → A)',
workspace_sort_created_desc: 'Data di creazione (più recenti prima)',
workspace_sort_modified_desc: 'Data di modifica (più recenti prima)',
workspace_sort_created_unavailable: 'L\'ora di creazione non è segnalata da questo server o piattaforma.',
workspace_show_hidden_files_desc: "Includi .DS_Store, .git, node_modules e altri file nascosti/di sistema nell'albero dei file.",
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -3994,12 +3994,12 @@ const LOCALES = {
workspace_empty_no_path: 'ワークスペースが選択されていません。設定 → ワークスペースで選択してください。',
workspace_empty_dir: 'このワークスペースは空です。',
workspace_show_hidden_files: '隠しファイルを表示',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: '並べ替え',
workspace_sort_name_asc: '名前 (A → Z)',
workspace_sort_name_desc: '名前 (Z → A)',
workspace_sort_created_desc: '作成日 (新しい順)',
workspace_sort_modified_desc: '更新日 (新しい順)',
workspace_sort_created_unavailable: 'このサーバーまたはプラットフォームでは作成時刻が報告されません。',
workspace_show_hidden_files_desc: '.DS_Store、.git、node_modules などの隠しファイル・システムファイルをファイルツリーに表示します。',
workspace_panel_show: 'ワークスペースパネルを表示',
workspace_panel_hide: 'ワークスペースパネルを非表示',
@@ -5217,9 +5217,9 @@ const LOCALES = {
cron_deliver_label: '出力先',
cron_deliver_local: 'ローカル (出力を保存のみ)',
cron_deliver_custom: 'カスタム配信先',
cron_profile_label: 'プロフィール',
cron_profile_label: 'プロファイル',
cron_profile_server_default: 'サーバーデフォルト',
cron_profile_server_default_hint: '実行時に WebUI サーバーのデフォルトプロフィールを使用します。プロフィールのない既存ジョブはこの従来の動作を維持します。',
cron_profile_server_default_hint: '実行時に WebUI サーバーのデフォルトプロファイルを使用します。プロファイルのない既存ジョブはこの従来の動作を維持します。',
cron_toast_notifications_label: '完了トースト',
cron_toast_notifications_hint: 'この Cron が完了したときにトーストを表示します。オフでもタスクバッジと新規実行マーカーは更新されます。',
cron_toast_notifications_enabled: '有効',
@@ -5664,12 +5664,12 @@ const LOCALES = {
settings_autosave_retry: 'Повторить',
workspace_empty_dir: 'Это рабочее пространство пусто.',
workspace_show_hidden_files: 'Показывать скрытые файлы',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Сортировать по',
workspace_sort_name_asc: 'Имя (A → Z)',
workspace_sort_name_desc: 'Имя (Z → A)',
workspace_sort_created_desc: 'Дата создания (сначала новые)',
workspace_sort_modified_desc: 'Дата изменения (сначала новые)',
workspace_sort_created_unavailable: 'Время создания не сообщается этим сервером или платформой.',
workspace_show_hidden_files_desc: 'Показывать .DS_Store, .git, node_modules и другие скрытые или системные файлы в дереве.',
workspace_panel_show: 'Показать панель рабочей области',
workspace_panel_hide: 'Скрыть панель рабочей области',
@@ -7395,12 +7395,12 @@ const LOCALES = {
workspace_empty_no_path: 'No hay espacio de trabajo seleccionado. Configure un espacio de trabajo en Ajustes \u2192 Workspace para explorar archivos.',
workspace_empty_dir: 'Este espacio de trabajo está vacío.',
workspace_show_hidden_files: 'Mostrar archivos ocultos',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Ordenar por',
workspace_sort_name_asc: 'Nombre (A → Z)',
workspace_sort_name_desc: 'Nombre (Z → A)',
workspace_sort_created_desc: 'Fecha de creación (más recientes primero)',
workspace_sort_modified_desc: 'Fecha de modificación (más recientes primero)',
workspace_sort_created_unavailable: 'Este servidor o plataforma no informa la hora de creación.',
workspace_show_hidden_files_desc: 'Include .DS_Store, .git, node_modules, and other hidden / system files in the file tree.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -9069,12 +9069,12 @@ const LOCALES = {
workspace_empty_no_path: 'Kein Workspace ausgewählt. Wähle einen Workspace unter Einstellungen \u2192 Workspace, um Dateien zu durchsuchen.',
workspace_empty_dir: 'Dieser Workspace ist leer.',
workspace_show_hidden_files: 'Versteckte Dateien anzeigen',
workspace_sort_by: 'Sort by',
workspace_sort_by: 'Sortieren nach',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_created_desc: 'Erstellungsdatum (neueste zuerst)',
workspace_sort_modified_desc: 'Änderungsdatum (neueste zuerst)',
workspace_sort_created_unavailable: 'Die Erstellungszeit wird von diesem Server oder dieser Plattform nicht gemeldet.',
workspace_show_hidden_files_desc: '.DS_Store, .git, node_modules und weitere versteckte / Systemdateien im Dateibaum anzeigen.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -10787,12 +10787,12 @@ const LOCALES = {
workspace_empty_no_path: '未选择工作区。请在 设置 → 工作区 中设置工作区以浏览文件。',
workspace_empty_dir: '此工作区为空。',
workspace_show_hidden_files: '显示隐藏文件',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: '排序方式',
workspace_sort_name_asc: '名称 (A → Z)',
workspace_sort_name_desc: '名称 (Z → A)',
workspace_sort_created_desc: '创建日期 (最新优先)',
workspace_sort_modified_desc: '修改日期 (最新优先)',
workspace_sort_created_unavailable: '此服务器或平台未报告创建时间。',
workspace_show_hidden_files_desc: '将 .DS_Store、.git、node_modules 以及其他隐藏文件/系统文件包含在文件树中。',
workspace_panel_show: '显示工作区面板',
workspace_panel_hide: '隐藏工作区面板',
@@ -12593,12 +12593,12 @@ const LOCALES = {
workspace_empty_no_path: '未選擇工作區。請在 設定 → 工作區 中設定工作區以瀏覽檔案。',
workspace_empty_dir: '此工作區為空。',
workspace_show_hidden_files: '顯示隱藏檔案',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: '排序方式',
workspace_sort_name_asc: '名稱 (A → Z)',
workspace_sort_name_desc: '名稱 (Z → A)',
workspace_sort_created_desc: '建立日期 (最新優先)',
workspace_sort_modified_desc: '修改日期 (最新優先)',
workspace_sort_created_unavailable: '此伺服器或平台未回報建立時間。',
workspace_show_hidden_files_desc: '在檔案樹中包含 .DS_Store、.git、node_modules 與其他隱藏/系統檔案。',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -14247,12 +14247,12 @@ const LOCALES = {
workspace_empty_no_path: 'Nenhum workspace selecionado. Configure em Configurações → Workspace.',
workspace_empty_dir: 'Este workspace está vazio.',
workspace_show_hidden_files: 'Mostrar arquivos ocultos',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Ordenar por',
workspace_sort_name_asc: 'Nome (A → Z)',
workspace_sort_name_desc: 'Nome (Z → A)',
workspace_sort_created_desc: 'Data de criação (mais recentes primeiro)',
workspace_sort_modified_desc: 'Data de modificação (mais recentes primeiro)',
workspace_sort_created_unavailable: 'A hora de criação não é informada por este servidor ou plataforma.',
workspace_show_hidden_files_desc: 'Include .DS_Store, .git, node_modules, and other hidden / system files in the file tree.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -15896,12 +15896,12 @@ const LOCALES = {
workspace_empty_no_path: 'No workspace selected. Set a workspace in Settings \u2192 Workspace to browse files.',
workspace_empty_dir: 'This workspace is empty.',
workspace_show_hidden_files: '숨김 파일 표시',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: '정렬 기준',
workspace_sort_name_asc: '이름 (A → Z)',
workspace_sort_name_desc: '이름 (Z → A)',
workspace_sort_created_desc: '생성 날짜 (최신순)',
workspace_sort_modified_desc: '수정 날짜 (최신순)',
workspace_sort_created_unavailable: '이 서버 또는 플랫폼에서 생성 시간을 보고하지 않습니다.',
workspace_show_hidden_files_desc: 'Include .DS_Store, .git, node_modules, and other hidden / system files in the file tree.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -17671,12 +17671,12 @@ const LOCALES = {
workspace_empty_no_path: 'Aucun espace de travail sélectionné. Définissez un espace de travail dans Paramètres → Espace de travail pour parcourir les fichiers.',
workspace_empty_dir: 'Cet espace de travail est vide.',
workspace_show_hidden_files: 'Afficher les fichiers cachés',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Trier par',
workspace_sort_name_asc: 'Nom (A → Z)',
workspace_sort_name_desc: 'Nom (Z → A)',
workspace_sort_created_desc: 'Date de création (plus récentes d\'abord)',
workspace_sort_modified_desc: 'Date de modification (plus récentes d\'abord)',
workspace_sort_created_unavailable: 'L\'heure de création n\'est pas indiquée par ce serveur ou cette plateforme.',
workspace_show_hidden_files_desc: 'Inclure .DS_Store, .git, node_modules et autres fichiers cachés/système dans l\'arborescence.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -20656,12 +20656,12 @@ const LOCALES = {
workspace_new_worktree_conversation_meta: 'Vytvořit izolovaný git worktree pro tento workspace.',
workspace_options: 'Možnosti pracovního prostoru',
workspace_show_hidden_files: 'Zobrazit skryté soubory',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Seřadit podle',
workspace_sort_name_asc: 'Název (A → Z)',
workspace_sort_name_desc: 'Název (Z → A)',
workspace_sort_created_desc: 'Datum vytvoření (nejnovější první)',
workspace_sort_modified_desc: 'Datum úpravy (nejnovější první)',
workspace_sort_created_unavailable: 'Čas vytvoření tento server nebo platforma neuvádí.',
workspace_show_hidden_files_desc: 'Do stromu souborů zahrňte soubory .DS_Store, .git, node_modules a další skryté / systémové soubory.',
workspace_switch_failed: 'Přepnutí workspace selhalo: ',
workspace_usage: 'Použití: /workspace <jméno>',
@@ -21116,12 +21116,12 @@ const LOCALES = {
workspace_empty_no_path: 'Çalışma alanı seçilmedi. Dosyalara göz atmak için Ayarlar \u2192 Çalışma Alanı\'nda bir çalışma alanı ayarlayın.',
workspace_empty_dir: 'Bu çalışma alanı boş.',
workspace_show_hidden_files: 'Gizli dosyaları göster',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Sırala',
workspace_sort_name_asc: 'Ad (A → Z)',
workspace_sort_name_desc: 'Ad (Z → A)',
workspace_sort_created_desc: 'Oluşturulma tarihi (en yeniler önce)',
workspace_sort_modified_desc: 'Değiştirilme tarihi (en yeniler önce)',
workspace_sort_created_unavailable: 'Oluşturma zamanı bu sunucu veya platform tarafından bildirilmiyor.',
workspace_show_hidden_files_desc: 'Dosya ağacına .DS_Store, .git, node_modules ve diğer gizli / sistem dosyalarını ekleyin.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -22889,12 +22889,12 @@ const LOCALES = {
workspace_empty_no_path: 'Nie wybrano obszaru roboczego. Ustaw obszar roboczy w Ustawienia → Obszar roboczy, aby przeglądać pliki.',
workspace_empty_dir: 'Ten obszar roboczy jest pusty.',
workspace_show_hidden_files: 'Pokaż ukryte pliki',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Sortuj według',
workspace_sort_name_asc: 'Nazwa (A → Z)',
workspace_sort_name_desc: 'Nazwa (Z → A)',
workspace_sort_created_desc: 'Data utworzenia (najnowsze pierwsze)',
workspace_sort_modified_desc: 'Data modyfikacji (najnowsze pierwsze)',
workspace_sort_created_unavailable: 'Czas utworzenia nie jest zgłaszany przez ten serwer lub platformę.',
workspace_show_hidden_files_desc: 'Dołącz .DS_Store, .git, node_modules i inne ukryte/systemowe pliki w drzewie plików.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
@@ -24609,12 +24609,12 @@ const LOCALES = {
workspace_empty_no_path: 'Chưa chọn workspace. Đặt workspace trong Settings → Workspace để duyệt file.',
workspace_empty_dir: 'Không gian làm việc này đang trống.',
workspace_show_hidden_files: 'Hiện file ẩn',
workspace_sort_by: 'Sort by',
workspace_sort_name_asc: 'Name (A → Z)',
workspace_sort_name_desc: 'Name (Z → A)',
workspace_sort_created_desc: 'Date created (newest first)',
workspace_sort_modified_desc: 'Date modified (newest first)',
workspace_sort_created_unavailable: 'Creation time is not reported by this server or platform.',
workspace_sort_by: 'Sắp xếp theo',
workspace_sort_name_asc: 'Tên (A → Z)',
workspace_sort_name_desc: 'Tên (Z → A)',
workspace_sort_created_desc: 'Ngày tạo (mới nhất trước)',
workspace_sort_modified_desc: 'Ngày sửa đổi (mới nhất trước)',
workspace_sort_created_unavailable: 'Thời gian tạo không được báo cáo bởi máy chủ hoặc nền tảng này.',
workspace_show_hidden_files_desc: 'Bao gồm .DS_Store, .git, node_modules và các file hệ thống / ẩn khác trong cây thư mục.',
workspace_panel_show: 'Show workspace panel',
workspace_panel_hide: 'Hide workspace panel',
+2 -2
View File
@@ -17,7 +17,7 @@
<meta name="apple-mobile-web-app-status-bar-style" content="black-translucent">
<meta name="apple-mobile-web-app-title" content="Hermes">
<link rel="apple-touch-icon" sizes="512x512" href="static/apple-touch-icon.png">
<script>(function(){try{var themes={light:1,dark:1,system:1},skins={codex:1,terracotta:1,default:1,ares:1,mono:1,graphite:1,github:1,slate:1,poseidon:1,sisyphus:1,charizard:1,sienna:1,catppuccin:1,hepburn:1,nous:1,'geist-contrast':1,neon:1,'neon-soft':1,'neon-paint':1,zeus:1,'verdigris':1},legacy={slate:['dark','slate'],solarized:['dark','poseidon'],monokai:['dark','sisyphus'],nord:['dark','slate'],oled:['dark','default']},t=(localStorage.getItem('hermes-theme')||'dark').toLowerCase(),s=(localStorage.getItem('hermes-skin')||'').toLowerCase(),m=legacy[t],theme=m?m[0]:(themes[t]?t:'dark');var pendingExt=s&&s!=='default'&&!skins[s]&&!legacy[s];var skin=skins[s]?s:(pendingExt?s:(m?m[1]:'default'));localStorage.setItem('hermes-theme',theme);localStorage.setItem('hermes-skin',skin);if(theme==='system')theme=window.matchMedia('(prefers-color-scheme:dark)').matches?'dark':'light';if(theme==='dark')document.documentElement.classList.add('dark');if(skin!=='default')document.documentElement.dataset.skin=skin;}catch(e){document.documentElement.classList.add('dark');}})()</script>
<script>(function(){try{var _hadAppearance=localStorage.getItem('hermes-theme')!==null||localStorage.getItem('hermes-skin')!==null;var themes={light:1,dark:1,system:1},skins={codex:1,terracotta:1,default:1,ares:1,mono:1,graphite:1,github:1,slate:1,poseidon:1,sisyphus:1,charizard:1,sienna:1,catppuccin:1,hepburn:1,nous:1,'geist-contrast':1,neon:1,'neon-soft':1,'neon-paint':1,zeus:1,'verdigris':1},legacy={slate:['dark','slate'],solarized:['dark','poseidon'],monokai:['dark','sisyphus'],nord:['dark','slate'],oled:['dark','default']},t=(localStorage.getItem('hermes-theme')||'dark').toLowerCase(),s=(localStorage.getItem('hermes-skin')||'').toLowerCase(),m=legacy[t],theme=m?m[0]:(themes[t]?t:'dark');var pendingExt=s&&s!=='default'&&!skins[s]&&!legacy[s];var skin=skins[s]?s:(pendingExt?s:(m?m[1]:'default'));if(_hadAppearance){localStorage.setItem('hermes-theme',theme);localStorage.setItem('hermes-skin',skin);}if(theme==='system')theme=window.matchMedia('(prefers-color-scheme:dark)').matches?'dark':'light';if(theme==='dark')document.documentElement.classList.add('dark');if(skin!=='default')document.documentElement.dataset.skin=skin;}catch(e){document.documentElement.classList.add('dark');}})()</script>
<script>(function(){try{var fs=localStorage.getItem('hermes-font-size');if(fs&&fs!=='default')document.documentElement.dataset.fontSize=fs;}catch(e){}})()</script>
<!-- theme-color: surfaces the active app chrome color to native status bars (Safari status bar, PWA, native WKWebView wrappers). Updated dynamically by boot.js when theme/skin changes. The light/dark default values match style.css :root --sidebar / :root.dark --sidebar. -->
<meta name="theme-color" content="#FAF7F0" media="(prefers-color-scheme: light)">
@@ -279,7 +279,7 @@
</div>
</div>
<div class="panel-head-sub" style="padding:0 12px 8px">
<select id="insightsPeriod" onchange="loadInsights()" style="width:100%;background:var(--input-bg);color:var(--text);border:1px solid var(--border);border-radius:6px;padding:4px 8px;font-size:12px">
<select id="insightsPeriod" class="insights-period-select" onchange="loadInsights()">
<option value="7">7 days</option>
<option value="30" selected>30 days</option>
<option value="90">90 days</option>
+318 -73
View File
@@ -1953,6 +1953,52 @@ async function send(){
}finally{ _sendInProgress=false; _sendInProgressSid=null; }
}
async function startRegeneration(sessionId, regenerationRevision){
const sid=String(sessionId||'');
if(!sid||!regenerationRevision||!S.session||S.session.session_id!==sid)return;
const snapshot=Array.isArray(S.messages)?S.messages.slice():[];
let assistantIndex=-1;
let userIndex=-1;
for(let i=snapshot.length-1;i>=0;i--){
if(assistantIndex<0&&snapshot[i]?.role==='assistant'){assistantIndex=i;continue;}
if(assistantIndex>=0&&snapshot[i]?.role==='user'){userIndex=i;break;}
}
if(userIndex<0)return;
const retained=Object.assign({},snapshot[userIndex],{_pending:true});
S.messages=snapshot.slice(0,userIndex+1);
S.messages[userIndex]=retained;
renderMessages();setBusy(true);
if(typeof ensureLiveWorklogShell==='function')ensureLiveWorklogShell();
else if(typeof appendThinking==='function')appendThinking('',{pending:true});
try{
const response=await api('/api/chat/start',{method:'POST',body:JSON.stringify({
session_id:sid,regenerate:true,regeneration_revision:regenerationRevision
})});
if(!S.session||S.session.session_id!==sid)return;
const streamId=response&&response.stream_id;
if(!streamId)throw new Error('Regeneration did not start a stream.');
S.activeStreamId=streamId;
S.session.active_stream_id=streamId;
S.session.regeneration_revision=null;
if(typeof response.pending_started_at==='number')S.session.pending_started_at=response.pending_started_at;
if(response.title&&typeof applySessionTitleUpdate==='function')applySessionTitleUpdate(sid,response.title);
if(!INFLIGHT[sid])INFLIGHT[sid]={messages:S.messages.slice(),uploaded:[],toolCalls:[]};
markInflight(sid,streamId);
if(typeof saveInflightState==='function')saveInflightState(sid,{streamId,messages:S.messages.slice(),uploaded:[],toolCalls:[]});
if(typeof showLiveRunStatus==='function')showLiveRunStatus(sid,{startedAt:S.session.pending_started_at||Date.now()/1000});
if(typeof updateSendBtn==='function')updateSendBtn();
if(typeof renderSessionList==='function')void renderSessionList();
attachLiveStream(sid,streamId,[]);
}catch(error){
if(S.session&&S.session.session_id===sid){
S.messages=snapshot;delete INFLIGHT[sid];
if(typeof clearInflightState==='function')clearInflightState(sid);
removeThinking();renderMessages();setBusy(false);setComposerStatus('');
}
throw error;
}
}
const LIVE_STREAMS={};
const _STREAM_NOTIFICATION_BACKGROUND={};
@@ -2613,7 +2659,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
if(streamId){
const st=await api(`/api/chat/stream/status?stream_id=${encodeURIComponent(streamId)}`);
if(st.active){
setComposerStatus('Reconnected');
setComposerStatus('Reconnected',1000);
_wireSSE(new EventSource(new URL(`api/chat/stream?stream_id=${encodeURIComponent(streamId)}${_runJournalReplayParams()}`,document.baseURI||location.href).href,{withCredentials:true}));
return;
}
@@ -4203,7 +4249,14 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
|| (text.includes('compressed')&&!text.includes('compressing'))
) return 'compressed';
if(
phase==='running'||phase==='compressing'
// NOT a bare phase==='running'. routes.py appends a placeholder
// "live anchor shell" row (role lifecycle, status running,
// source_event_type runtime_journal_snapshot) whenever a stream has
// events but no visible rows yet. That falls through the source check
// above, and a bare running phase then classified every such shell as
// a compression start - a permanent phantom "Compressing context"
// divider on sessions that never compressed anything.
phase==='compressing'
|| text.includes('compressing context')
|| text.includes('compacting context')
|| text.includes('preflight compression')
@@ -4333,7 +4386,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
s=s.replace(/<(?:\s*|\s*DSML\s*[||]\s*)?function_calls(?:>|$)[\s\S]*$/i,'');
// Remove malformed DSML tag fragments like "<|DSML |" that can leak in tokens.
s=s.replace(/<\s*|\s*DSML\s*[||]\s*/gi,'');
return s.trim();
return s.replace(/^\s+/, '');
}
function _streamDisplay(){
return _extractInlineThinkingFromContent(_stripXmlToolCalls(assistantText), liveReasoningText, {streaming:true}).content;
@@ -5030,6 +5083,13 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
};
_walk(rootEl);
}
// Exposed for the transparent-stream fade prose reconciler in ui.js
// (same pattern as __anchorProseIncrementalNode above): the no-cursor
// rebuild branch of _refreshTransparentFadeProseRow snapshots the rendered
// text before clearing and re-applies this mute so only genuinely-new tail
// words animate (#7082 review). The helper is stateless, so unlike
// __anchorProseIncrementalNode it never needs to be cleared per-stream.
if(typeof window!=='undefined') window.__streamFadeMuteRenderedPrefix=_streamFadeMuteRenderedPrefix;
function _streamFadePauseAfter(text, paragraphBreakIndex){
if(paragraphBreakIndex>=0) return 90;
const trimmed=String(text||'').trimEnd();
@@ -5226,6 +5286,10 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
const force=!!(options&&options.force);
const skipAnchorProcessProse=!!(options&&options.skipAnchorProcessProse);
if(!assistantBody||(!force&&!_renderPending)) return;
// #6449: guard — this stream's session is no longer the active pane.
// Callers already gate on _isActiveSession(), but add the guard here too
// so any future call-site cannot leak rendering into the wrong session.
if(!_isActiveSession()) return;
if(_renderPending) _cancelAnimationFramePendingStreamRender();
const displayText=segmentStart===0
? _parseStreamState().displayText
@@ -5551,6 +5615,11 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
}
if(_renderPending) return;
if(_streamFinalized) return; // Bug A: don't schedule new rAF after stream finalized
// #6449: guard — this stream's session is no longer the active frontend pane.
// Drop the scheduled render instead of writing into a detached or wrong-session DOM.
// Callers (token/interim_assistant handlers) already gate on _isActiveSession(), but
// the rAF/setTimeout window between schedule and execution can outlive a session switch.
if(!_isActiveSession()) return;
_renderPending=true;
// Cap render rate to ~15fps. The browser's rAF fires at 60fps, but each DOM
// update takes 50-150ms on large sessions. During GC pauses, rAF callbacks
@@ -5566,6 +5635,9 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
_renderPending=false;
// Guard: a pending setTimeout+rAF can outlive stream finalization.
if(_streamFinalized) return;
// #6449: guard — the frontend session changed between rAF schedule and execution.
// Writing DOM into this stream's assistantBody would leak text into the wrong pane.
if(!_isActiveSession()) return;
// Mobile scroll-jank guard: temporarily disable overflow-anchor before DOM
// writes to suppress Chromium scroll re-anchoring during streaming growth.
if(typeof window._fixMobileScrollJank==='function') window._fixMobileScrollJank();
@@ -6127,7 +6199,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
const _prevCost=(S.session&&S.session.estimated_cost)||0;
const _prevCacheRead=(S.session&&S.session.cache_read_tokens)||0;
const _prevCacheWrite=(S.session&&S.session.cache_write_tokens)||0;
S.session=d.session;S.messages=_carryForwardEphemeralTurnFields(S.messages||[], d.session.messages||[]);if(typeof _messagesTruncated!=='undefined')_messagesTruncated=!!d.session._messages_truncated;
S.session=d.session;S.messages=_carryForwardEphemeralTurnFields(S.messages||[], d.session.messages||[]);if(typeof _adoptRegenerationRevision==='function')_adoptRegenerationRevision(d.session);if(typeof _messagesTruncated!=='undefined')_messagesTruncated=!!d.session._messages_truncated;
// #4720: reset _oldestIdx (full-load symmetry; keeps the #4613 anchor aligned).
if(typeof _oldestIdx!=='undefined')_oldestIdx=d.session._messages_offset||0;
S.messages=_filterRecoveryControlMessages(S.messages || []);
@@ -6238,7 +6310,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
if(hasMessageToolMetadata) S._settledLiveToolMetadata=S.toolCalls.map(tc=>({...tc,done:true}));
S.toolCalls=hasMessageToolMetadata?[]:S.toolCalls.map(tc=>({...tc,done:true}));
}
if(typeof renderSessionArtifacts==='function') renderSessionArtifacts();
if(typeof projectSessionArtifactsForOwner==='function') projectSessionArtifactsForOwner(completedSid);
if(uploaded.length){
const lastUser=[...S.messages].reverse().find(m=>m.role==='user');
if(lastUser)lastUser.attachments=uploaded;
@@ -6306,6 +6378,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
}else if(_doneLiveScrollSnapshot&&typeof _restoreMessageScrollSnapshotSameFrame==='function'){
_restoreMessageScrollSnapshotSameFrame(_doneLiveScrollSnapshot);
}
if(typeof _restoreMessageRenderWindowAfterSettledRender==='function') _restoreMessageRenderWindowAfterSettledRender();
if(shouldFollowOnDone&&typeof scrollToBottom==='function') scrollToBottom();
if(typeof noteWorkspaceMutationsFromToolCalls==='function') noteWorkspaceMutationsFromToolCalls(S.toolCalls);
loadDir('.', { preservePreview: true });
@@ -6566,6 +6639,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
} else if(d.session&&typeof d.session==='object'){
S.session=d.session;
const _nextMsgs3018=(d.session.messages||[]).filter(m=>m&&m.role);
if(typeof _adoptRegenerationRevision==='function')_adoptRegenerationRevision(d.session);
_attachProjectedAnchorSceneToLastAssistant(_nextMsgs3018);
S.messages=_carryForwardEphemeralTurnFields(S.messages||[], _nextMsgs3018);
if(S.session&&S.session.session_id){
@@ -6636,9 +6710,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
return;
}
// Show as a small inline notice, not a full error
setComposerStatus(`${d.message||'Warning'}`);
// If it's a fallback notice, show it briefly then clear
if(d.type==='fallback') setTimeout(()=>setComposerStatus(''),4000);
setComposerStatus(`${d.message||'Warning'}`,d.type==='fallback'?4000:undefined);
}catch(_){}
});
@@ -6687,7 +6759,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
try{
const st=await api(`/api/chat/stream/status?stream_id=${encodeURIComponent(streamId)}`);
if(st&&st.active){
setComposerStatus('Reconnected');
setComposerStatus('Reconnected',1000);
_wireSSE(new EventSource(new URL(`api/chat/stream?stream_id=${encodeURIComponent(streamId)}${_runJournalReplayParams()}`,document.baseURI||location.href).href,{withCredentials:true}));
return;
}
@@ -6808,6 +6880,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
: (typeof _messageUserUnpinned!=='undefined' && _messageUserUnpinned));
S.session=sessionPayload;
const _nextMsgs3018=(sessionPayload.messages||[]).filter(m=>m&&m.role);
if(typeof _adoptRegenerationRevision==='function')_adoptRegenerationRevision(sessionPayload);
_attachProjectedAnchorSceneToLastAssistant(_nextMsgs3018);
S.messages=_carryForwardEphemeralTurnFields(S.messages||[], _nextMsgs3018);
if(typeof _hydrateTodosFromSession==='function') _hydrateTodosFromSession(S.session);
@@ -6980,6 +7053,7 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
try{localStorage.setItem('hermes-webui-session',S.session.session_id);}catch(_){}
if(typeof _setActiveSessionUrl==='function') _setActiveSessionUrl(S.session.session_id);
}
if(typeof _adoptRegenerationRevision==='function')_adoptRegenerationRevision(session);
const _markerOnlyAssistantError=_replaceMarkerOnlyAssistantWithStreamError(S.messages);
if(_markerOnlyAssistantError&&typeof showToast==='function') showToast('No response received after context compression. Please retry.',5000,'error');
const hasMessageToolMetadata=S.messages.some(m=>{
@@ -7005,6 +7079,8 @@ function attachLiveStream(activeSid, streamId, uploaded=[], options={}){
_messageRenderWindowSize=Math.max(typeof _currentMessageRenderWindowSize==='function'?_currentMessageRenderWindowSize():50, _messageRenderableMessageCount());
}
syncTopbar();renderMessages({preserveScroll:true});
if(typeof _restoreMessageRenderWindowAfterSettledRender==='function') _restoreMessageRenderWindowAfterSettledRender();
if(typeof projectSessionArtifactsForOwner==='function') projectSessionArtifactsForOwner(completedSid);
}
if(_isActiveSession()) _queueDrainSid=activeSid;
renderSessionList();
@@ -7160,7 +7236,28 @@ function autoResize(){
}
const el=$('msg');
const _nextValue=String(el.value||'');
if(typeof CSS!=='undefined'&&typeof CSS.supports==='function'&&CSS.supports('field-sizing','content')){
if(el.style.height) el.style.height='';
_composerLastResizeValue=_nextValue;
updateSendBtn();
return;
}
const _isAppendOnly=_nextValue.length>_composerLastResizeValue.length&&_nextValue.startsWith(_composerLastResizeValue);
// An EMPTY composer has no content to measure, so clear any inline height and
// let the CSS `min-height` define the resting size. Measuring instead would
// read the PLACEHOLDER's scrollHeight — a long busy/compression hint wraps to
// two or three lines and would grow the empty composer (71px for the English
// busy hint, 97px for the French compression one) purely because of hint text.
// That made the empty height history-dependent on this path: 44px on a fresh
// send, but grown after any later resize while empty. The native
// `field-sizing` path above always holds the resting height, so clearing here
// keeps both paths on the same contract.
if(!_nextValue){
if(el.style.height) el.style.height='';
_composerLastResizeValue=_nextValue;
updateSendBtn();
return;
}
const _fitsCurrentHeight=el.scrollHeight<=el.offsetHeight;
// Only a direct append at the natural one-row height can skip the height
// round trip. Replacements and an already-tall composer must remeasure so the
@@ -7253,18 +7350,9 @@ function _updateYoloPill() {
}
async function toggleYoloFromApproval() {
const sid = S.session && S.session.session_id;
if (!sid) return;
try {
await api('/api/session/yolo', {
method: 'POST',
body: JSON.stringify({ session_id: sid, enabled: true }),
});
_yoloEnabled = true;
_updateYoloPill();
hideApprovalCard(true);
showToast(t('yolo_enabled'));
} catch (e) { showToast('YOLO: ' + e.message); }
const owner = _captureApprovalResponseOwner();
if (!owner) return false;
return !!(await respondApproval('once', {yolo: true, owner}));
}
// ── Approval polling ──
@@ -7326,8 +7414,13 @@ function hideApprovalCard(force=false) {
return;
}
}
const preserveDisplayedOwner = _approvalOwnerIdentityMatches(
_approvalDisplayedOwner,
_approvalResponding,
);
_approvalSessionId = null;
_resetApprovalCardState();
if (!preserveDisplayedOwner) _approvalDisplayedOwner = null;
card.classList.remove("visible");
card.classList.remove("collapsed");
_setPromptFlyoutHidden(card, true);
@@ -7341,6 +7434,8 @@ let _approvalSessionId = null;
let _approvalCurrentId = null; // approval_id of the card currently shown
let _approvalPendingBySession = new Map();
let _approvalResponding = null;
let _approvalClearedOwner = null;
let _approvalDisplayedOwner = null;
const _DISMISSED_APPROVALS_KEY = 'hermes_dismissed_approvals';
@@ -7429,20 +7524,113 @@ function _renderPendingApprovalForActiveSession() {
if (entry) showApprovalCard(entry.pending, entry.pendingCount);
}
function _approvalResponseMatches(sid, approvalId) {
function _approvalMirrorOwnerFor(sid, approvalId) {
const entry = _approvalPendingBySession.get(sid);
const pending = entry && entry.pending;
if (!pending || pending.approval_id !== approvalId) return {runId: '', mirrorToken: ''};
const runId = String(pending.run_id || '').trim();
const mirrorToken = String(pending._gateway_mirror_token || '').trim();
return runId && mirrorToken ? {runId, mirrorToken} : {runId: '', mirrorToken: ''};
}
function _approvalOwnerForPending(sid, pending) {
if (!pending) return null;
const approvalId = pending.approval_id || null;
if (!sid || !approvalId) return null;
const runId = String(pending.run_id || '').trim();
const mirrorToken = String(pending._gateway_mirror_token || '').trim();
return {
sid,
approvalId,
runId: runId && mirrorToken ? runId : '',
mirrorToken: runId && mirrorToken ? mirrorToken : '',
};
}
function _approvalOwnerIdentityMatches(left, right) {
return !!(
left &&
right &&
left.sid === right.sid &&
left.approvalId === right.approvalId &&
left.runId === right.runId &&
left.mirrorToken === right.mirrorToken
);
}
function _captureApprovalResponseOwner() {
const card = $("approvalCard");
const sid = _approvalSessionId;
const approvalId = _approvalCurrentId;
if (!card || !card.classList.contains("visible") || !sid || !approvalId) return null;
if (!S.session || S.session.session_id !== sid) return null;
if (
!_approvalDisplayedOwner ||
_approvalDisplayedOwner.sid !== sid ||
_approvalDisplayedOwner.approvalId !== approvalId
) return null;
return {..._approvalDisplayedOwner, generation: _loadSessionGeneration};
}
function _approvalResponseOwnerIsCurrent(owner) {
return !!(
owner &&
S.session &&
S.session.session_id === owner.sid &&
_loadSessionGeneration === owner.generation &&
_approvalOwnerIdentityMatches(_approvalDisplayedOwner, owner)
);
}
function _approvalResponseMatches(
sid,
approvalId,
generation = _loadSessionGeneration,
mirrorOwner = _approvalMirrorOwnerFor(sid, approvalId),
) {
return !!(
_approvalResponding &&
_approvalResponding.sid === sid &&
(_approvalResponding.approvalId || null) === (approvalId || null)
_approvalResponding.generation === generation &&
_approvalResponding.approvalId === approvalId &&
_approvalResponding.runId === mirrorOwner.runId &&
_approvalResponding.mirrorToken === mirrorOwner.mirrorToken
);
}
function _releaseApprovalResponseOwner(owner) {
if (_approvalResponseMatches(owner.sid, owner.approvalId, owner.generation, owner)) {
_approvalResponding = null;
}
const card = $("approvalCard");
if (
_approvalOwnerIdentityMatches(_approvalDisplayedOwner, owner) &&
(!card || !card.classList.contains("visible"))
) {
_approvalDisplayedOwner = null;
}
}
function _approvalClearedOwnerMayRefresh(owner) {
return !!(
_approvalClearedOwner === owner &&
S.session &&
S.session.session_id === owner.sid &&
_loadSessionGeneration === owner.generation &&
_approvalSessionId === null &&
_approvalCurrentId === null
);
}
function _setApprovalControlsDisabled(choice, disabled) {
["approvalBtnOnce","approvalBtnSession","approvalBtnAlways","approvalBtnDeny"].forEach(id => {
const loadingId = choice === "skipAll"
? "approvalSkipAll"
: (choice ? "approvalBtn" + choice.charAt(0).toUpperCase() + choice.slice(1) : null);
["approvalBtnOnce","approvalBtnSession","approvalBtnAlways","approvalBtnDeny","approvalSkipAll"].forEach(id => {
const b = $(id);
if (!b) return;
b.disabled = !!disabled;
if (disabled && choice && b.id === "approvalBtn" + choice.charAt(0).toUpperCase() + choice.slice(1)) {
if (disabled && b.id === loadingId) {
b.classList.add("loading");
} else {
b.classList.remove("loading");
@@ -7460,16 +7648,25 @@ function showApprovalCard(pending, pendingCount) {
const sid = _rememberApprovalPending(pending, pendingCount);
if (!_approvalPromptBelongsToActiveSession(sid)) return;
if (pending && pending.approval_id && _isApprovalDismissed(sid, pending.approval_id)) return;
_approvalClearedOwner = null;
const keys = pending.pattern_keys || (pending.pattern_key ? [pending.pattern_key] : []);
const desc = (pending.description || "") + (keys.length ? " [" + keys.join(", ") + "]" : "");
const cmd = pending.command || "";
const sig = JSON.stringify({desc, cmd, sid: pending._session_id || (S.session && S.session.session_id) || null, approval_id: pending.approval_id || null});
const sig = JSON.stringify({
desc,
cmd,
sid: pending._session_id || (S.session && S.session.session_id) || null,
approval_id: pending.approval_id || null,
run_id: pending.run_id || null,
mirror_token: pending._gateway_mirror_token || null,
});
const card = $("approvalCard");
const sameApproval = card.classList.contains("visible") && _approvalSignature === sig;
$("approvalDesc").textContent = desc;
$("approvalCmd").textContent = cmd;
_approvalSessionId = sid;
_approvalCurrentId = pending.approval_id || null;
_approvalDisplayedOwner = _approvalOwnerForPending(sid, pending);
_approvalSignature = sig;
// Show "1 of N" counter when multiple approvals are queued
const counter = $("approvalCounter");
@@ -7492,7 +7689,7 @@ function showApprovalCard(pending, pendingCount) {
}
const responding = _approvalResponseMatches(sid, _approvalCurrentId);
_setApprovalControlsDisabled(
responding ? _approvalResponding.choice : null,
responding ? (_approvalResponding.controlChoice || _approvalResponding.choice) : null,
responding,
);
_setPromptFlyoutHidden(card, false);
@@ -7564,14 +7761,22 @@ function _syncApprovalTranscriptSpace(card, opts) {
setTimeout(measure, 420);
}
function _restoreFailedApprovalResponse(sid, errMsg) {
_approvalResponding = null;
function _restoreFailedApprovalResponse(owner, errMsg) {
const isCurrent = _approvalResponseOwnerIsCurrent(owner);
_releaseApprovalResponseOwner(owner);
if (!isCurrent) return;
_setApprovalControlsDisabled(null, false);
if (_approvalPromptBelongsToActiveSession(sid)) _renderPendingApprovalForActiveSession();
_renderPendingApprovalForActiveSession();
if (typeof showToast === "function") showToast(errMsg, 5000);
if (typeof setStatus === "function") setStatus(errMsg);
}
function _applyApprovalYoloProjection(result) {
if (!result || typeof result.yolo_enabled !== "boolean") return;
_yoloEnabled = result.yolo_enabled;
_updateYoloPill();
}
function toggleApprovalCardCollapsed(forceCollapsed) {
const card = $("approvalCard");
if (!card) return;
@@ -7581,59 +7786,84 @@ function toggleApprovalCardCollapsed(forceCollapsed) {
_syncApprovalTranscriptSpace(card, {immediate: true});
}
async function respondApproval(choice) {
const sid = _approvalSessionId || (S.session && S.session.session_id);
if (!sid) return;
const approvalId = _approvalCurrentId;
if (_approvalResponseMatches(sid, approvalId)) return;
async function respondApproval(choice, options = {}) {
const owner = options.owner || _captureApprovalResponseOwner();
if (!_approvalResponseOwnerIsCurrent(owner)) return false;
const {sid, approvalId} = owner;
if (_approvalResponseMatches(sid, approvalId, owner.generation, owner)) return false;
_approvalClearedOwner = null;
_unmarkApprovalDismissed(sid, approvalId);
_approvalResponding = {sid, approvalId: approvalId || null, choice};
_setApprovalControlsDisabled(choice, true);
const controlChoice = options.yolo ? "skipAll" : choice;
_approvalResponding = {...owner, choice};
_approvalResponding.controlChoice = controlChoice;
_setApprovalControlsDisabled(controlChoice, true);
try {
const result = await api("/api/approval/respond", {
method: "POST",
body: JSON.stringify({ session_id: sid, choice, approval_id: approvalId })
body: JSON.stringify({
session_id: sid,
choice,
approval_id: approvalId,
...(owner.runId ? {run_id: owner.runId} : {}),
...(owner.mirrorToken ? {mirror_token: owner.mirrorToken} : {}),
...(options.yolo ? {yolo: true} : {}),
})
});
if (!_approvalResponseOwnerIsCurrent(owner)) {
_releaseApprovalResponseOwner(owner);
return false;
}
if (result && result.ok) {
_approvalResponding = null;
_releaseApprovalResponseOwner(owner);
if (options.yolo) _applyApprovalYoloProjection(result);
const pendingEntry = _approvalPendingBySession.get(sid);
const samePending = !!(pendingEntry && pendingEntry.pending && (pendingEntry.pending.approval_id || null) === (approvalId || null));
// `stale_cleared` means the server found nothing pending for this session
// (the approval already resolved or its stream ended while the card was
// up). The orphan card must be cleared unconditionally so it can never
// get stuck — even if the displayed id has since drifted. (#4948 local
// variant: previously surfaced as a stuck "Approval response not
// accepted." toast.)
if (result.stale_cleared || (_approvalSessionId === sid && _approvalCurrentId === approvalId)) {
_approvalSessionId = null;
_approvalCurrentId = null;
hideApprovalCard(true);
}
if (samePending || result.stale_cleared) _clearApprovalPendingForSession(sid);
// Hardening for the narrow stale-clear race: a brand-new approval could
// have been parked server-side after the server's empty-check but before
// we processed this stale response. The unconditional clear above would
// hide that fresh card. Re-query the authoritative server pending state
// (same endpoint the fallback poll uses) so any approval that arrived in
// the window re-surfaces immediately instead of waiting for the next
// SSE/poll tick. Best-effort; poll/SSE remain the backstop. (Opus review
// nit on the #4948 fix.)
const pendingOwner = _approvalMirrorOwnerFor(sid, approvalId);
const samePending = !!(
pendingEntry &&
pendingEntry.pending &&
pendingEntry.pending.approval_id === approvalId &&
pendingOwner.runId === owner.runId &&
pendingOwner.mirrorToken === owner.mirrorToken
);
if (samePending) _clearApprovalPendingForSession(sid);
_approvalSessionId = null;
_approvalCurrentId = null;
_approvalClearedOwner = owner;
hideApprovalCard(true);
if (result.stale_cleared) {
api("/api/approval/pending?session_id=" + encodeURIComponent(sid), {timeoutToast: false})
.then(data => {
if (data && data.pending && _approvalPromptBelongsToActiveSession(sid)) {
showApprovalForSession(sid, data.pending, data.pending_count || 1);
}
})
.catch(() => {});
void (async () => {
if (!_approvalClearedOwnerMayRefresh(owner)) return;
try {
const data = await api("/api/approval/pending?session_id=" + encodeURIComponent(sid), {timeoutToast: false});
if (!_approvalClearedOwnerMayRefresh(owner)) return;
_approvalClearedOwner = null;
if (data && data.pending) showApprovalForSession(sid, data.pending, data.pending_count || 1);
} catch (_) {
if (_approvalClearedOwner === owner) _approvalClearedOwner = null;
}
})();
}
return;
if (options.yolo) showToast(t(_yoloEnabled ? 'yolo_enabled' : 'yolo_disabled'));
return options.yolo ? result : true;
}
const errMsg = (result && result.error) || "Approval response not accepted.";
_restoreFailedApprovalResponse(sid, errMsg);
_restoreFailedApprovalResponse(owner, errMsg);
return false;
} catch(e) {
const errMsg = (e && e.message) || (t("approval_responding") + " failed");
_restoreFailedApprovalResponse(sid, errMsg);
let errorPayload = null;
if (e && typeof e.body === 'string') {
try { errorPayload = JSON.parse(e.body); } catch (_) { /* non-JSON HTTP error */ }
}
const errMsg = (errorPayload && (errorPayload.error || errorPayload.message))
|| (e && e.message)
|| (t("approval_responding") + " failed");
if (!_approvalResponseOwnerIsCurrent(owner)) {
_releaseApprovalResponseOwner(owner);
return false;
}
if (options.yolo) _applyApprovalYoloProjection(errorPayload);
_restoreFailedApprovalResponse(owner, errMsg);
return false;
}
}
@@ -8012,6 +8242,20 @@ function startSessionStream(sid) {
// the hidden-tab return (loadSession force + keepStaleUntilLoaded): the new
// transcript replaces the old in a single render frame — NO clear+refetch,
// so the #5177/#5189 blank-gap "jump" is not reintroduced.
// #6999: focusing a backgrounded tab also fires the visibility-recovery
// probe in sessions.js (refreshActiveSessionIfExternallyUpdated), which
// holds the shared _activeSessionExternalRefreshInFlight guard while it
// probes + force-reloads this same session. This handler honors that guard
// so the two paths never start two concurrent loadSession(force) calls
// (double full-transcript fetch + double renderMessages pass = the OOM
// pattern on long sessions). The probe side carries its own
// _loadingSessionId guard; loadSession() keeps its legitimate
// newest-wins supersede semantics untouched.
// #6999 re-gate: while the probe owns the guard, frames are COALESCED via
// _coalesceSessionUpdatedWhileRefreshHeld — the max announced count is
// latched per SID and the owner's finally runs ONE guarded follow-up when
// local state is still behind (a bare return dropped the update; production
// does not guarantee a second event).
es.addEventListener('session-updated', e => {
try {
const d = JSON.parse(e.data || '{}');
@@ -8024,12 +8268,13 @@ function startSessionStream(sid) {
: (S.session && S.session.session_id === sid);
if (!isCurrent) return;
if (S.activeStreamId) return;
const serverCount = Number(d.message_count);
if (typeof _coalesceSessionUpdatedWhileRefreshHeld === 'function' && _coalesceSessionUpdatedWhileRefreshHeld(sid, serverCount)) return;
// Re-check against our CURRENT known count — a concurrent load may have
// already caught us up between the server's emit and now.
const localCount = (S.session && S.session.session_id === sid && Number.isFinite(Number(S.session.message_count)))
? Number(S.session.message_count)
: (Array.isArray(S.messages) ? S.messages.length : 0);
const serverCount = Number(d.message_count);
if (!Number.isFinite(serverCount) || serverCount <= localCount) return;
if (typeof loadSession === 'function') {
void loadSession(sid, {force: true, externalRefreshReason: 'session-updated', keepStaleUntilLoaded: true});
@@ -8397,7 +8642,7 @@ function _clarifyExpiryMs(pending) {
if (Number.isFinite(expiresAt) && expiresAt > 0) return expiresAt * 1000;
const requestedAt = Number(pending && pending.requested_at);
const timeoutSeconds = Number(pending && pending.timeout_seconds);
if (Number.isFinite(requestedAt) && Number.isFinite(timeoutSeconds)) {
if (Number.isFinite(requestedAt) && Number.isFinite(timeoutSeconds) && timeoutSeconds > 0) {
return (requestedAt + timeoutSeconds) * 1000;
}
return 0;
+1 -1
View File
@@ -705,7 +705,7 @@ async function startCodexOAuth(){
<p style="margin-top:8px"><strong>${t('oauth_codex_step2')}</strong></p>
<div style="display:flex;gap:8px;align-items:center;flex-wrap:wrap;margin-top:4px">
<code style="display:inline-block;font-size:18px;letter-spacing:0.1em;background:rgba(255,255,255,.08);padding:6px 14px;border-radius:8px;user-select:all">${esc(user_code)}</code>
<button class="sm-btn" type="button" onclick="copyCodexOAuthCode('${esc(user_code)}')">Copy code</button>
<button class="sm-btn" type="button" onclick="copyCodexOAuthCode(${jsArg(user_code)})">Copy code</button>
<button class="sm-btn" type="button" onclick="cancelCodexOAuth()">Cancel</button>
</div>
<p style="margin-top:8px;color:var(--muted);font-size:13px">${t('oauth_codex_polling')}</p>
+205 -79
View File
@@ -1217,7 +1217,7 @@ function _cronAgentPromptCardHtml(job){
return `<div class="detail-card">
<div class="detail-card-title detail-card-title-row">
<span>${esc(t('cron_prompt_label') || 'Prompt')}</span>
<button type="button" class="detail-expand-toggle" onclick="toggleCronPromptExpanded('${esc(job.id)}')" title="${esc(promptToggleLabel)}" aria-label="${esc(promptToggleLabel)}">${esc(promptExpanded ? '▴' : '▾')}</button>
<button type="button" class="detail-expand-toggle" onclick="toggleCronPromptExpanded(${jsArg(job.id)})" title="${esc(promptToggleLabel)}" aria-label="${esc(promptToggleLabel)}">${esc(promptExpanded ? '▴' : '▾')}</button>
</div>
<div class="detail-prompt ${promptExpanded ? 'expanded' : ''}">${esc(job.prompt || '')}</div>
</div>`;
@@ -1366,11 +1366,11 @@ async function _loadCronDetailRuns(jobId, detailKey){
const usageStrip = isScriptJob ? '' : _formatCronRunUsageStrip(run.usage);
const runExpanded = _cronExpansionGet(_cronRunExpandKey(jobId, run.filename));
const runToggleLabel = runExpanded ? (t('cron_collapse_output') || 'Collapse output') : (t('cron_expand_output') || 'Expand output');
return `<div class="detail-run-item" id="${rid}">
<div class="detail-run-head" onclick="_loadRunContent('${esc(jobId)}','${esc(run.filename)}','${rid}')">
return `<div class="detail-run-item" id="${esc(rid)}">
<div class="detail-run-head" onclick="_loadRunContent(${jsArg(jobId)},${jsArg(run.filename)},${jsArg(rid)})">
<span><span style="opacity:.7">${esc(ts)}</span> <span style="opacity:.4;font-size:11px">${esc(sizeStr)}</span>${usageStrip ? ` <span class="cron-run-usage-strip">${esc(usageStrip)}</span>` : ''}</span>
<span class="detail-run-actions">
<button type="button" class="detail-expand-toggle" onclick="event.stopPropagation();toggleCronRunExpanded('${esc(jobId)}','${esc(run.filename)}','${rid}')" title="${esc(runToggleLabel)}" aria-label="${esc(runToggleLabel)}">${esc(runExpanded ? '▴' : '▾')}</button>
<button type="button" class="detail-expand-toggle" onclick="event.stopPropagation();toggleCronRunExpanded(${jsArg(jobId)},${jsArg(run.filename)},${jsArg(rid)})" title="${esc(runToggleLabel)}" aria-label="${esc(runToggleLabel)}">${esc(runExpanded ? '▴' : '▾')}</button>
<span style="opacity:.6">▸</span>
</span>
</div>
@@ -1889,17 +1889,19 @@ function cancelCronForm(){
_clearCronDetail();
}
function _cronModelBareName(model, provider) {
function _modelBareNameForProvider(model, provider) {
// Strip @provider: prefix from a model value when provider is stored separately.
// The model dropdown may contain values like "@custom:9router:chat" (from
// _apply_provider_prefix) but cron jobs store model and provider separately,
// so the model should be just "chat".
if (model && provider && model.startsWith('@' + provider + ':')) {
return model.slice(('@' + provider + ':').length);
}
return model;
}
function _cronModelBareName(model, provider) {
// Cron jobs store model and provider separately, just like auxiliary slots.
return _modelBareNameForProvider(model, provider);
}
async function saveCronForm(){
const nameEl=$('cronFormName');
const schEl=$('cronFormSchedule');
@@ -2211,7 +2213,7 @@ function _kanbanRenderSidebar(columns){
}
list.innerHTML = tasks.map(task => {
const meta = _kanbanTaskMeta(task);
return `<button class="kanban-list-item" onclick="loadKanbanTask('${esc(task.id)}')">
return `<button class="kanban-list-item" onclick="loadKanbanTask(${jsArg(task.id)})">
<span class="kanban-list-status">${esc(_kanbanColumnLabel(task.status))}</span>
<span class="kanban-list-title">${esc(_kanbanTaskTitle(task))}</span>
${meta.length ? `<span class="kanban-meta">${esc(meta.join(' · '))}</span>` : ''}
@@ -2429,10 +2431,9 @@ function _kanbanCardStalenessClass(task){
}
function _kanbanCardQuickActions(task){
const id = esc(task.id || '');
const status = task.status || '';
const complete = status !== 'done' && status !== 'archived' ? `<button type="button" class="kanban-card-action" onclick="quickKanbanCardAction(event,'${id}','done')">${esc(t('kanban_card_complete'))}</button>` : '';
const archive = status !== 'archived' ? `<button type="button" class="kanban-card-action danger" onclick="quickKanbanCardAction(event,'${id}','archived')">${esc(t('kanban_card_archive'))}</button>` : '';
const complete = status !== 'done' && status !== 'archived' ? `<button type="button" class="kanban-card-action" onclick="quickKanbanCardAction(event,${jsArg(task.id)},'done')">${esc(t('kanban_card_complete'))}</button>` : '';
const archive = status !== 'archived' ? `<button type="button" class="kanban-card-action danger" onclick="quickKanbanCardAction(event,${jsArg(task.id)},'archived')">${esc(t('kanban_card_archive'))}</button>` : '';
return `<div class="kanban-card-actions" onclick="event.stopPropagation()">${complete}${archive}</div>`;
}
@@ -2515,7 +2516,7 @@ function _kanbanLaneNames(columns){
function _kanbanRenderColumn(col){
const tasks = col.tasks || [];
return `<section class="kanban-column" data-status="${esc(col.name)}" data-kanban-status="${esc(col.name)}" ondragover="allowKanbanDrop(event)" ondragenter="event.currentTarget.classList.add('drop-target')" ondragleave="clearKanbanDrop(event)" ondrop="dropKanbanTask(event, '${esc(col.name)}')">
return `<section class="kanban-column" data-status="${esc(col.name)}" data-kanban-status="${esc(col.name)}" ondragover="allowKanbanDrop(event)" ondragenter="event.currentTarget.classList.add('drop-target')" ondragleave="clearKanbanDrop(event)" ondrop="dropKanbanTask(event, ${jsArg(col.name)})">
<div class="kanban-column-head">
<span>${esc(_kanbanColumnLabel(col.name))}</span>
<span class="kanban-count">${tasks.length}</span>
@@ -2583,7 +2584,7 @@ function _kanbanCard(task, status){
const stale = _kanbanCardStalenessClass(task);
const body = _kanbanTaskBody(task);
const assignee = task.assignee ? `<span class="kanban-card-assignee">@${esc(task.assignee)}</span>` : `<span class="kanban-card-unassigned">${esc(t('kanban_unassigned'))}</span>`;
return `<article class="kanban-card ${esc(stale)}" data-kanban-task-id="${esc(task.id)}" draggable="true" ondragstart="dragKanbanTask(event, '${esc(task.id)}')" ondragend="finishKanbanDrag(event)" onclick="return openKanbanCard(event, '${esc(task.id)}')" tabindex="0" role="button" onkeydown="if(event.key==='Enter'||event.key===' '){event.preventDefault();loadKanbanTask('${esc(task.id)}')}">
return `<article class="kanban-card ${esc(stale)}" data-kanban-task-id="${esc(task.id)}" draggable="true" ondragstart="dragKanbanTask(event, ${jsArg(task.id)})" ondragend="finishKanbanDrag(event)" onclick="return openKanbanCard(event, ${jsArg(task.id)})" tabindex="0" role="button" onkeydown="if(event.key==='Enter'||event.key===' '){event.preventDefault();loadKanbanTask(${jsArg(task.id)})}">
<div class="kanban-card-topline"><span class="kanban-card-id">${esc(task.id || '')}</span>${priority ? `<span class="kanban-badge priority">P${priority}</span>` : ''}${task.tenant ? `<span class="kanban-badge tenant">${esc(task.tenant)}</span>` : ''}</div>
<div class="kanban-card-title">${esc(_kanbanTaskTitle(task))}</div>
${body ? `<div class="kanban-card-body">${_kanbanRenderMarkdown(body)}</div>` : ''}
@@ -3092,14 +3093,6 @@ function _kanbanRunHtml(run){
</div>`;
}
function _kanbanJsArg(s){
// Encode a value for safe interpolation inside an inline on* handler's JS
// string literal. JSON.stringify quotes/escapes for JS context; esc() then
// makes it safe inside the HTML attribute. Without this, a task id containing
// a quote breaks out of the handler (esc() alone is HTML-escaping, which the
// browser decodes BEFORE executing the inline handler). (#3797)
return esc(JSON.stringify(String(s == null ? '' : s)));
}
function _kanbanLinkableTaskOptions(excludeId){
// Datalist of existing task ids (with title as the option label) so the
// dependency field is a pick-from-real-tasks autocomplete rather than a blind
@@ -3124,7 +3117,7 @@ function _kanbanLinksHtml(links){
const item = (id, isParent) => {
const parentId = isParent ? id : taskId;
const childId = isParent ? taskId : id;
return `<code>${esc(id)} <button class="btn mini" onclick="removeKanbanDependency(${_kanbanJsArg(parentId)},${_kanbanJsArg(childId)})" data-i18n="kanban_remove_dependency" title="${esc(t('kanban_remove_dependency') || 'Remove')}">✕</button></code>`;
return `<code>${esc(id)} <button class="btn mini" onclick="removeKanbanDependency(${jsArg(parentId)},${jsArg(childId)})" data-i18n="kanban_remove_dependency" title="${esc(t('kanban_remove_dependency') || 'Remove')}">✕</button></code>`;
};
const hasLinks = parents.length || children.length;
return `<div class="kanban-detail-links-section">
@@ -3135,7 +3128,7 @@ function _kanbanLinksHtml(links){
<div class="kanban-detail-links-controls">
<input type="text" id="kanbanDependencyInput" class="kanban-detail-links-input" list="kanbanDependencyOptions" maxlength="255" autocomplete="off" data-i18n-placeholder="kanban_dependency_placeholder" placeholder="Task ID to link">
<datalist id="kanbanDependencyOptions">${_kanbanLinkableTaskOptions(taskId)}</datalist>
<button class="btn secondary" onclick="addKanbanDependency(${_kanbanJsArg(taskId)})" data-i18n="kanban_add_dependency">Add dependency</button>
<button class="btn secondary" onclick="addKanbanDependency(${jsArg(taskId)})" data-i18n="kanban_add_dependency">Add dependency</button>
</div>
</div>`;
}
@@ -3763,12 +3756,12 @@ function _kanbanRenderTaskDetail(data){
// dashboard plugin's contract. UI users want to claim/promote a ready task
// via the dispatcher Nudge button, not flip it to running by hand.
const statusButtons = ['triage', 'todo', 'ready', 'blocked', 'done', 'archived'].map(status =>
`<button class="btn secondary" onclick="updateKanbanTask('${esc(task.id)}',{status:'${status}'})">${esc(_kanbanColumnLabel(status))}</button>`
).join('') + `<button class="btn secondary" onclick="blockKanbanTask('${esc(task.id)}')">${esc(t('kanban_block'))}</button><button class="btn secondary" onclick="unblockKanbanTask('${esc(task.id)}')">${esc(t('kanban_unblock'))}</button>`;
`<button class="btn secondary" onclick="updateKanbanTask(${jsArg(task.id)},{status:'${status}'})">${esc(_kanbanColumnLabel(status))}</button>`
).join('') + `<button class="btn secondary" onclick="blockKanbanTask(${jsArg(task.id)})">${esc(t('kanban_block'))}</button><button class="btn secondary" onclick="unblockKanbanTask(${jsArg(task.id)})">${esc(t('kanban_unblock'))}</button>`;
return `<div class="kanban-task-preview-header">
<button class="btn secondary kanban-back-btn" onclick="closeKanbanTaskDetail()">${esc(t('kanban_back_to_board'))}</button>
<div class="kanban-task-preview-title">${esc(title)}</div>
<button class="btn secondary kanban-edit-btn" onclick="openKanbanEdit('${esc(task.id)}')" data-i18n="kanban_edit_task" title="${esc(t('kanban_edit_task') || 'Edit task')}">${esc(t('kanban_edit_task') || 'Edit task')}</button>
<button class="btn secondary kanban-edit-btn" onclick="openKanbanEdit(${jsArg(task.id)})" data-i18n="kanban_edit_task" title="${esc(t('kanban_edit_task') || 'Edit task')}">${esc(t('kanban_edit_task') || 'Edit task')}</button>
</div>
<div class="kanban-task-preview-body">${_kanbanRenderMarkdown(body)}</div>
${meta.length ? `<div class="kanban-meta">${esc(meta.join(' · '))}</div>` : ''}
@@ -3782,7 +3775,7 @@ function _kanbanRenderTaskDetail(data){
</div>
<div class="kanban-comment-form">
<textarea id="kanbanCommentInput" rows="2" placeholder="${esc(t('kanban_add_comment'))}"></textarea>
<button class="btn primary" onclick="addKanbanComment('${esc(task.id)}')">${esc(t('kanban_add_comment'))}</button>
<button class="btn primary" onclick="addKanbanComment(${jsArg(task.id)})">${esc(t('kanban_add_comment'))}</button>
</div>`;
}
@@ -3972,7 +3965,7 @@ function _renderKanbanBoardMenu(boards, current){
const icon = b.icon ? esc(b.icon) : '';
const safeColor = _kanbanSafeColor(b.color);
const colorStyle = safeColor ? `color:${safeColor}` : '';
return `<button type="button" class="kanban-board-switcher-item ${isCurrent ? 'is-current' : ''}" role="menuitem" data-board-slug="${esc(b.slug)}" onclick="switchKanbanBoard('${esc(b.slug)}')">
return `<button type="button" class="kanban-board-switcher-item ${isCurrent ? 'is-current' : ''}" role="menuitem" data-board-slug="${esc(b.slug)}" onclick="switchKanbanBoard(${jsArg(b.slug)})">
<span class="kanban-board-switcher-item-icon" style="${colorStyle}">${icon || (isCurrent ? '✓' : '')}</span>
<span class="kanban-board-switcher-item-name">${esc(b.name || b.slug)}</span>
<span class="kanban-board-switcher-item-count">${esc(String(total))}</span>
@@ -5374,13 +5367,13 @@ function _renderExternalNotesSources() {
<div class="notes-source-card-head notes-ai-recent-head"><strong>${li('bot', 14)}${esc(t('external_notes_recent_ai'))}</strong><span class="detail-badge">${esc(t('external_notes_auto'))}</span></div>
<div class="notes-ai-recent-list">${recentAiNotes.map(note => {
const updated = note.updated_time ? new Date(Number(note.updated_time)).toLocaleString() : '';
return `<button type="button" class="notes-result-card notes-ai-recent-item" onclick="previewExternalNote('${esc(note.source||'joplin')}','${esc(note.id||'')}')"><strong>${esc(note.title||note.label||'Untitled')}</strong><span>${li('clock', 14)}${esc(note.label||t('external_notes_recent_ai_reason'))}${updated ? ` · ${esc(updated)}` : ''}</span></button>`;
return `<button type="button" class="notes-result-card notes-ai-recent-item" onclick="previewExternalNote(${jsArg(note.source||'joplin')},${jsArg(note.id||'')})"><strong>${esc(note.title||note.label||'Untitled')}</strong><span>${li('clock', 14)}${esc(note.label||t('external_notes_recent_ai_reason'))}${updated ? ` · ${esc(updated)}` : ''}</span></button>`;
}).join('')}</div>
</section>`
: '';
const searchError = _notesSearchError ? `<div class="detail-form-error">${esc(_notesSearchError)}</div>` : '';
const resultHtml = _notesSearchResults.length
? `<div class="notes-search-results">${_notesSearchResults.map(note => `<button type="button" class="notes-result-card" onclick="previewExternalNote('${esc(note.source||_notesSelectedSource)}','${esc(note.id||'')}')"><strong>${esc(note.title||'Untitled')}</strong>${note.snippet?`<span>${esc(note.snippet)}</span>`:''}</button>`).join('')}</div>`
? `<div class="notes-search-results">${_notesSearchResults.map(note => `<button type="button" class="notes-result-card" onclick="previewExternalNote(${jsArg(note.source||_notesSelectedSource)},${jsArg(note.id||'')})"><strong>${esc(note.title||'Untitled')}</strong>${note.snippet?`<span>${esc(note.snippet)}</span>`:''}</button>`).join('')}</div>`
: `<div class="memory-empty">${esc(t('external_notes_search_empty'))}</div>`;
const previewHtml = _notesPreviewNote
? `<section class="notes-source-card notes-preview-card"><div class="notes-source-card-head"><strong>${esc(_notesPreviewNote.title||'Untitled')}</strong><span class="detail-badge">${esc(_notesPreviewNote.source||_notesSelectedSource)}</span></div><div class="memory-content preview-md">${renderMd(_notesPreviewNote.body||'')}</div></section>`
@@ -5832,9 +5825,11 @@ function renderWorkspaceDropdownInto(dd, workspaces, currentWs){
const sc=searchRow.querySelector('.ws-search-clear');
dd.appendChild(searchRow);
// ── Workspace list ──────────────────────────────────────────────────────
// Sort alphabetically by name (case-insensitive) before rendering.
const sorted=[...workspaces].sort((a,b)=>(a.name||'').localeCompare(b.name||''));
// Render in the server's stored order — the same order shown in the
// Workspaces settings panel (user-controlled via drag-and-drop reorder,
// with the default "Home" workspace first). No client-side re-sorting.
// Shallow copy so later in-place mutations can't touch the caller's array.
const sorted=[...workspaces];
const listContainer=document.createElement('div');
listContainer.className='ws-list-container';
dd.appendChild(listContainer);
@@ -9881,7 +9876,18 @@ function _extensionSettingsControls(entry){
</div>`;
}
function _extensionInstalledList(extensions,extensionDirConfigured){
function _extensionConfigureButton(entry,surface){
if(surface!=='installed'||!(entry&&entry.effective_enabled)) return '';
const id=(entry&&entry.id)||'';
const runtime=window.HermesExtensionSettings;
if(!id||!runtime||typeof runtime._configureStateForExtension!=='function') return '';
const state=runtime._configureStateForExtension(id);
if(!state||!state.available) return '';
const pending=state.pending===true;
return `<button class="sm-btn extension-configure-btn" type="button" data-extension-configure-id="${esc(id)}" aria-busy="${pending?'true':'false'}"${pending?' disabled':''}>${pending?'Opening…':'Configure'}</button>`;
}
function _extensionInstalledList(extensions,extensionDirConfigured,surface){
const list=Array.isArray(extensions)?extensions:[];
if(!list.length){
if(!extensionDirConfigured) return '<div class="extension-url-empty">No extension directory is configured.</div>';
@@ -9898,6 +9904,7 @@ function _extensionInstalledList(extensions,extensionDirConfigured){
const note=canToggle
? 'Toggles the WebUI-managed override for the next app load.'
: 'Manifest-disabled entries cannot be enabled from WebUI.';
const configureButton=_extensionConfigureButton(entry,surface);
return `<div class="extension-installed-row" data-extension-id="${esc(id)}">
<div class="extension-installed-main">
<div class="extension-installed-title-row">
@@ -9907,7 +9914,10 @@ function _extensionInstalledList(extensions,extensionDirConfigured){
<div class="extension-installed-meta"><code>${esc(id)}</code><span>${esc(note)}</span></div>
${_extensionSettingsControls(entry)}
</div>
<button class="sm-btn extension-toggle-btn" type="button" data-extension-toggle-id="${esc(id)}" data-extension-next-enabled="${nextEnabled}"${disabledAttr}>${esc(buttonText)}</button>
<div class="extension-installed-actions">
${configureButton}
<button class="sm-btn extension-toggle-btn" type="button" data-extension-toggle-id="${esc(id)}" data-extension-next-enabled="${nextEnabled}"${disabledAttr}>${esc(buttonText)}</button>
</div>
</div>`;
}).join('')}</div>`;
}
@@ -10115,6 +10125,7 @@ function _renderExtensionsPanel(data,seq){
const copyBtn=$('extensionsCopyDiagnosticsBtn');
if(!target) return;
_extensionsStatusData=data||null;
if(_extensionsGalleryData) _extensionsGalleryData.statusData=data||null;
_configureExtensionSettingsFromStatus(data);
if(copyBtn) copyBtn.disabled=!data;
const manifest=(data&&data.manifest)||{};
@@ -10165,7 +10176,7 @@ function _renderExtensionsPanel(data,seq){
</div>
</div>
<div class="provider-card-body extension-card-body">
${_extensionInstalledList(extensions,!!(data&&data.extension_dir_configured))}
${_extensionInstalledList(extensions,!!(data&&data.extension_dir_configured),'diagnostics')}
</div>
</div>
<div class="provider-card extension-assets-card">
@@ -10201,6 +10212,37 @@ function _renderExtensionsPanel(data,seq){
_monitorExtensionSidecars(sidecars,seq);
}
function _bindExtensionConfigureButtons(root){
if(!root) return;
root.querySelectorAll('[data-extension-configure-id]').forEach(btn=>{
btn.addEventListener('click',()=>handleExtensionConfigure(btn));
});
}
function _syncExtensionConfigureButtonState(id){
const runtime=window.HermesExtensionSettings;
if(!runtime||typeof runtime._configureStateForExtension!=='function') return;
const state=runtime._configureStateForExtension(id);
document.querySelectorAll('[data-extension-configure-id]').forEach(btn=>{
if(!btn.dataset||btn.dataset.extensionConfigureId!==id) return;
const pending=!!(state&&state.available&&state.pending);
btn.disabled=pending;
btn.setAttribute('aria-busy',pending?'true':'false');
btn.textContent=pending?'Opening…':'Configure';
});
}
function handleExtensionConfigure(btn){
if(!btn||btn.disabled) return;
const id=btn.dataset.extensionConfigureId||'';
const runtime=window.HermesExtensionSettings;
if(!id||!runtime||typeof runtime._invokeConfigure!=='function') return;
runtime._invokeConfigure(id,{
opener:btn,
onError:()=>showToast('Extension configuration failed.',4200,'error'),
});
}
function _bindExtensionToggleButtons(root){
if(!root) return;
root.querySelectorAll('[data-extension-toggle-id]').forEach(btn=>{
@@ -10362,6 +10404,21 @@ function switchExtensionsTab(tab){
if(tab==='gallery'&&!_extensionsGalleryLoaded) loadExtensionsGallery();
}
function _handleExtensionConfigureChange(change){
if(!change||!change.id) return;
if(change.reason==='pending'){
_syncExtensionConfigureButtonState(change.id);
return;
}
if(_extensionsGalleryData&&_extensionsGalleryData.statusData){
_renderInstalledExtensionsSurface(_extensionsGalleryData.statusData);
}
}
if(window.HermesExtensionSettings&&typeof window.HermesExtensionSettings._onConfigureChange==='function'){
window.HermesExtensionSettings._onConfigureChange(_handleExtensionConfigureChange);
}
function _extensionSafeHttpUrl(value){
if(!value) return '';
const raw=String(value).trim();
@@ -10517,6 +10574,19 @@ function _extensionPostInstallNote(entry,isInstalled){
</div>`;
}
function _renderInstalledExtensionsSurface(statusData){
const installedEl=$('extensionsInstalled');
if(!installedEl) return;
installedEl.innerHTML=_extensionInstalledList(
statusData&&statusData.extensions,
!!(statusData&&statusData.extension_dir_configured),
'installed'
);
_bindExtensionToggleButtons(installedEl);
_bindExtensionSettingsButtons(installedEl);
_bindExtensionConfigureButtons(installedEl);
}
async function loadExtensionsGallery(){
_extensionsGalleryLoaded=true;
const galleryEl=$('extensionsGallery');
@@ -10540,7 +10610,6 @@ async function loadExtensionsGallery(){
function _renderExtensionsGallery(entries,statusData){
const galleryEl=$('extensionsGallery');
const installedEl=$('extensionsInstalled');
_configureExtensionSettingsFromStatus(statusData);
const installedIds=new Set();
if(statusData&&statusData.gallery_installed){
@@ -10551,11 +10620,7 @@ function _renderExtensionsGallery(entries,statusData){
}
if(!Array.isArray(entries)||entries.length===0){
if(galleryEl) galleryEl.innerHTML='<div class="extensions-empty">No extensions found in the registry.</div>';
if(installedEl){
installedEl.innerHTML=_extensionInstalledList(statusData&&statusData.extensions,!!(statusData&&statusData.extension_dir_configured));
_bindExtensionToggleButtons(installedEl);
_bindExtensionSettingsButtons(installedEl);
}
_renderInstalledExtensionsSurface(statusData);
return;
}
const galleryCards=[];
@@ -10599,11 +10664,7 @@ function _renderExtensionsGallery(entries,statusData){
galleryCards.push(card);
}
if(galleryEl) galleryEl.innerHTML=galleryCards.length?galleryCards.join(''):'<div class="extensions-empty">No extensions found.</div>';
if(installedEl){
installedEl.innerHTML=_extensionInstalledList(statusData&&statusData.extensions,!!(statusData&&statusData.extension_dir_configured));
_bindExtensionToggleButtons(installedEl);
_bindExtensionSettingsButtons(installedEl);
}
_renderInstalledExtensionsSurface(statusData);
_bindExtensionGalleryButtons(entries);
}
@@ -11971,7 +12032,7 @@ async function loadPasskeys(){
}
const creds=(data&&data.credentials)||[];
if(!creds.length){list.textContent='No passkeys registered.';return;}
list.innerHTML=creds.map(c=>`<div style="display:flex;align-items:center;justify-content:space-between;gap:8px;border:1px solid var(--border);border-radius:8px;padding:8px;margin-top:6px"><span>${esc(c.label||'Passkey')}</span><button class="btn-tiny" onclick="deletePasskey('${esc(c.id)}')">Remove</button></div>`).join('');
list.innerHTML=creds.map(c=>`<div style="display:flex;align-items:center;justify-content:space-between;gap:8px;border:1px solid var(--border);border-radius:8px;padding:8px;margin-top:6px"><span>${esc(c.label||'Passkey')}</span><button class="btn-tiny" onclick="deletePasskey(${jsArg(c.id)})">Remove</button></div>`).join('');
}catch(e){list.textContent='Failed to load passkeys: '+e.message;}
}
@@ -12128,7 +12189,7 @@ async function checkUpdatesNow(channelOverride){
// saved setting. (Fable UX gate.)
const _checkBody={force:true};
if(channelOverride==='stable'||channelOverride==='experimental') _checkBody.channel=channelOverride;
const data=await api('/api/updates/check',{method:'POST',body:JSON.stringify(_checkBody),timeoutMs:60000});
const data=await api('/api/updates/check',{method:'POST',body:JSON.stringify(_checkBody),timeoutMs:300000});
if(data.disabled){
if(status){status.textContent=t('settings_updates_disabled');status.style.color='var(--muted)';}
} else {
@@ -12254,12 +12315,23 @@ function _buildAuxProviderOptions(sel,providers,currentProvider){
autoOpt.value='auto';autoOpt.textContent='auto ('+t('settings_aux_provider_auto')+')';
if(currentProvider==='auto'||!currentProvider) autoOpt.selected=true;
sel.appendChild(autoOpt);
let matched=currentProvider==='auto'||!currentProvider;
for(const p of providers){
const opt=document.createElement('option');
opt.value=p.slug;opt.textContent=p.name;
if(p.slug===currentProvider) opt.selected=true;
if(p.slug===currentProvider){opt.selected=true;matched=true;}
sel.appendChild(opt);
}
// The configured provider can be absent from the /api/models catalog (e.g. its
// group exposes no models). Keep it selectable: with no matching option the
// select falls back to its first entry ('auto') and the next Apply would
// persist that, silently discarding the configured value.
if(!matched&&currentProvider){
const configuredOpt=document.createElement('option');
configuredOpt.value=currentProvider;configuredOpt.textContent=currentProvider+' (configured)';
configuredOpt.selected=true;
sel.appendChild(configuredOpt);
}
}
function _buildAuxModelOptions(sel,provider,providers,currentModel){
@@ -12267,17 +12339,35 @@ function _buildAuxModelOptions(sel,provider,providers,currentModel){
const emptyOpt=document.createElement('option');
emptyOpt.value='';emptyOpt.textContent=t('settings_aux_model_auto')||'auto (use provider default)';
sel.appendChild(emptyOpt);
const canonicalCurrent=_modelBareNameForProvider(currentModel,provider)||'';
if(!provider||provider==='auto'){
sel.value=currentModel||'';
return;
sel.value=canonicalCurrent;
return canonicalCurrent;
}
// Find matching provider in cached list
const pData=providers.find(p=>p.slug===provider);
// A provider kept in the list only because its models endpoint failed would
// otherwise render as a silent, empty model select. Echo the same hint the
// main picker shows instead of implying "no models to choose from". (#7521)
if(pData&&pData.modelsEndpointError){
const errOpt=document.createElement('option');
errOpt.value='';errOpt.disabled=true;
errOpt.dataset.modelsEndpointError='1';
errOpt.textContent='\u26a0 '+(pData.modelsEndpointError.message||'Models endpoint could not be reached for this provider.');
sel.appendChild(errOpt);
}
const modelValues=new Set();
if(pData&&pData.models){
for(const mId of pData.models){
for(const modelEntry of pData.models){
const routeId=typeof modelEntry==='string'?modelEntry:String(modelEntry?.id||'');
const mId=_modelBareNameForProvider(routeId,provider)||'';
if(!mId||modelValues.has(mId)) continue;
modelValues.add(mId);
const routeLabel=typeof modelEntry==='string'?'':String(modelEntry?.label||'');
const modelLabel=_modelBareNameForProvider(routeLabel,provider)||mId;
const opt=document.createElement('option');
opt.value=mId;opt.textContent=mId;
if(mId===currentModel) opt.selected=true;
opt.value=mId;opt.textContent=modelLabel;
if(mId===canonicalCurrent) opt.selected=true;
sel.appendChild(opt);
}
}
@@ -12286,12 +12376,13 @@ function _buildAuxModelOptions(sel,provider,providers,currentModel){
customOpt.value='__custom__';customOpt.textContent=t('settings_aux_model_custom')||'Custom model…';
sel.appendChild(customOpt);
// If currentModel not in list and not empty, add it as a custom option
if(currentModel&&!pData?.models?.includes(currentModel)){
if(canonicalCurrent&&!modelValues.has(canonicalCurrent)){
const existingOpt=document.createElement('option');
existingOpt.value=currentModel;existingOpt.textContent=currentModel+' (configured)';
existingOpt.value=canonicalCurrent;existingOpt.textContent=canonicalCurrent+' (configured)';
existingOpt.selected=true;
sel.insertBefore(existingOpt,customOpt);
}
return canonicalCurrent;
}
function _onAuxProviderChange(taskKey,providers){
@@ -12309,13 +12400,16 @@ async function _onAuxModelChange(taskKey){
if(modelSel.value==='__custom__'){
const customModel=await showPromptDialog({title:t('settings_aux_model_custom')||'Custom model',message:t('settings_aux_model_custom_prompt')||'Enter model ID:',placeholder:'model/provider:model-id',confirmLabel:t('settings_btn_apply_aux_models')||'Apply'});
if(customModel&&customModel.trim()){
const provider=$('aux-prov-'+taskKey)?.value||'';
const enteredModel=customModel.trim();
const canonicalModel=_modelBareNameForProvider(enteredModel,provider)||enteredModel;
// Insert custom model option before the __custom__ option
const opt=document.createElement('option');
opt.value=customModel.trim();opt.textContent=customModel.trim();
opt.value=canonicalModel;opt.textContent=canonicalModel;
// Remove __custom__ selection
const customIdx=[...modelSel.options].findIndex(o=>o.value==='__custom__');
if(customIdx>=0) modelSel.insertBefore(opt,modelSel.options[customIdx]);
modelSel.value=customModel.trim();
modelSel.value=canonicalModel;
}else{
modelSel.value='';
}
@@ -12505,6 +12599,22 @@ function _bindMainAdvancedOptionsButton(){
btn.addEventListener('click',()=>{if(_mainAdvancedConfig!==null)_openAuxAdvancedOptions('__main__',_mainAdvancedConfig||{});});
}
// Build the auxiliary picker provider list from /api/models groups.
// A named custom provider whose /v1/models probe failed still reaches the UI as
// a group with an empty ``models`` list plus ``models_endpoint_error``
// (api/config.py). Zero-model groups used to be filtered out here, which made
// the provider vanish from every auxiliary select even though the main model
// picker renders that same group together with its unreachable-endpoint hint. (#7521)
function _auxProvidersFromModelGroups(groups){
const list=Array.isArray(groups)?groups:[];
return list.filter(g=>g&&g.provider&&((g.models&&g.models.length>0)||(g.extra_models&&g.extra_models.length>0)||g.models_endpoint_error)).map(g=>({
slug:g.provider_id||g.provider,
name:g.provider,
modelsEndpointError:g.models_endpoint_error||null,
models:[...(g.models||[]),...(g.extra_models||[])].map(m=>({id:m.id,label:m.label||m.id})),
}));
}
async function _loadAuxiliaryModels(){
const container=$('auxModelsContainer');
if(!container) return;
@@ -12519,11 +12629,7 @@ async function _loadAuxiliaryModels(){
// Build provider list from /api/models groups
// /api/models returns: { groups: [{ provider: str, provider_id: str, models: [{id,label}] }] }
const groups=(modelsData&&modelsData.groups)||[];
_auxProviders=groups.filter(g=>g.provider&&((g.models&&g.models.length>0)||(g.extra_models&&g.extra_models.length>0))).map(g=>({
slug:g.provider_id||g.provider,
name:g.provider,
models:[...(g.models||[]),...(g.extra_models||[])].map(m=>m.id),
}));
_auxProviders=_auxProvidersFromModelGroups(groups);
if(auxData&&Object.prototype.hasOwnProperty.call(auxData,'main')){
_mainAdvancedConfig=auxData.main||{};
}else{
@@ -12537,6 +12643,7 @@ async function _loadAuxiliaryModels(){
_auxOriginalConfig=JSON.parse(JSON.stringify(taskMap));
container.innerHTML='';
let needsCanonicalSave=false;
for(const task of _auxTasks){
const cfg=taskMap[task.task]||{provider:'auto',model:''};
const row=document.createElement('div');
@@ -12560,7 +12667,8 @@ async function _loadAuxiliaryModels(){
const modelSel=document.createElement('select');
modelSel.id='aux-model-'+task.task;
modelSel.style.cssText=_auxSelectStyle();
_buildAuxModelOptions(modelSel,cfg.provider,_auxProviders,cfg.model);
const canonicalModel=_buildAuxModelOptions(modelSel,cfg.provider,_auxProviders,cfg.model);
if(canonicalModel!==cfg.model) needsCanonicalSave=true;
modelSel.addEventListener('change',()=>_onAuxModelChange(task.task));
row.appendChild(modelSel);
@@ -12577,9 +12685,9 @@ async function _loadAuxiliaryModels(){
container.appendChild(row);
}
// Hide apply button (no changes yet)
// Matching legacy @provider:model values can be repaired with one explicit Apply.
const applyBtn=$('btnApplyAuxModels');
if(applyBtn) applyBtn.style.display='none';
if(applyBtn) applyBtn.style.display=needsCanonicalSave?'':'none';
// Reset button
const resetBtn=$('btnResetAuxModels');
@@ -12624,7 +12732,13 @@ async function _applyAuxModels(){
saved++;
}catch(e){
console.warn('[settings] failed to save aux task',task.task,e);
if(typeof showToast==='function') showToast(t('settings_aux_save_failed')||'Failed to save auxiliary model');
// Surface the server's actionable message (e.g. an ambiguous custom-provider
// slug collision: rename one provider so its slug is unique) instead of a
// generic failure, and abort the loop so the dirty selection is retained for
// the user to fix and retry — the reload that would clear it is skipped.
const _msg=(e&&e.message)?e.message:'';
const _base=t('settings_aux_save_failed')||'Failed to save auxiliary model';
if(typeof showToast==='function') showToast(_msg?(_base+': '+_msg):_base,6000,'error');
return;
}
}
@@ -12738,11 +12852,17 @@ async function saveSettings(andClose){
const saved=await _enqueueSettingsPost({method:'POST',body:JSON.stringify(payload)});
if(modelChanged && model){
try{
await api('/api/default-model',{method:'POST',body:JSON.stringify({model,provider:modelState.model_provider||null})});
body.default_model=model;
body.default_model_provider=(modelState&&modelState.model===model)?(modelState.model_provider||null):null;
await api('/api/default-model',{method:'POST',body:JSON.stringify({model,provider:modelState.model_provider||null})});
body.default_model=model;
body.default_model_provider=(modelState&&modelState.model===model)?(modelState.model_provider||null):null;
}catch(_modelErr){
if(typeof showToast==='function') showToast('Failed to update default model — settings saved');
// A 400 here (e.g. an ambiguous custom-provider slug collision: rename
// one provider) is user-fixable, not a partial success. Surface the
// message, abort before "settings saved", and retain dirty state so the
// user can fix and retry instead of the error being swallowed.
const _msg=(_modelErr&&_modelErr.message)?_modelErr.message:'';
if(typeof showToast==='function') showToast('Failed to update default model'+(_msg?(': '+_msg):''),6000,'error');
return;
}
}
_applySavedSettingsUi(saved, body, {sendKey,showTokenUsage,showQuotaChip,showConversationOutline,showBusyPlaceholderHint,showTps,fadeTextEffect,showCliSessions,theme,skin,language,sidebarDensity,fontSize});
@@ -12771,9 +12891,15 @@ async function saveSettings(andClose){
await api('/api/default-model',{method:'POST',body:JSON.stringify({model,provider:modelState.model_provider||null})});
body.default_model=model;
body.default_model_provider=(modelState&&modelState.model===model)?(modelState.model_provider||null):null;
}catch(_modelErr){
if(typeof showToast==='function') showToast('Failed to update default model — settings saved');
}
}catch(_modelErr){
// A 400 here (e.g. an ambiguous custom-provider slug collision: rename
// one provider) is user-fixable, not a partial success. Surface the
// message, abort before "settings saved", and retain dirty state so the
// user can fix and retry instead of the error being swallowed.
const _msg=(_modelErr&&_modelErr.message)?_modelErr.message:'';
if(typeof showToast==='function') showToast('Failed to update default model'+(_msg?(': '+_msg):''),6000,'error');
return;
}
}
_applySavedSettingsUi(saved, body, {sendKey,showTokenUsage,showQuotaChip,showConversationOutline,showBusyPlaceholderHint,showTps,fadeTextEffect,showCliSessions,theme,skin,language,sidebarDensity,fontSize});
showToast(t('settings_saved'));
@@ -13187,7 +13313,7 @@ function loadMcpTools(){
let _gatewayActionInFlight=false;
function _gatewayActionButton(action){
const labels={start:t('gateway_start'),stop:t('gateway_stop'),restart:t('gateway_restart')};
return `<button class="sm-btn gateway-action-btn" data-gateway-action="${esc(action)}" onclick="_gatewayAction('${esc(action)}')" ${_gatewayActionInFlight?'disabled':''} style="padding:5px 10px;font-size:12px">${esc(labels[action]||action)}</button>`;
return `<button class="sm-btn gateway-action-btn" data-gateway-action="${esc(action)}" onclick="_gatewayAction(${jsArg(action)})" ${_gatewayActionInFlight?'disabled':''} style="padding:5px 10px;font-size:12px">${esc(labels[action]||action)}</button>`;
}
function _gatewayActionControls(r){
const actions=(r&&r.running)?['stop','restart']:['start'];
@@ -13280,10 +13406,10 @@ async function _loadCheckpoints(workspace){
</div>
</div>
<div style="display:flex;gap:4px;flex-shrink:0;margin-left:8px">
<button class="panel-head-btn" title="${esc(t('checkpoint_view_diff'))}" onclick="event.stopPropagation();_viewCheckpointDiff('${esc(workspace)}','${esc(ck.id)}')">
<button class="panel-head-btn" title="${esc(t('checkpoint_view_diff'))}" onclick="event.stopPropagation();_viewCheckpointDiff(${jsArg(workspace)},${jsArg(ck.id)})">
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="16" height="16"><path d="M12 20h9"/><path d="M16.5 3.5a2.121 2.121 0 0 1 3 3L7 19l-4 1 1-4L16.5 3.5z"/></svg>
</button>
<button class="panel-head-btn" title="${esc(t('checkpoint_restore'))}" onclick="event.stopPropagation();_restoreCheckpoint('${esc(workspace)}','${esc(ck.id)}','${esc(msg.replace(/'/g,"\\'"))}')">
<button class="panel-head-btn" title="${esc(t('checkpoint_restore'))}" onclick="event.stopPropagation();_restoreCheckpoint(${jsArg(workspace)},${jsArg(ck.id)},${jsArg(msg)})">
<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" width="16" height="16"><polyline points="1 4 1 10 7 10"/><path d="M3.51 15a9 9 0 1 0 2.13-9.36L1 10"/></svg>
</button>
</div>
+152 -17
View File
@@ -189,6 +189,15 @@ function _clearRememberedNewChatDraftSession(sid) {
} catch (_) {}
}
function _adoptRegenerationRevision(sessionPayload){
if(!S||!S.session||!sessionPayload||typeof sessionPayload!=='object') return;
if(Object.prototype.hasOwnProperty.call(sessionPayload,'regeneration_revision')){
S.session.regeneration_revision=sessionPayload.regeneration_revision;
}else{
delete S.session.regeneration_revision;
}
}
async function _restoreRememberedNewChatDraftSession() {
let sid = '';
try { sid = localStorage.getItem(NEW_CHAT_DRAFT_SESSION_KEY) || ''; } catch (_) { sid = ''; }
@@ -1400,15 +1409,20 @@ async function newSession(flash, options={}){
_messagesTruncated=false;
_oldestIdx=0;
clearLiveToolCards();
// One-shot profile-switch workspace wins first; otherwise prefer the profile default.
// Explicit profile switch wins, then the current conversation, then the profile default.
// Provenance lets the server recover only a deleted inherited path; explicit paths stay strict.
const switchWs=S._profileSwitchWorkspace;
S._profileSwitchWorkspace=null;
const inheritWs=switchWs||(S._profileDefaultWorkspace||null)||(S.session?S.session.workspace:null);
const sessionWs=(!switchWs&&S.session)?S.session.workspace:null;
const inheritWs=switchWs||sessionWs||(S._profileDefaultWorkspace||null);
const reqBody={
workspace:inheritWs,
profile:S.activeProfile||'default',
};
if(S.session&&S.session.session_id) reqBody.prev_session_id=S.session.session_id;
if(S.session&&S.session.session_id){
reqBody.prev_session_id=S.session.session_id;
if(sessionWs) reqBody.workspace_inherited_from_prev_session=true;
}
// Three-value worktree contract (#6022): explicit true/false is forwarded
// verbatim; an ABSENT key lets the server apply the agent's config-level
// `worktree:` default. Auto-bind paths pass worktree:false explicitly so a
@@ -1444,7 +1458,7 @@ async function newSession(flash, options={}){
}
if(newModelState&&newModelState.model){
reqBody.model=newModelState.model;
// Cold-start / picker-without-provider fallback: when the dropdown option's
// Cold-start / picker-without-provider fallback (#2518): when the dropdown option's
// data-provider is empty/'default' or the persisted state predates provider
// tracking, newModelState.model_provider is null. POST /api/session/new's
// fast path in _resolve_compatible_session_model_state requires both model
@@ -1491,7 +1505,7 @@ async function newSession(flash, options={}){
if(consumedExplicitModelOverride&&typeof _clearEmptyComposerModelOverride==='function'){
_clearEmptyComposerModelOverride();
}
S.session=data.session;S.messages=data.session.messages||[];
S.session=data.session;if(typeof _adoptRegenerationRevision==='function') _adoptRegenerationRevision(data.session);S.messages=data.session.messages||[];
S._pendingSessionToolsets=null;
if(_sessionSourceFilter==='cli') _sessionSourceFilter='webui';
if(typeof _hydrateTodosFromSession==='function') _hydrateTodosFromSession(S.session);
@@ -1703,6 +1717,12 @@ async function loadSession(sid){
// #2971: idempotent re-arm before the no-op guard revives a stream a prior
// failed loadSession killed; no-ops on real switches.
_rearmActiveSessionStream();
// #6999: same-session force-reload coordination lives in the refresh paths
// (refreshActiveSessionIfExternallyUpdated guard + session-updated SSE
// handler in messages.js), NOT here: a second loadSession(sid,{force:true})
// for the same sid is a legitimate supersede (generation bump below) that
// cross-session ordering tests rely on. Coalescing at the entry point would
// drop the superseding fetch and leave a stale first load in charge.
if(currentSid===sid && !forceReload && (!_loadingSessionId || _loadingSessionId===sid)){
// Re-selecting the already-open session is a no-op for transcript/scroll, but
// it is still a *visit*: clear a stale sidebar unread dot (e.g. one a
@@ -1968,6 +1988,7 @@ async function loadSession(sid){
return loadSession(continuationSid,{...opts,skipLineageResolve:true,skipContinuationResolve:true,force:true,_preloadNotified:true});
}
S.session=data.session;
if(typeof _adoptRegenerationRevision==='function') _adoptRegenerationRevision(data.session);
if(typeof _clearEmptyComposerModelOverride==='function') _clearEmptyComposerModelOverride();
// Loading a real existing session abandons any pre-session toolset override
// staged on the empty composer before any deferred refresh work runs.
@@ -2381,7 +2402,7 @@ async function loadSession(sid){
);
}
if(typeof renderSessionArtifacts==='function') renderSessionArtifacts();
if(typeof projectSessionArtifactsForOwner==='function') projectSessionArtifactsForOwner(sid);
// ── Cross-channel handoff hint ──
// After session fully loaded, check if this is a messaging session with
@@ -2403,7 +2424,7 @@ const _HANDOFF_THRESHOLD = 10; // conversation rounds
const _HANDOFF_STORAGE_PREFIX = 'handoff:';
const _HANDOFF_SUFFIX_DISMISSED_AT = 'dismissed_at';
const _HANDOFF_SUFFIX_SUMMARY_HANDLED_AT = 'summary_handled_at';
const _MESSAGING_RAW_SOURCES = new Set(['weixin', 'telegram', 'discord', 'slack', 'email', 'wecom', 'wecom_callback', 'matrix']);
const _MESSAGING_RAW_SOURCES = new Set(['weixin', 'telegram', 'discord', 'slack', 'email', 'wecom', 'wecom_callback', 'matrix', 'signal']);
const _MESSAGING_SOURCE_LABELS = {
weixin: 'WeChat',
telegram: 'Telegram',
@@ -2413,6 +2434,7 @@ const _MESSAGING_SOURCE_LABELS = {
wecom: 'WeCom',
wecom_callback: 'WeCom Callback',
matrix: 'Matrix',
signal: 'Signal',
};
function _isMessagingSession(session) {
@@ -2785,7 +2807,7 @@ function _showHandoffHint(sid, rounds) {
</div>
<div class="handoff-hint-actions">
<button class="handoff-hint-action" type="button">View summary</button>
<button class="handoff-hint-dismiss" type="button" onclick="event.stopPropagation(); _dismissHandoffHint('${esc(sid)}')" title="Dismiss">
<button class="handoff-hint-dismiss" type="button" onclick="event.stopPropagation(); _dismissHandoffHint(${jsArg(sid)})" title="Dismiss">
Close
</button>
</div>
@@ -2902,12 +2924,20 @@ async function _generateHandoffSummary(sid, rounds) {
} catch (e) {
console.warn('Handoff summary failed:', e);
if (S.session && S.session.session_id === sid && typeof setHandoffUi === 'function') {
// A 400 carries an actionable, user-fixable message (e.g. an ambiguous
// custom-provider slug collision: rename one provider so its slug is
// unique). Surface it verbatim rather than degrading to the generic
// "try again" card — previously the server answered 200 with a warning
// that this handler ignored, hiding the fix from the user.
const errorText = (e && e.status === 400 && e.message)
? e.message
: ('Summary generation failed: ' + (e && e.message ? e.message : 'unknown error'));
setHandoffUi({
sessionId: sid,
phase: 'error',
channel,
rounds,
errorText: 'Summary generation failed: ' + e.message,
errorText,
});
}
}
@@ -3187,9 +3217,23 @@ async function _ensureMessagesLoaded(sid, opts) {
// Expand render window to cover all loaded messages so the next
// renderMessages() doesn't hide most of them behind a tiny window.
if(typeof _messageRenderableMessageCount==='function'&&typeof _currentMessageRenderWindowSize==='function'){
_messageRenderWindowSize=Math.max(_currentMessageRenderWindowSize(), _messageRenderableMessageCount());
// #6999: bound the auto-expansion. This number gates
// _messageHiddenBeforeCount() (load-older / jump-to-start affordances)
// and the non-virtualized fallback render width; the virtualized DOM tail
// is independently capped at MESSAGE_RENDER_WINDOW_DEFAULT via
// _messageVirtualKeepTailCount(). Growing the window to the FULL loaded
// transcript on every force reload (tab focus, SSE catch-up) zeroes the
// hidden-before count for long sessions and widens the effective render
// window for non-virtualized paths. Keep the #3686 intent (don't collapse
// back to the 50-row default after a load) but cap the growth to a small
// multiple of the default window.
_messageRenderWindowSize=Math.max(
_currentMessageRenderWindowSize(),
Math.min(_messageRenderableMessageCount(), (typeof MESSAGE_RENDER_WINDOW_DEFAULT==='number'?MESSAGE_RENDER_WINDOW_DEFAULT:50)*4)
);
}
if(S.session&&S.session.session_id===sid){
if(typeof _adoptRegenerationRevision==='function') _adoptRegenerationRevision(data.session);
S.session.message_count=Number(data.session.message_count || msgs.length);
S.lastUsage={...(data.session.last_usage||S.lastUsage||{})};
// Phase 2: the messages=1 response carries the canonical cold-load
@@ -3886,6 +3930,11 @@ async function _ensureAllMessagesLoaded() {
_syncToolCallsForLoadedMessages(msgs, data.session.tool_calls);
if (S.session && S.session.session_id === sid) {
S.session.message_count = Number(data.session.message_count || msgs.length);
if (Object.prototype.hasOwnProperty.call(data.session, 'regeneration_revision')) {
S.session.regeneration_revision = data.session.regeneration_revision;
} else {
delete S.session.regeneration_revision;
}
}
} finally {
_loadingOlder = false;
@@ -5730,6 +5779,12 @@ let _sessionTimeRefreshVisibilityHandler = null;
let _activeSessionExternalRefreshTimer = null;
let _activeSessionExternalRefreshInFlight = false;
let _deferredActiveSessionExternalRefreshReason = '';
// #6999 re-gate: per-SID latch of the maximum message_count announced by a
// `session-updated` frame while the external-refresh guard was held. The
// refresh owner runs ONE guarded follow-up in its finally instead of letting
// the event die — production does not guarantee a second event, so discarding
// would leave the transcript stale until the next focus/poll.
let _pendingSessionUpdatedCounts = null;
let _sessionEventsSSE = null;
let _sessionEventsRefreshTimer = 0;
let _sessionEventsRefreshPendingRequest = null;
@@ -5822,6 +5877,51 @@ function _flushDeferredActiveSessionExternalRefresh(){
void refreshActiveSessionIfExternallyUpdated(reason);
}
// #6999 re-gate: coalesce, never discard. Called by the messages.js
// `session-updated` handler when it finds the external-refresh guard held:
// instead of dropping the update (production does not guarantee a second
// event), latch the MAXIMUM announced count per SID. The refresh owner's
// finally drains it with ONE guarded follow-up.
function _latchSessionUpdatedPendingCount(sid, count){
const n = Number(count);
if(!sid || !Number.isFinite(n)) return;
if(!_pendingSessionUpdatedCounts) _pendingSessionUpdatedCounts = {};
const prev = _pendingSessionUpdatedCounts[sid];
if(prev === undefined || n > prev) _pendingSessionUpdatedCounts[sid] = n;
}
// Single-line entry used by the messages.js `session-updated` handler:
// returns true when the external-refresh probe owns the guard (the frame was
// latched for the owner's finally to drain); returns false when the handler
// should take its normal direct-load path.
function _coalesceSessionUpdatedWhileRefreshHeld(sid, count){
if(typeof _activeSessionExternalRefreshInFlight === 'undefined' || !_activeSessionExternalRefreshInFlight) return false;
_latchSessionUpdatedPendingCount(sid, count);
return true;
}
// Drain the per-SID latch after the external-refresh owner releases its
// guard. Runs ONE guarded follow-up only when local state is still behind
// the latched count — a same-SID load during the refresh window already
// caught us up, a switch-away moved the active session elsewhere, and
// duplicate events coalesce into this single follow-up.
function _drainSessionUpdatedPendingCount(){
const pending = _pendingSessionUpdatedCounts;
_pendingSessionUpdatedCounts = null;
if(!pending || !S.session || !S.session.session_id) return;
const sid = S.session.session_id;
const latched = pending[sid];
if(latched === undefined) return;
const localCount = Number(S.session.message_count || (Array.isArray(S.messages)?S.messages.length:0) || 0);
if(Number.isFinite(localCount) && localCount >= latched) return;
// The follow-up re-enters refreshActiveSessionIfExternallyUpdated, which
// re-probes server metadata and only force-reloads when the count actually
// changed — so the OOM guards (busy/stream/loading/document.hidden) all
// apply, and a metadata-read-only probe that already ran during the refresh
// window gets a second chance to observe the latched growth.
void refreshActiveSessionIfExternallyUpdated('session-updated');
}
// Reconcile the active session against server-side metadata. Returns a status
// string so callers (notably the post-stream idle reconcile) can decide how to
// react:
@@ -5846,6 +5946,10 @@ async function refreshActiveSessionIfExternallyUpdated(reason){
if(_activeSessionExternalRefreshInFlight) return 'skipped';
if(!S.session || !S.session.session_id) return 'skipped';
if(S.busy || S.activeStreamId) return 'skipped';
// #6999: if a load for this exact session is already in flight, it owns the
// refresh — probing now would duplicate the fetch and the O(N) render work
// that exhausts the tab's JS heap when the tab comes back into focus.
if(_loadingSessionId === S.session.session_id) return 'skipped';
if(typeof _isMessageReaderUnpinned==='function'&&_isMessageReaderUnpinned()){
_deferActiveSessionExternalRefresh(reason||'poll');
return 'skipped';
@@ -5930,6 +6034,10 @@ async function refreshActiveSessionIfExternallyUpdated(reason){
return 'failed';
}finally{
_activeSessionExternalRefreshInFlight = false;
// #6999 re-gate: any `session-updated` frames that arrived while we owned
// the guard were latched (coalesced), not dropped — run ONE guarded
// follow-up when local state is still behind the latched count.
_drainSessionUpdatedPendingCount();
}
}
@@ -7015,6 +7123,8 @@ function _attachChildSessionsToSidebarRows(collapsedRows, rawSessions, rawRefere
};
const orphans=[];
const renderableChildIds=new Set((rawSessions||[]).map(s=>s&&s.session_id).filter(Boolean));
const childAttachOrderById=new Map();
let childAttachCursor=0;
const attachQueueById=new Map();
for(const candidate of [...(rawSessions||[]),...(referenceSessions||[])]){
if(candidate&&candidate.session_id&&!attachQueueById.has(candidate.session_id)) attachQueueById.set(candidate.session_id,candidate);
@@ -7024,6 +7134,12 @@ function _attachChildSessionsToSidebarRows(collapsedRows, rawSessions, rawRefere
const childRenderable=!!(child&&child.session_id&&renderableChildIds.has(child.session_id));
if(child&&child.session_id&&visibleBySid.has(child.session_id)) continue;
const isForkChild=_isForkWithResolvableParent(child, sessionIdsInList)&&!(child&&child.pinned);
const childRawRole=[
child&&child.raw_source,
child&&child.source_tag,
child&&child.source,
].map(source=>String(source||'').trim().toLowerCase()).find(Boolean)||'';
const childIsDelegatedSubagent=_isChildSession(child)&&childRawRole==='subagent';
const childLineageKey=child&&(child._lineage_root_id||child.lineage_root_id||child.parent_session_id);
const isHiddenLineageReferenceChild=!!(child&&child.archived&&child.parent_session_id&&childLineageKey&&!child.pinned&&!childRenderable);
if(!_isChildSession(child)&&!isForkChild&&!isHiddenLineageReferenceChild) continue;
@@ -7048,11 +7164,10 @@ function _attachChildSessionsToSidebarRows(collapsedRows, rawSessions, rawRefere
hiddenArchivedChildTree.add(child.session_id);
continue;
}
// Cross-surface rows (for example a WebUI continuation from a Telegram
// conversation) should remain top-level when there is no WebUI-owned parent
// row to stack under. But if the parent is visible in this same sidebar
// render, attach normally — delegated subagent rows are also cross-source
// relative to their WebUI parent and should not be forced into orphans.
// Independent cross-surface rows (for example a WebUI continuation from a
// Telegram conversation) remain top-level instead of nesting under an
// external parent. Delegated subagents are also cross-source, but they are
// parent-owned work and should still attach to the visible parent row.
const parentSourceMarker=String(parentRow&&(
parentRow.session_source||parentRow.raw_source||parentRow.source_tag||parentRow.source
)||'').toLowerCase();
@@ -7063,12 +7178,15 @@ function _attachChildSessionsToSidebarRows(collapsedRows, rawSessions, rawRefere
parentRow.session_source==='messaging'||
(parentSourceMarker&&parentSourceMarker!=='webui'&&parentSourceMarker!=='subagent'&&parentSourceMarker!=='other'&&parentSourceMarker!=='fork')
);
if(parentRow&&child._cross_surface_child_session&&parentIsExternal){
if(parentRow&&child._cross_surface_child_session&&parentIsExternal&&!childIsDelegatedSubagent){
if(childRenderable) orphans.push({...child,_orphan_child_session:true});
continue;
}
if(parentRow){
const childCopy={...child};
if(!childAttachOrderById.has(childCopy.session_id)){
childAttachOrderById.set(childCopy.session_id, childAttachCursor++);
}
if(parentSegment){
childCopy._parent_segment_id=parentSegment.session_id;
childCopy._parent_segment_title=_sessionDisplayTitle(parentSegment)||child.parent_title||'Untitled';
@@ -7097,6 +7215,23 @@ function _attachChildSessionsToSidebarRows(collapsedRows, rawSessions, rawRefere
orphans.push({...child,_orphan_child_session:true});
}
}
const resolveReadOnlySession = typeof _isReadOnlySession === 'function'
? _isReadOnlySession
: ((session) => !!(session && session.read_only));
for(const row of rows){
if(Array.isArray(row._child_sessions)&&row._child_sessions.length>1){
row._child_sessions.sort((a,b)=>{
const readOnlyCmp = Number(resolveReadOnlySession(a))-Number(resolveReadOnlySession(b));
if(readOnlyCmp!==0) return readOnlyCmp;
const aOrder = childAttachOrderById.get(a&&a.session_id) ?? Number.MAX_SAFE_INTEGER;
const bOrder = childAttachOrderById.get(b&&b.session_id) ?? Number.MAX_SAFE_INTEGER;
return aOrder-bOrder;
});
}
if(Array.isArray(row._child_sessions)){
row._child_session_count=row._child_sessions.length;
}
}
return [...rows,...orphans];
}
@@ -8214,7 +8349,7 @@ function renderSessionListFromCache(){
const childList=document.createElement('div');
childList.className='session-child-sessions';
['pointerdown','pointerup','click','touchstart','touchmove','touchend','touchcancel'].forEach(ev=>childList.addEventListener(ev,e=>e.stopPropagation()));
const sortedChildren=[...s._child_sessions].sort((a,b)=>_sessionTimestampMs(b)-_sessionTimestampMs(a));
const sortedChildren=[...s._child_sessions];
const openChildSession=async(childSession)=>{
await _openSidebarSession(childSession, {skipLineageResolve:true});
};
+1 -1
View File
@@ -6,7 +6,7 @@
<title>Shared Conversation - Hermes</title>
<meta name="robots" content="noindex,nofollow,noarchive">
<meta name="referrer" content="same-origin">
<script>(function(){try{var themes={light:1,dark:1,system:1},skins={codex:1,terracotta:1,default:1,ares:1,mono:1,graphite:1,github:1,slate:1,poseidon:1,sisyphus:1,charizard:1,sienna:1,catppuccin:1,hepburn:1,nous:1,'geist-contrast':1,neon:1,zeus:1,'verdigris':1},legacy={slate:['dark','slate'],solarized:['dark','poseidon'],monokai:['dark','sisyphus'],nord:['dark','slate'],oled:['dark','default']},t=(localStorage.getItem('hermes-theme')||'dark').toLowerCase(),s=(localStorage.getItem('hermes-skin')||'').toLowerCase(),m=legacy[t],theme=m?m[0]:(themes[t]?t:'dark');var pendingExt=s&&s!=='default'&&!skins[s]&&!legacy[s];var skin=skins[s]?s:(pendingExt?s:(m?m[1]:'default'));localStorage.setItem('hermes-theme',theme);localStorage.setItem('hermes-skin',skin);if(theme==='system')theme=window.matchMedia('(prefers-color-scheme:dark)').matches?'dark':'light';if(theme==='dark')document.documentElement.classList.add('dark');if(skin!=='default')document.documentElement.dataset.skin=skin;}catch(e){document.documentElement.classList.add('dark');}})()</script>
<script>(function(){try{var themes={light:1,dark:1,system:1},skins={codex:1,terracotta:1,default:1,ares:1,mono:1,graphite:1,github:1,slate:1,poseidon:1,sisyphus:1,charizard:1,sienna:1,catppuccin:1,hepburn:1,nous:1,'geist-contrast':1,neon:1,'neon-soft':1,'neon-paint':1,zeus:1,'verdigris':1},legacy={slate:['dark','slate'],solarized:['dark','poseidon'],monokai:['dark','sisyphus'],nord:['dark','slate'],oled:['dark','default']},t=(localStorage.getItem('hermes-theme')||'dark').toLowerCase(),s=(localStorage.getItem('hermes-skin')||'').toLowerCase(),m=legacy[t],theme=m?m[0]:(themes[t]?t:'dark');var pendingExt=s&&s!=='default'&&!skins[s]&&!legacy[s];var skin=skins[s]?s:(pendingExt?s:(m?m[1]:'default'));var _hadAppearance=localStorage.getItem('hermes-theme')!==null||localStorage.getItem('hermes-skin')!==null;if(_hadAppearance){localStorage.setItem('hermes-theme',theme);localStorage.setItem('hermes-skin',skin);}if(theme==='system')theme=window.matchMedia('(prefers-color-scheme:dark)').matches?'dark':'light';if(theme==='dark')document.documentElement.classList.add('dark');if(skin!=='default')document.documentElement.dataset.skin=skin;}catch(e){document.documentElement.classList.add('dark');}})()</script>
<link rel="stylesheet" href="/static/style.css">
<style>
html,body{height:auto;min-height:100%}
+14 -6
View File
@@ -26,6 +26,7 @@
--border-subtle:rgba(0,0,0,.08);--border-muted:rgba(0,0,0,.12);
font-family:var(--font-ui);font-size:14px;line-height:1.6;
}
button,input,select,textarea{font-family:var(--font-ui);}
/* ── Font size modifiers ── */
/* ── Font size preference: scale key UI text elements ── */
/* Default is 14px (no attribute needed). Small=12px, Large=16px, Extra Large=18px. */
@@ -2233,6 +2234,7 @@
.notes-preview-card .memory-content{max-height:420px;overflow:auto;}
.field-label{font-size:10px;font-weight:700;text-transform:uppercase;letter-spacing:.08em;color:var(--muted);margin-bottom:5px;opacity:.8;}
select{width:100%;background:var(--input-bg);border:1px solid var(--border2);border-radius:8px;color:var(--text);padding:8px 28px 8px 10px;font-size:12px;outline:none;appearance:none;margin-bottom:6px;cursor:pointer;background-image:url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='10' height='6' viewBox='0 0 10 6'%3E%3Cpath d='M1 1l4 4 4-4' stroke='%238888aa' stroke-width='1.5' fill='none' stroke-linecap='round'/%3E%3C/svg%3E");background-repeat:no-repeat;background-position:right 10px center;}
.insights-period-select{width:auto;border-radius:6px;padding:6px 28px 6px 10px;font-size:12px;margin-bottom:0;}
select:focus{border-color:var(--accent);box-shadow:0 0 0 2px var(--accent-bg);}
optgroup{color:var(--muted);font-size:11px;font-weight:700;}
option{background:var(--bg);color:var(--text);padding:6px;}
@@ -2415,7 +2417,7 @@
.msg-body th{background:rgba(255,255,255,.07);padding:6px 10px;text-align:left;font-weight:600;border:1px solid var(--border2);}
.msg-body td{padding:5px 10px;border:1px solid rgba(255,255,255,.06);}
.msg-body tr:nth-child(even){background:rgba(255,255,255,.03);}
.markdown-table-filter{display:block;width:min(260px,100%);margin:8px 0 4px;padding:5px 8px;border:1px solid var(--border2);border-radius:6px;background:var(--input-bg);color:var(--text);font:inherit;font-size:12px;}
.markdown-table-filter{display:block;width:min(260px,100%);margin:8px 0 4px;padding:5px 8px;border:1px solid var(--border2);border-radius:6px;background:var(--input-bg);color:var(--text);font:inherit;font-family:var(--font-ui);font-size:12px;}
.markdown-table-filter:focus{outline:1px solid var(--accent);outline-offset:1px;}
.markdown-table-sort{display:flex;align-items:center;justify-content:space-between;gap:8px;width:100%;min-height:20px;padding:0;border:0;background:transparent;color:inherit;font:inherit;font-weight:inherit;text-align:left;cursor:pointer;}
.markdown-table-sort:focus-visible{outline:1px solid var(--accent);outline-offset:2px;}
@@ -2592,7 +2594,8 @@
@media (hover: hover) {
.attach-thumb:hover{filter:brightness(1.05);transform:scale(1.04);}
}
textarea#msg{width:100%;background:transparent;border:none;outline:none;color:var(--text);font-size:16px;line-height:1.65;padding:12px 16px 6px;resize:none;min-height:44px;max-height:200px;font-family:inherit;}
textarea#msg{width:100%;background:transparent;border:none;outline:none;color:var(--text);font-size:16px;line-height:1.65;padding:12px 16px 6px;resize:none;min-height:44px;max-height:200px;overflow-y:auto;field-sizing:content;font-family:inherit;}
textarea#msg:placeholder-shown{field-sizing:fixed;}
textarea#msg::placeholder{color:var(--muted);}
.composer-footer{display:flex;align-items:center;justify-content:space-between;gap:10px;padding:6px 10px 10px;position:relative;container-type:inline-size;container-name:composer-footer;}
.composer-left{display:flex;align-items:center;gap:4px;min-width:0;flex:1;overflow-x:auto;overflow-y:hidden;scrollbar-width:none;}
@@ -5490,8 +5493,9 @@ main.main > #mainPlugin{display:none;}
.extension-installed-title{font-size:13px;font-weight:600;color:var(--text);overflow-wrap:anywhere;min-width:0;}
.extension-installed-meta{display:flex;align-items:center;gap:8px;flex-wrap:wrap;font-size:11px;color:var(--muted);}
.extension-installed-meta code{font-family:var(--font-mono);font-size:11px;color:var(--text);overflow-wrap:anywhere;word-break:break-word;}
.extension-toggle-btn{white-space:nowrap;flex:0 0 auto;}
.extension-toggle-btn:disabled{opacity:.55;cursor:not-allowed;}
.extension-installed-actions{display:flex;align-items:center;justify-content:flex-end;gap:6px;flex:0 0 auto;flex-wrap:wrap;}
.extension-toggle-btn,.extension-configure-btn{white-space:nowrap;flex:0 0 auto;}
.extension-toggle-btn:disabled,.extension-configure-btn:disabled{opacity:.55;cursor:not-allowed;}
.extension-settings-box{margin-top:8px;padding:9px 10px;border:1px solid var(--border);border-radius:8px;background:var(--card-bg,var(--sidebar-bg));}
.extension-settings-head{display:flex;align-items:flex-start;justify-content:space-between;gap:10px;margin-bottom:8px;}
.extension-settings-title{font-size:12px;font-weight:700;color:var(--text);}
@@ -5564,7 +5568,8 @@ main.main > #mainPlugin{display:none;}
.extensions-copy-diagnostics-btn{width:100%;}
.extension-summary-grid{grid-template-columns:1fr;}
.extension-installed-row{align-items:flex-start;flex-direction:column;}
.extension-toggle-btn{width:100%;}
.extension-installed-actions{width:100%;}
.extension-installed-actions .extension-toggle-btn,.extension-installed-actions .extension-configure-btn{width:auto;flex:1 1 140px;}
.extension-sidecar-row-head{align-items:flex-start;flex-direction:column;}
.extension-sidecar-fields>div{flex-direction:column;gap:2px;}
.extension-sidecar-fields span{flex:0 0 auto;}
@@ -6848,7 +6853,7 @@ main.main.showing-insights > #mainInsights{display:flex;overflow-y:auto;}
feels native to the WebUI rather than a one-off bridge UI. */
.kanban-modal-overlay{
position:fixed;inset:0;background:rgba(7,12,19,.62);backdrop-filter:blur(6px);
display:flex;align-items:center;justify-content:center;
display:flex;align-items:center;align-items:safe center;justify-content:center;overflow-y:auto;
z-index:1100;padding:24px;
}
.kanban-modal-overlay[hidden]{display:none;}
@@ -6861,6 +6866,9 @@ main.main.showing-insights > #mainInsights{display:flex;overflow-y:auto;}
padding:18px 18px 16px;
color:var(--text);
box-sizing:border-box;
max-height:calc(100vh - 48px);
max-height:calc(100dvh - 48px);
overflow-y:auto;
}
:root:not(.dark) .kanban-modal{
background:linear-gradient(180deg,#fff,#f5f0e8);
+765 -154
View File
File diff suppressed because it is too large Load Diff
+16 -3
View File
@@ -609,6 +609,13 @@ function renderSessionArtifacts(){
}).join('');
}
function projectSessionArtifactsForOwner(sessionId){
if(!sessionId||!S.session||S.session.session_id!==sessionId) return false;
if(typeof _isSessionCurrentPane!=='function'||!_isSessionCurrentPane(sessionId)) return false;
renderSessionArtifacts();
return true;
}
async function _workspacePathExists(path){
if(!S.session||!path) return false;
const parts=String(path).replace(/\\/g,'/').split('/').filter(Boolean);
@@ -1418,16 +1425,22 @@ async function _collectFilesFromEntry(entry, relPrefix) {
async function _collectOsDropUploads(dataTransfer) {
const out = [];
const items = dataTransfer.items ? [...dataTransfer.items] : [];
if (items.length && typeof items[0].webkitGetAsEntry === 'function') {
const files = dataTransfer.files ? [...dataTransfer.files] : [];
if (items.length) {
const entries = [];
for (const item of items) {
if (item.kind !== 'file') continue;
const entry = item.webkitGetAsEntry();
const getAsEntry = item.getAsEntry || item.webkitGetAsEntry;
const entry = typeof getAsEntry === 'function' ? getAsEntry.call(item) : null;
if (!entry) continue;
entries.push(entry);
}
for (const entry of entries) {
out.push(...await _collectFilesFromEntry(entry, ''));
}
if (out.length) return out;
}
for (const file of dataTransfer.files) {
for (const file of files) {
out.push({ file, relDir: '' });
}
return out;
+33 -1
View File
@@ -33,13 +33,45 @@ def auxiliary_client_modules():
The stubs are fresh objects owned by this helper, so the ``auxiliary_client``
attribute is set on our own module and no pre-existing ``agent`` module is
mutated. Restoration is therefore entirely ``patch.dict``'s job.
``agent.portal_tags`` is stubbed with a real ``ContextVar`` behind the same
set/reset/get surface as the Agent runtime (tokens are truthy objects, and
reset restores the previous value), so title-generation tests can assert
conversation-context publication/restoration without a real Agent runtime
(#7470).
"""
import contextvars as _contextvars
agent_stub = types.ModuleType("agent")
auxiliary_client_stub = types.ModuleType("agent.auxiliary_client")
portal_tags_stub = types.ModuleType("agent.portal_tags")
agent_stub.auxiliary_client = auxiliary_client_stub
portal_state = {"set_calls": [], "reset_calls": []}
_conversation_cv = _contextvars.ContextVar("webui_test_conversation_id", default=None)
def _set_conversation_context(conversation_id):
portal_state["set_calls"].append(conversation_id)
return _conversation_cv.set(conversation_id or None)
def _reset_conversation_context(token):
portal_state["reset_calls"].append(token)
_conversation_cv.reset(token)
def _get_conversation_context():
return _conversation_cv.get()
portal_tags_stub.set_conversation_context = _set_conversation_context
portal_tags_stub.reset_conversation_context = _reset_conversation_context
portal_tags_stub.get_conversation_context = _get_conversation_context
portal_tags_stub._portal_state = portal_state
agent_stub.portal_tags = portal_tags_stub
return mock.patch.dict(
sys.modules,
{"agent": agent_stub, "agent.auxiliary_client": auxiliary_client_stub},
{
"agent": agent_stub,
"agent.auxiliary_client": auxiliary_client_stub,
"agent.portal_tags": portal_tags_stub,
},
)
+22
View File
@@ -0,0 +1,22 @@
import json
from pathlib import Path
_FIXTURE_PATH = Path(__file__).parent / "fixtures" / "issue6611_regeneration_rows.json"
_EXPECTED_ROWS = [
{"role": "user", "content": "same prompt"},
{"role": "assistant", "content": "provider failed", "_error": True},
]
def load_issue6611_fixture():
fixture = json.loads(_FIXTURE_PATH.read_text(encoding="utf-8"))
assert set(fixture) == {"issue", "rows"}
assert fixture["issue"] == 6611
assert isinstance(fixture["rows"], list)
assert fixture["rows"] == _EXPECTED_ROWS
assert all(set(row) == {"role", "content"} for row in fixture["rows"][:1])
assert set(fixture["rows"][1]) == {"role", "content", "_error"}
assert [row["role"] for row in fixture["rows"]] == ["user", "assistant"]
assert fixture["rows"][1]["_error"] is True
return fixture
+2
View File
@@ -5,6 +5,8 @@ This test boots the real WebUI server with isolated state, drives the real chat
composer in Chromium, and supplies deterministic runtime events through the
existing Hermes Gateway Runs API. It proves that one assistant turn keeps the
same semantic activity across live streaming, settlement, and a hard reload.
Proposed in #6247; first implementation slice merged as #6251.
"""
from __future__ import annotations
+92
View File
@@ -231,6 +231,44 @@ def _reset_password_hash_cache():
_invalidate_password_hash_cache()
def _strip_leaked_webui_password_env() -> None:
"""Remove a leaked HERMES_WEBUI_PASSWORD between tests (#7168 review).
bootstrap.py runs _load_repo_dotenv() at import time, which copies values
from the developer's real repo .env straight into os.environ. When any
test imports bootstrap mid-session (e.g. tests/test_bootstrap_foreground.py
via its import_bootstrap fixture), a local .env containing
HERMES_WEBUI_PASSWORD leaks into the process environment OUTSIDE
monkeypatch's undo scope. Every later test then sees is_auth_enabled()
True and no-handler cookie helpers raise spurious
"build_profile_cookie requires a request handler" errors — exactly the
#5588 failure shape, but sourced from the repo .env instead of the hash
cache. Tests that legitimately enable auth set the var themselves AFTER
this strip; an intentionally-empty value ("") is preserved so
ctl.sh-style override semantics keep working.
HERMES_COMMAND gets the same treatment (#7168 re-gate round 7): a local
.env carrying HERMES_COMMAND leaks past bootstrap imports and redirects
gateway_restart._resolve_hermes_command() away from its mocked
shutil.which result, failing every later active-profile-restart test
with a machine-specific CLI path. Upstream code has no
HERMES_COMMAND override, so stripping a leaked value restores exact
upstream semantics.
"""
if os.environ.get("HERMES_WEBUI_PASSWORD") == "":
pass # intentional empty override preserved for the password var
else:
os.environ.pop("HERMES_WEBUI_PASSWORD", None)
os.environ.pop("HERMES_COMMAND", None)
@pytest.fixture(autouse=True)
def _strip_leaked_webui_password():
_strip_leaked_webui_password_env()
yield
_strip_leaked_webui_password_env()
@pytest.fixture(autouse=True)
def _invalidate_providers_cache():
"""Clear the /api/providers TTL cache around every test (#6010).
@@ -651,6 +689,7 @@ def pytest_collection_modifyitems(config, items):
'test_delivery_options_structure',
'test_delivery_options_includes_common_platforms',
'test_delivery_options_local_label',
'test_delivery_options_survives_the_authority_module_move',
# Skills endpoints (need tools.skills_tool module)
'test_skills_list',
'test_skills_list_has_required_fields',
@@ -1192,6 +1231,24 @@ _AGENT_PATH_ENV_KEYS = ("HERMES_WEBUI_AGENT_DIR", "PYTHONPATH", "HERMES_WEBUI_PY
_REAL_AGENT_ENV = {k: os.environ.get(k) for k in _AGENT_PATH_ENV_KEYS}
_REAL_SYS_PATH = list(sys.path)
# Keep the Windows restart seams inert after the suite isolation snapshots.
from api import updates as _updates
_real_windows_restart_spawn = _updates._windows_restart_spawn
_real_windows_restart_exit = _updates._windows_restart_exit
def _pytest_session_safe_windows_restart_spawn(_args, **_kwargs): # pragma: no cover
return None
def _pytest_session_safe_windows_restart_exit(_code): # pragma: no cover
return None
_updates._windows_restart_spawn = _pytest_session_safe_windows_restart_spawn
_updates._windows_restart_exit = _pytest_session_safe_windows_restart_exit
def _hermes_cli_is_healthy() -> bool:
mod = sys.modules.get("hermes_cli")
@@ -1246,6 +1303,41 @@ def _restore_hermes_cli_module():
sys.path[:] = _REAL_SYS_PATH
# ── hermes-agent API-drift compat: tools.approval._ApprovalEntry ─────────────
# The core agent (commit 14791b4d4e, released in v2026.9.7) split tools/approval.py
# into smart/human-wait/gateway-wait modules and dropped 43 back-compat facade
# re-exports — including the module-level ``tools.approval._ApprovalEntry`` alias.
# The class itself still exists, unchanged, at ``tools.approval_gateway_wait._ApprovalEntry``
# (identical ``__init__(self, data)`` contract: threading.Event + data dict + an
# auto-stamped request_id). Several webui gateway-approval tests reference
# ``_ApprovalEntry`` through ``tools.approval`` (its pre-split public location):
# on a box with the NEW agent installed the attribute is gone (AttributeError via
# the module's PEP-562 ``__getattr__``); on an OLD agent it is still present. No
# PRODUCTION webui code imports the dropped facade (verified: api/route_approvals.py
# and api/streaming.py import only surviving names), so this is purely test-side
# coupling to a moved internal symbol.
#
# This runs as an AUTOUSE fixture (not at conftest import time) because the agent
# dir is only added to sys.path by the earlier autouse guards / server fixture —
# ``tools.approval`` is not importable at conftest module load. The backfill is
# idempotent and only aliases the class when the installed agent lacks it, using
# the exact class object the agent uses internally. It keeps the tests working
# across both the pre-split and post-split agent without touching ~22 call sites,
# and is a no-op when tools.approval isn't importable (agent absent → tests skip).
@pytest.fixture(autouse=True)
def _backfill_approval_entry_facade():
try:
import tools.approval as _approval_mod
except Exception:
return # agent not importable here → agent-dependent tests skip anyway
if getattr(_approval_mod, "_ApprovalEntry", None) is None:
try:
from tools.approval_gateway_wait import _ApprovalEntry as _entry_cls
except Exception:
return # unexpected layout → let tests skip/fail loudly, don't fake it
_approval_mod._ApprovalEntry = _entry_cls
# ── Per-test session cleanup ──────────────────────────────────────────────────
@pytest.fixture(autouse=True)
+14
View File
@@ -0,0 +1,14 @@
{
"issue": 6611,
"rows": [
{
"role": "user",
"content": "same prompt"
},
{
"role": "assistant",
"content": "provider failed",
"_error": true
}
]
}
+13
View File
@@ -0,0 +1,13 @@
{
"source": "https://github.com/nesquena/hermes-webui/issues/7195#issuecomment-5381720224",
"issue": "https://github.com/nesquena/hermes-webui/issues/7195",
"base": "3b9c632a1fa339abcfd457973dcf10810640e760",
"command": "python -m pytest tests/test_updates.py tests/test_update_checker.py tests/test_update_channels.py -q --tb=short -p no:cacheprovider",
"launch": {
"platform": "win32",
"frozen": false,
"argv_shape": ["C:\\Users\\david\\AppData\\Local\\hermes\\hermes-agent\\venv\\Lib\\site-packages\\pytest\\__main__.py", "tests/test_updates.py", "tests/test_update_checker.py", "tests/test_update_channels.py", "-q", "--tb=short", "-p", "no:cacheprovider"],
"restart_payload_before_fix": ["C:\\Users\\david\\AppData\\Local\\hermes\\hermes-agent\\venv\\Scripts\\python.exe", "C:\\Users\\david\\AppData\\Local\\hermes\\hermes-agent\\venv\\Lib\\site-packages\\pytest\\__main__.py", "tests/test_updates.py", "tests/test_update_checker.py", "tests/test_update_channels.py", "-q", "--tb=short", "-p", "no:cacheprovider"],
"observed_symptom": "Each pytest generation starts another pytest generation and its server fixtures, producing recursive children and leaked processes."
}
}
+3 -1
View File
@@ -606,7 +606,9 @@ def test_session_compact_includes_parent():
# the top of compact() which pushed the parent_session_id field beyond a
# 1500-char window — widen the scan to 3000 chars to cover the full
# return-dict body without re-tightening every time compact() grows.
compact_def_match = re.search(r"def compact\(self", src)
# compact() accepts optional projection flags on separate lines; match the
# method name and first parameter without pinning formatting.
compact_def_match = re.search(r"def compact\(\s*self", src)
assert compact_def_match, "Could not find compact() method"
snippet = src[compact_def_match.start():compact_def_match.start() + 3000]
assert "'parent_session_id'" in snippet, \
+6 -1
View File
@@ -344,7 +344,12 @@ console.log(JSON.stringify({
prefill_pos = BOOT_JS.find(
"const prefillIntent=(typeof _composerPrefillIntentFromLocation==='function')?_composerPrefillIntentFromLocation():null;"
)
first_await_pos = BOOT_JS.find("const s=await api('/api/settings');", prefill_pos)
first_await_match = re.search(
r"const s=await api\('/api/settings'(?:,\{[^)]*\})?\);",
BOOT_JS[prefill_pos:],
)
assert first_await_match, "settings fetch with optional options not found after prefill intent"
first_await_pos = prefill_pos + first_await_match.start()
active_profile_pos = BOOT_JS.find(
"const activeProfileState = await _resolveActiveProfileBootstrapState();",
first_await_pos,
+7 -2
View File
@@ -19,6 +19,7 @@ flipped the default to 'steer'; this file was updated for the rename while
preserving the persistence-mirror guarantees (the load-failure path must still
honor the saved preference, not clobber it with a hardcoded default).
"""
import re
from pathlib import Path
ROOT = Path(__file__).parent.parent
@@ -47,8 +48,12 @@ class TestEagerDefault:
eager_idx = BOOT_JS.find("window._defaultMessageMode=_readPersistedDefaultMessageMode()")
assert eager_idx >= 0, "eager default assignment not found"
# The async IIFE awaits /api/settings; the success-path assignment lives inside it.
await_idx = BOOT_JS.find("const s=await api('/api/settings')")
assert await_idx >= 0, "async settings fetch not found"
await_match = re.search(
r"const s=await api\('/api/settings'(?:,\{[^)]*\})?\)",
BOOT_JS,
)
assert await_match, "async settings fetch with optional options not found"
await_idx = await_match.start()
assert eager_idx < await_idx, (
"the eager window._defaultMessageMode default must be set BEFORE the async "
"/api/settings fetch — otherwise the boot-window race (#5167) persists"
@@ -0,0 +1,188 @@
"""Regression checks for #6808: the boot script must not fabricate a theme choice.
`syncSettings` in static/boot.js deliberately lets the server win on a first
visit, because (its own comment) "empty (new-browser) state is indistinguishable
from a user who chose the defaults". The inline bootstrap in static/index.html
used to defeat that: it resolved a theme with
`localStorage.getItem('hermes-theme')||'dark'` and then wrote the result back
unconditionally, so a brand-new browser stored `hermes-theme=dark` before any
request was made. syncSettings then saw an "explicit" value, ignored the
server's SETTINGS_DEFAULTS, and POSTed the fabricated value back — making the
server-side appearance default unreachable for deployments that change it.
These cases pin the distinction the fix introduces, which nothing covered
before: ABSENT appearance state is not the same as EXPLICIT appearance state.
- both keys absent -> no setItem at all (and the dark pre-paint is UNCHANGED)
- either key present -> both values are normalised and persisted
- legacy names -> migration still happens, exactly as before
The third is the reason the guard is on the PAIR rather than per key. The write
is not pointless: it canonicalises legacy names (`solarized` -> dark+poseidon).
With `hermes-theme=solarized` and no `hermes-skin`, a per-key guard would never
write the derived skin and the mapping would be lost on the next load.
"""
import json
import re
import shutil
import subprocess
import tempfile
from pathlib import Path
import pytest
ROOT = Path(__file__).resolve().parent.parent
INDEX_HTML = ROOT / "static" / "index.html"
NODE = shutil.which("node")
pytestmark = pytest.mark.skipif(
NODE is None,
reason="node is required to execute the appearance bootstrap harness",
)
def _bootstrap_script() -> str:
"""The inline appearance bootstrap, lifted out of index.html."""
html = INDEX_HTML.read_text(encoding="utf-8")
for m in re.finditer(r"<script>(.*?)</script>", html, re.S):
body = m.group(1)
if "hermes-theme" in body and "legacy" in body and "skins" in body:
return body
raise AssertionError("appearance bootstrap <script> not found in index.html")
_DRIVER = r"""
const fs = require('fs');
const script = fs.readFileSync(process.argv[2], 'utf8');
const store = JSON.parse(process.argv[3] || '{}');
const writes = [];
globalThis.localStorage = {
getItem(k) { return Object.prototype.hasOwnProperty.call(store, k) ? store[k] : null; },
setItem(k, v) { writes.push([k, String(v)]); store[k] = String(v); },
removeItem(k) { delete store[k]; },
};
const classes = new Set();
globalThis.document = {
documentElement: {
classList: { add: (c) => classes.add(c), remove: (c) => classes.delete(c) },
dataset: {},
},
querySelectorAll: () => [],
};
globalThis.window = globalThis;
globalThis.matchMedia = () => ({ matches: false });
globalThis.window.matchMedia = globalThis.matchMedia;
(0, eval)(script);
process.stdout.write(JSON.stringify({
writes,
store,
classes: [...classes],
skin: globalThis.document.documentElement.dataset.skin || null,
}));
"""
def _run(initial_store: dict) -> dict:
"""Execute the real bootstrap in node against a stubbed localStorage.
Driver and script go to FILES and are invoked as `node <driver> <script>`.
Passing them via `node -e ... -- <script>` silently shifts process.argv, so
the driver eval'd the JSON scenario instead of the bootstrap — and because
`eval('{}')` is a valid empty block rather than a syntax error, the
empty-store case passed while executing nothing at all.
"""
with tempfile.TemporaryDirectory() as td:
driver = Path(td) / "driver.js"
script = Path(td) / "bootstrap.js"
driver.write_text(_DRIVER, encoding="utf-8")
script.write_text(_bootstrap_script(), encoding="utf-8")
proc = subprocess.run(
[NODE, str(driver), str(script), json.dumps(initial_store)],
capture_output=True, text=True, timeout=60, check=False,
)
assert proc.returncode == 0, f"harness failed: {proc.stderr[:400]}"
out = json.loads(proc.stdout)
# The harness must have actually run the bootstrap: it always paints.
assert out["classes"] or out["skin"] or out["writes"], (
"harness produced no effects at all — it probably executed the wrong "
"source; check the node argv indices"
)
return out
def test_fresh_browser_writes_nothing():
"""The bug: an empty store must not become an 'explicit' user choice."""
out = _run({})
assert out["writes"] == [], (
"a fresh visit must not write appearance keys — those writes are what "
f"made syncSettings treat the fallback as explicit; got {out['writes']}"
)
assert out["store"] == {}
def test_fresh_browser_still_paints_dark():
"""The fix changes persistence only, never the pre-paint result."""
out = _run({})
assert "dark" in out["classes"], (
"the dark pre-paint fallback must be unchanged for an empty store"
)
assert out["skin"] is None
@pytest.mark.parametrize(
"initial,expected_theme,expected_skin",
[
({"hermes-theme": "dark"}, "dark", "default"),
({"hermes-skin": "mono"}, "dark", "mono"),
({"hermes-theme": "light", "hermes-skin": "mono"}, "light", "mono"),
],
)
def test_any_prior_state_still_normalises_and_persists(initial, expected_theme, expected_skin):
"""Either key present is enough to make this an explicit choice."""
out = _run(dict(initial))
assert out["store"]["hermes-theme"] == expected_theme
assert out["store"]["hermes-skin"] == expected_skin
assert [k for k, _ in out["writes"]] == ["hermes-theme", "hermes-skin"]
@pytest.mark.parametrize(
"legacy,theme,skin",
[
("solarized", "dark", "poseidon"),
("slate", "dark", "slate"),
("monokai", "dark", "sisyphus"),
("nord", "dark", "slate"),
("oled", "dark", "default"),
],
)
def test_legacy_theme_migration_survives(legacy, theme, skin):
"""A per-key guard would drop the derived skin here. The pair guard must not.
Only `hermes-theme` is stored, and the skin is DERIVED from the legacy
mapping — so the write of `hermes-skin` is the only thing that persists the
migration.
"""
out = _run({"hermes-theme": legacy})
assert out["store"]["hermes-theme"] == theme
assert out["store"]["hermes-skin"] == skin, (
f"legacy {legacy!r} must persist its derived skin {skin!r}"
)
def test_guard_is_on_the_pair_not_per_key():
"""Source assertion, so the intent survives a future refactor of the block."""
script = _bootstrap_script()
assert "_hadAppearance" in script
assert re.search(
r"_hadAppearance\s*=\s*localStorage\.getItem\('hermes-theme'\)\s*!==\s*null\s*\|\|"
r"\s*localStorage\.getItem\('hermes-skin'\)\s*!==\s*null",
script,
), "the guard must be true when EITHER appearance key is present"
assert re.search(r"if\(_hadAppearance\)\{[^}]*setItem\('hermes-theme'", script), (
"both appearance writes must sit behind the guard"
)
+406
View File
@@ -0,0 +1,406 @@
"""Regression tests for #7027 — UID/GID auto-detection probe order.
Background: ``docker_init.bash`` auto-detects the UID/GID to remap the
``hermeswebui`` user to, by stat-ing mounted directories. Before this fix the
probe list was:
priority 1: /home/hermeswebui/.hermes, $HERMES_HOME, /opt/data
priority 2: /workspace
fallback: 1024
In a stock single-container image none of the priority-1 candidates exist, but
``/workspace`` *does* — owned by the image's build-time ``1024:1024``. Detection
therefore returned the image's own owner, which carries no information about the
host, while the one directory that carries the host UID by definition — the
state-dir bind mount — was never probed. With a host-owned state mount and no
explicit ``WANTED_UID``, the container remapped to 1024, failed its own
writability check on the state dir, and restart-looped:
touch: cannot touch '/app/data/.testfile': Permission denied
!! ERROR: Failed to verify state directory at /app/data
The second half of the bug: ``1024`` was both the fallback sentinel and a
legitimate UID, so ``WANTED_UID=1024`` supplied explicitly by an operator was
overwritten by detection anyway.
The behavioural tests below extract the UID/GID resolution block from
``docker_init.bash`` and run it under real bash, with ``stat`` stubbed so
ownership can be simulated without root. Source-level assertions alone cannot
catch shell-quoting or ordering regressions inside the block (the pattern
mirrors ``tests/test_docker_env_readonly_vars.py``).
The startup-health half of this contract — a real container coming up on a
host-owned state mount — is covered by the ``state-dir-uid`` job in
``.github/workflows/docker-smoke.yml``.
"""
import shutil
import subprocess
import textwrap
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parents[1]
INIT_SH = (REPO_ROOT / "docker_init.bash").read_text(encoding="utf-8")
# Logical probe name -> path basename used inside the sandbox. The extracted
# block's absolute paths are rewritten to these so the test can create/own them.
PROBE_PATHS = {
"state": "state",
"app_data": "app-data",
"hermes_home": "hermes-home",
"opt_data": "opt-data",
"workspace": "workspace",
}
_BLOCK_START = "it=$itdir/hermeswebui_user_uid\n"
_BLOCK_END = 'echo "-- WANTED_GID: \\"${WANTED_GID}\\""'
# ── source-level invariants ───────────────────────────────────────────────────
def _probe_loop_lines():
"""Both `for _probe_dir in ...` lines (UID block, then GID block)."""
return [ln for ln in INIT_SH.splitlines() if "for _probe_dir in" in ln]
def test_7027_both_probe_loops_include_the_state_dir():
"""UID *and* GID detection must probe the configured state directory."""
loops = _probe_loop_lines()
assert len(loops) == 2, f"expected one probe loop for UID and one for GID, found {len(loops)}"
for line in loops:
assert "HERMES_WEBUI_STATE_DIR" in line, (
"probe loop must include the configured state dir — it is the only "
f"path that is a bind mount by definition (#7027): {line.strip()!r}"
)
def test_7027_state_dir_is_probed_first_in_both_loops():
"""The state dir must be the first candidate, ahead of hermes-home/workspace."""
for line in _probe_loop_lines():
candidates = line.split("for _probe_dir in", 1)[1].rstrip("; do").strip()
first = candidates.split()[0]
assert "HERMES_WEBUI_STATE_DIR" in first, (
"the state dir must be probed before any image-owned path, otherwise "
f"a stock /workspace wins with the build-time 1024 (#7027): {first!r}"
)
def test_7027_workspace_probe_still_exists_as_lower_priority():
"""/workspace stays as a fallback signal — it is not removed (#569, #668)."""
assert 'if [ -d "/workspace" ]' in INIT_SH, (
"/workspace must remain a lower-priority probe for setups that bind-mount it"
)
state_pos = INIT_SH.find("HERMES_WEBUI_STATE_DIR:-/app/data")
workspace_pos = INIT_SH.find('if [ -d "/workspace" ]')
assert state_pos != -1 and workspace_pos != -1
assert state_pos < workspace_pos, "state-dir probe must precede the /workspace probe"
def test_7027_explicit_value_guard_present_on_every_detection_branch():
"""Each detection branch must be gated on the value not being explicit.
Without this, `WANTED_UID=1024` set deliberately by an operator is
indistinguishable from the unset fallback and gets overwritten.
"""
uid_guards = INIT_SH.count('[ "$_wanted_uid_source" != "explicit" ]')
gid_guards = INIT_SH.count('[ "$_wanted_gid_source" != "explicit" ]')
assert uid_guards == 2, f"both UID detection branches must be guarded, found {uid_guards}"
assert gid_guards == 2, f"both GID detection branches must be guarded, found {gid_guards}"
def test_7027_id_source_is_persisted_for_the_runtime_pass():
"""The explicit/detected origin must be persisted next to the value.
`su` drops the environment when the script re-enters as the runtime user,
so without persistence the second pass cannot tell an operator's explicit
1024 from the fallback.
"""
assert "hermeswebui_user_uid_source" in INIT_SH
assert "hermeswebui_user_gid_source" in INIT_SH
assert INIT_SH.count('write_privtmpfile $it_source') == 2, (
"both the UID and GID source markers must be written to $itdir"
)
def test_7027_fallback_default_preserved():
"""The 1024 fallback must survive (guard shared with #569)."""
assert "WANTED_UID=${WANTED_UID:-1024}" in INIT_SH
assert "WANTED_GID=${WANTED_GID:-1024}" in INIT_SH
# ── behavioural harness ───────────────────────────────────────────────────────
@pytest.mark.skipif(shutil.which("bash") is None, reason="bash not available")
class TestIdResolutionBehaviour:
"""Run the real resolution block under bash against simulated mounts."""
@staticmethod
def _extract_block(init_sh: str) -> str:
start = init_sh.find(_BLOCK_START)
assert start != -1, "UID/GID resolution block not found in docker_init.bash"
end = init_sh.find(_BLOCK_END, start)
assert end != -1, "end of the UID/GID resolution block not found"
return init_sh[start:end + len(_BLOCK_END)]
def _run(self, tmp_path, *, owners, env=None, itdir=None, state_dir_env=True):
"""Execute the resolution block.
owners: {logical name: (uid, gid)} — each named directory is created and
its simulated ownership is returned by the stubbed `stat`.
Directories not listed are simply absent.
state_dir_env: when False, HERMES_WEBUI_STATE_DIR is left unset so the
block falls back to its built-in default path.
"""
sandbox = tmp_path / "sandbox"
sandbox.mkdir(exist_ok=True)
itdir = itdir or (tmp_path / "itdir")
Path(itdir).mkdir(exist_ok=True)
paths = {name: sandbox / base for name, base in PROBE_PATHS.items()}
for name in owners:
paths[name].mkdir(parents=True, exist_ok=True)
block = self._extract_block(INIT_SH)
# Rewrite the block's absolute probe paths into the sandbox. Longest
# first so /app/data is not clipped by a shorter prefix.
for absolute, name in (
("/home/hermeswebui/.hermes", "hermes_home"),
("/opt/data", "opt_data"),
("/app/data", "app_data"),
("/workspace", "workspace"),
):
block = block.replace(absolute, str(paths[name]))
# `stat -c FMT PATH` stub: $2 is the format, $3 the path. Unknown paths
# exit nonzero so the block's `|| echo ""` yields an empty detection.
cases = []
for name, (uid, gid) in owners.items():
cases.append(
f' {str(paths[name])!r}) '
f'if [ "$_fmt" = "%u" ]; then echo "{uid}"; else echo "{gid}"; fi; return 0 ;;'
)
stat_cases = "\n".join(cases) if cases else " __never__) return 1 ;;"
script = textwrap.dedent("""\
set -e
itdir={itdir}
write_privtmpfile() {{
tmpfile=$1
if [ -f "$tmpfile" ]; then rm -f "$tmpfile"; fi
printf '%s' "$2" > "$tmpfile"
chmod 600 "$tmpfile"
}}
error_exit() {{ echo "!! ERROR: $*"; exit 1; }}
stat() {{
_fmt=$2
_path=$3
case "$_path" in
{stat_cases}
esac
return 1
}}
{block}
echo "RESULT_UID=${{WANTED_UID}}"
echo "RESULT_GID=${{WANTED_GID}}"
echo "RESULT_UID_SOURCE=$(cat "$itdir/hermeswebui_user_uid_source" 2>/dev/null || echo missing)"
echo "RESULT_GID_SOURCE=$(cat "$itdir/hermeswebui_user_gid_source" 2>/dev/null || echo missing)"
""").format(
itdir=str(itdir),
stat_cases=stat_cases,
block=block,
)
run_env = {"PATH": "/usr/bin:/bin:/usr/sbin:/sbin"}
if state_dir_env:
run_env["HERMES_WEBUI_STATE_DIR"] = str(paths["state"])
run_env.update(env or {})
result = subprocess.run(
["bash", "-c", script],
capture_output=True,
text=True,
timeout=20,
env=run_env,
)
assert result.returncode == 0, (
f"resolution block failed: rc={result.returncode}\n"
f"stdout: {result.stdout}\nstderr: {result.stderr}"
)
parsed = dict(
line.split("=", 1)
for line in result.stdout.splitlines()
if line.startswith("RESULT_")
)
parsed["_stdout"] = result.stdout
parsed["_state_dir"] = str(paths["state"])
return parsed
# ── the reported bug ──────────────────────────────────────────────────────
def test_7027_state_mount_beats_stock_workspace(self, tmp_path):
"""The reproduction from the issue: host-owned state dir, stock /workspace.
Host directory owned by 1001:1001 bind-mounted as the state dir, no
explicit IDs, and an image-owned /workspace at 1024:1024. Detection must
return 1001 — the identity that makes the state dir writable.
"""
out = self._run(
tmp_path,
owners={"state": (1001, 1001), "workspace": (1024, 1024)},
)
assert out["RESULT_UID"] == "1001", (
"UID must be detected from the state bind mount, not from the "
f"image-owned /workspace (#7027). stdout:\n{out['_stdout']}"
)
assert out["RESULT_GID"] == "1001", (
f"GID must be detected from the state bind mount too. stdout:\n{out['_stdout']}"
)
assert f"(from {out['_state_dir']})" in out["_stdout"], (
"the log line must name the state dir as the detection source so the "
"operator can see which path decided the identity"
)
def test_7027_non_1024_state_mount_with_no_workspace_at_all(self, tmp_path):
"""Same, in the single-container shape where /workspace is not mounted."""
out = self._run(tmp_path, owners={"state": (501, 20)})
assert out["RESULT_UID"] == "501", "macOS-style host UID must be detected"
assert out["RESULT_GID"] == "20"
def test_7027_default_state_path_probed_when_env_unset(self, tmp_path):
"""With HERMES_WEBUI_STATE_DIR unset, the built-in default is still probed."""
out = self._run(
tmp_path,
owners={"app_data": (1001, 1001), "workspace": (1024, 1024)},
state_dir_env=False,
)
assert out["RESULT_UID"] == "1001"
assert out["RESULT_GID"] == "1001"
# ── explicit operator values ──────────────────────────────────────────────
def test_7027_explicit_1024_is_not_overwritten(self, tmp_path):
"""`WANTED_UID=1024` set deliberately must survive detection.
1024 is both the fallback default and a legitimate UID; before the fix
the two were indistinguishable and detection clobbered the explicit value.
"""
out = self._run(
tmp_path,
owners={"state": (1001, 1001), "workspace": (1001, 1001)},
env={"WANTED_UID": "1024", "WANTED_GID": "1024"},
)
assert out["RESULT_UID"] == "1024", (
"an explicitly supplied 1024 must be preserved, not treated as unset (#7027)"
)
assert out["RESULT_GID"] == "1024"
assert out["RESULT_UID_SOURCE"] == "explicit"
assert out["RESULT_GID_SOURCE"] == "explicit"
def test_7027_explicit_non_default_still_wins(self, tmp_path):
"""The pre-existing contract: any explicit value beats detection."""
out = self._run(
tmp_path,
owners={"state": (1001, 1001)},
env={"WANTED_UID": "1500", "WANTED_GID": "1500"},
)
assert out["RESULT_UID"] == "1500"
assert out["RESULT_GID"] == "1500"
def test_7027_explicit_choice_survives_the_privilege_drop(self, tmp_path):
"""Second pass (as the runtime user) must not re-detect over an explicit 1024.
docker_init.bash runs twice: once as root, then `exec su` re-enters it as
hermeswebui with the environment dropped. The second pass reads the value
back from $itdir — so the *origin* has to persist too, or the explicit
1024 gets auto-detected away and the UID no longer matches the running
user, which is a hard startup failure.
"""
itdir = tmp_path / "itdir-shared"
first = self._run(
tmp_path,
owners={"state": (1001, 1001)},
env={"WANTED_UID": "1024", "WANTED_GID": "1024"},
itdir=itdir,
)
assert first["RESULT_UID"] == "1024"
second = self._run( # no WANTED_* in env — `su` dropped it
tmp_path,
owners={"state": (1001, 1001)},
itdir=itdir,
)
assert second["RESULT_UID"] == "1024", (
"the runtime pass must honour the persisted explicit choice; "
"re-detecting here would make WANTED_UID disagree with the user the "
"script is already running as"
)
assert second["RESULT_GID"] == "1024"
def test_7027_persisted_default_is_still_re_detected(self, tmp_path):
"""A *detected*/fallback 1024 stays re-detectable on a later run.
This is what the `= 1024` sentinel was for: a container that fell back to
1024 with nothing mounted must pick up the right identity once a state
volume appears. Only explicit values are frozen.
"""
itdir = tmp_path / "itdir-shared"
first = self._run(tmp_path, owners={}, itdir=itdir)
assert first["RESULT_UID"] == "1024", "nothing mounted -> fallback"
assert first["RESULT_UID_SOURCE"] == "detected"
second = self._run(tmp_path, owners={"state": (1001, 1001)}, itdir=itdir)
assert second["RESULT_UID"] == "1001", (
"a persisted fallback 1024 must not freeze detection for later runs"
)
assert second["RESULT_GID"] == "1001"
# ── previously shipped behaviour that must not regress ────────────────────
def test_7027_workspace_fallback_preserved(self, tmp_path):
"""With no state mount, /workspace still decides (#569)."""
out = self._run(tmp_path, owners={"workspace": (501, 20)})
assert out["RESULT_UID"] == "501"
assert out["RESULT_GID"] == "20"
assert "from /workspace" in out["_stdout"] or "workspace UID" in out["_stdout"]
def test_7027_hermes_home_still_beats_workspace(self, tmp_path):
"""Two-container setups keep detecting from the shared hermes-home (#668)."""
out = self._run(
tmp_path,
owners={"hermes_home": (1001, 1001), "workspace": (1024, 1024)},
)
assert out["RESULT_UID"] == "1001"
assert out["RESULT_GID"] == "1001"
def test_7027_hermes_home_env_probe_preserved(self, tmp_path):
"""$HERMES_HOME remains a probe candidate (#668)."""
sandbox = tmp_path / "sandbox"
sandbox.mkdir(exist_ok=True)
hermes_home = sandbox / "hermes-home"
out = self._run(
tmp_path,
owners={"hermes_home": (1001, 1001)},
env={"HERMES_HOME": str(hermes_home)},
)
assert out["RESULT_UID"] == "1001"
def test_7027_root_owned_state_dir_is_skipped(self, tmp_path):
"""A root-owned state dir (fresh named volume) is not a host signal.
Docker creates a brand-new named volume owned by 0:0. Detection must skip
it — as it already does for every probe — and fall through.
"""
out = self._run(
tmp_path,
owners={"state": (0, 0), "workspace": (501, 20)},
)
assert out["RESULT_UID"] == "501", "root-owned probes must be ignored"
assert out["RESULT_GID"] == "20"
def test_7027_nothing_mounted_falls_back_to_1024(self, tmp_path):
"""No probe resolves -> the documented 1024 default."""
out = self._run(tmp_path, owners={})
assert out["RESULT_UID"] == "1024"
assert out["RESULT_GID"] == "1024"
+20 -15
View File
@@ -8,6 +8,11 @@ at 1, producing "1. 1. 1." instead of "1. 2. 3.".
Fix: emit value="N" on every <li> so the correct ordinal is preserved even when
items end up in separate <ol> containers after the paragraph split.
The list stage was rewritten for #6700 into a single-pass, mixed-marker parser
(a stack keyed by indentation + marker kind). These static assertions pin the
parts of that implementation that keep #886 working: the ordered-marker digit
capture and the value="N" emission on <li>.
"""
import os
import re
@@ -21,15 +26,15 @@ def get_ui_js():
class TestOrderedListNumbering:
def _ordered_list_block(self, src: str) -> str:
start = src.find("function _renderListBlock(lines, ordered){")
start = src.find("function _renderListBlock(lines){")
assert start != -1, "_renderListBlock helper not found in ui.js"
end = src.find("function _renderLists(src, ordered){", start)
end = src.find("function _renderLists(src){", start)
assert end != -1, "_renderLists helper not found after _renderListBlock"
return src[start:end]
def _ordered_list_dispatch_block(self, src: str) -> str:
start = src.find("s=_renderLists(s,true);")
assert start != -1, "ordered-list dispatch not found in ui.js"
start = src.find("s=_renderLists(s);")
assert start != -1, "list dispatch not found in ui.js"
return src[max(0, start - 260):start + 80]
def test_li_value_attr_present_in_ordered_list_block(self):
@@ -42,18 +47,18 @@ class TestOrderedListNumbering:
)
def test_li_value_uses_parsed_number(self):
"""The value= must be derived from parseInt of the captured digit, not hardcoded."""
"""The value= must be derived from parseInt of the captured digits, not hardcoded."""
src = get_ui_js()
ol_block = self._ordered_list_block(src)
assert 'parseInt' in ol_block, (
"Ordered-list block should use parseInt() to parse the list number (#886)"
)
def test_numMatch_variable_present(self):
"""The ordered-list branch must still capture digits from the markdown marker."""
def test_marker_digit_capture_present(self):
"""The ordered branch must still capture digits from the markdown marker."""
src = get_ui_js()
ol_block = self._ordered_list_block(src)
assert "const marker=ordered?'\\\\d+\\\\. ':'[-*+] ';" in ol_block, (
assert "\\d+\\." in ol_block or re.search(r"\(\\d\+\)\\\.", ol_block), (
"Ordered-list block should keep a digit marker pattern for numbered items (#886)"
)
@@ -68,18 +73,18 @@ class TestOrderedListNumbering:
)
def test_ordered_list_comment_references_issue(self):
"""A comment near the OL fix should reference the issue (#886) or the symptom."""
"""A comment near the list dispatch should reference #886 or the symptom."""
src = get_ui_js()
context = self._ordered_list_dispatch_block(src)
has_comment = '#886' in context or '1. 1. 1.' in context or 'blank lines' in context.lower()
assert has_comment, (
"Expected a comment near the OL fix explaining the blank-line issue (#886)"
"Expected a comment near the list dispatch explaining the blank-line issue (#886)"
)
def test_list_without_blank_lines_unaffected(self):
"""A compact list should still flow through the ordered-list helper."""
def test_lists_route_through_single_shared_helper(self):
"""All list rendering must flow through the single-pass shared helper."""
src = get_ui_js()
ol_block = self._ordered_list_dispatch_block(src)
assert "s=_renderLists(s,true);" in ol_block, (
"Ordered-list rendering should still route through the shared helper"
dispatch = self._ordered_list_dispatch_block(src)
assert "s=_renderLists(s);" in dispatch, (
"List rendering should route through the single-pass _renderLists helper"
)
@@ -175,3 +175,49 @@ def test_acp_session_appears_in_projection_as_cli(tmp_path):
assert row['is_cli_session'] is True
assert row['session_source'] == 'cli'
assert row['source_label'] == 'ACP'
def test_newer_kanban_rows_do_not_evict_cli_sessions_from_projection(tmp_path):
"""Kanban has a bounded window separate from interactive conversations."""
kanban_count = models.KANBAN_PROJECT_CHIP_LIMIT + 5
sessions = [
{
'id': 'older-cli-session',
'source': 'cli',
'title': 'CLI conversation',
},
*[
{
'id': f'newer-kanban-{i:03d}',
'source': 'kanban',
'title': f'Kanban worker {i:03d}',
}
for i in range(kanban_count)
],
]
results = _call_uncached(tmp_path, sessions)
result_ids = {row['session_id'] for row in results}
kanban_rows = [row for row in results if row.get('source_tag') == 'kanban']
assert 'older-cli-session' in result_ids
assert len(kanban_rows) == models.KANBAN_PROJECT_CHIP_LIMIT
assert {row['session_id'] for row in kanban_rows} == {
f'newer-kanban-{i:03d}'
for i in range(kanban_count - models.KANBAN_PROJECT_CHIP_LIMIT, kanban_count)
}
from api.routes import _dedupe_cli_sidebar_sessions_for_api
hidden = _dedupe_cli_sidebar_sessions_for_api(
results,
set(),
show_kanban_sessions=False,
)
visible = _dedupe_cli_sidebar_sessions_for_api(
results,
set(),
show_kanban_sessions=True,
)
assert not any(row.get('source_tag') == 'kanban' for row in hidden)
assert sum(row.get('source_tag') == 'kanban' for row in visible) == models.KANBAN_PROJECT_CHIP_LIMIT
+90
View File
@@ -0,0 +1,90 @@
"""Regression: Agent ``_row_id`` must reconcile with WebUI state provenance."""
def test_agent_and_webui_row_id_aliases_resolve_to_the_same_sqlite_row():
from api.models import _state_db_row_identity_details
agent_identity, agent_valid = _state_db_row_identity_details({"_row_id": 780107})
webui_identity, webui_valid = _state_db_row_identity_details(
{"_state_db_row_id": 780107}
)
assert agent_valid is True
assert webui_valid is True
assert agent_identity == webui_identity
assert agent_identity is not None
def test_agent_row_id_alias_dedupes_restamped_prompt_before_final_answer():
from api.models import merge_session_messages_append_only
sidecar_messages = [
{
"role": "assistant",
"content": "earlier activity",
"timestamp": 100.0,
"_state_db_row_id": 700,
"api_content": "earlier activity wire",
},
{
"role": "user",
"content": "original maintenance question",
"timestamp": 110.0,
"_row_id": 780107,
"api_content": "original maintenance question wire",
},
{
"role": "assistant",
"content": "settled final answer",
"timestamp": 130.0,
},
]
state_messages = [
{
"role": "user",
"content": "original maintenance question",
"timestamp": 120.0,
"_state_db_row_id": 780107,
"api_content": "original maintenance question wire",
}
]
merged = merge_session_messages_append_only(sidecar_messages, state_messages)
assert [message["content"] for message in merged] == [
"earlier activity",
"original maintenance question",
"settled final answer",
]
def test_imported_row_id_is_not_trusted_as_state_identity():
"""JSON import must not keep a caller-supplied ``_row_id`` as provenance.
Replay merge now treats ``_row_id`` as an Agent/WebUI identity alias. Import
must strip that field the same way it strips ``_state_db_row_id``, or a
crafted import can collide with a later genuine state.db turn.
"""
from api.helpers import strip_public_internal_fields
from api.models import _state_db_row_identity_details
imported = strip_public_internal_fields(
{
"messages": [
{
"role": "user",
"content": "same prompt",
"timestamp": 10.0,
"_row_id": 780107,
"_state_db_row_id": 780107,
}
]
}
)
imported_msg = imported["messages"][0]
assert isinstance(imported_msg, dict)
assert "_row_id" not in imported_msg
assert "_state_db_row_id" not in imported_msg
identity, valid = _state_db_row_identity_details(imported_msg)
assert valid is True
assert identity is None
+374 -3
View File
@@ -6,6 +6,7 @@ import os
from pathlib import Path
import subprocess
import sys
import time
import types
import pytest
@@ -25,6 +26,195 @@ def _git(repo: Path, *args: str) -> str:
return result.stdout.strip()
@pytest.fixture
def changed_agent_checkout(monkeypatch, tmp_path):
"""Exercise the real revision guard and scheduler without replacing a process."""
from api import agent_runtime, routes, updates
agent_root = tmp_path / "agent"
agent_root.mkdir()
module_path = agent_root / "run_agent.py"
module_path.write_text("class AIAgent: pass\n", encoding="utf-8")
_git(agent_root, "init", "-q")
_git(agent_root, "add", "run_agent.py")
_git(agent_root, "commit", "-qm", "loaded")
monkeypatch.setattr(agent_runtime, "_AGENT_SOURCE_DIR", agent_root)
monkeypatch.setattr(agent_runtime, "_AGENT_MODULE_PATH", module_path)
monkeypatch.setattr(agent_runtime, "_AGENT_REVISION", _git(agent_root, "rev-parse", "HEAD"))
module_path.write_text("class AIAgent: revision = 'changed'\n", encoding="utf-8")
_git(agent_root, "commit", "-qam", "changed")
hermes_home = tmp_path / "home"
hermes_home.mkdir()
venv_root = tmp_path / "separate-install"
(venv_root / "venv" / "bin").mkdir(parents=True)
monkeypatch.setattr(agent_runtime, "_HERMES_HOME", hermes_home)
monkeypatch.setattr(agent_runtime, "_AGENT_PYTHON", venv_root / "venv" / "bin" / "python")
workers = []
class CapturedThread:
def __init__(self, *, target, daemon):
self.target = target
def start(self):
workers.append(self.target)
monkeypatch.setattr(updates, "threading", types.SimpleNamespace(Thread=CapturedThread))
monkeypatch.setattr(routes, "get_config", lambda: {})
monkeypatch.setattr(routes, "webui_gateway_chat_enabled", lambda _cfg: False)
monkeypatch.setattr(
routes, "j",
lambda _handler, payload, status=200: {"status": status, "payload": payload},
)
def no_session_mutation(*_args, **_kwargs):
pytest.fail("stale runtime reached session materialization")
monkeypatch.setattr(routes, "_get_or_materialize_session", no_session_mutation)
return types.SimpleNamespace(
root=agent_root, home=hermes_home, venv_root=venv_root, workers=workers,
request=lambda: routes._handle_chat_start(object(), {"session_id": "stale-session"}),
)
@pytest.mark.parametrize(
("marker_state", "diagnostic"),
[
("absent", "unverified"),
("failed-exit", "unverified"),
("interrupted-exit", "unverified"),
("active", "active"),
("dead", "stale"),
("over-age", "stale"),
("malformed", "unknown"),
("invalid-encoding", "unknown"),
("future", "unknown"),
("nonfinite", "unknown"),
("oversized-pid", "unknown"),
("dangling", "unknown"),
("unreadable", "unknown"),
("update-incomplete", "incomplete"),
("lazy-refresh-incomplete", "incomplete"),
("separate-venv-incomplete", "incomplete"),
("symlink-venv-incomplete", "incomplete"),
],
)
def test_unverified_update_keeps_manual_409_without_restart(
changed_agent_checkout, marker_state, diagnostic,
):
"""Readable HEAD and lifecycle markers never authorize process replacement."""
checkout = changed_agent_checkout
marker = checkout.home / ".hermes-update-in-progress"
if marker_state in {"active", "failed-exit", "interrupted-exit"}:
marker.write_text(f"{os.getpid()}\n{time.time()}\n", encoding="utf-8")
if marker_state != "active":
marker.unlink() # Agent removes its marker on failed/interrupted exits.
elif marker_state == "dead":
process = subprocess.Popen([sys.executable, "-c", "pass"])
process.wait(timeout=10)
marker.write_text(f"{process.pid}\n{time.time()}\n", encoding="utf-8")
elif marker_state == "over-age":
marker.write_text(f"{os.getpid()}\n{time.time() - 86400}\n", encoding="utf-8")
elif marker_state == "malformed":
marker.write_text("not-a-pid\nnot-a-time\n", encoding="utf-8")
elif marker_state == "invalid-encoding":
marker.write_bytes(b"\xff\xfe")
elif marker_state == "oversized-pid":
marker.write_text(f"{2 ** 100}\n{time.time()}\n", encoding="utf-8")
elif marker_state in {"future", "nonfinite"}:
timestamp = time.time() + 86400 if marker_state == "future" else "nan"
marker.write_text(f"{os.getpid()}\n{timestamp}\n", encoding="utf-8")
elif marker_state == "dangling":
marker.symlink_to(checkout.home / "missing")
elif marker_state == "unreadable":
marker.mkdir()
elif marker_state in {"update-incomplete", "lazy-refresh-incomplete"}:
(checkout.root / f".{marker_state}").touch()
elif marker_state == "separate-venv-incomplete":
(checkout.venv_root / ".update-incomplete").touch()
elif marker_state == "symlink-venv-incomplete":
(checkout.venv_root / "venv/bin/python").symlink_to(sys.executable)
(checkout.venv_root / ".update-incomplete").touch()
response = checkout.request()
assert checkout.workers == [], "revision mismatch scheduled an automatic restart"
assert response["status"] == 409
payload = response["payload"]
assert payload["type"] == "agent_runtime_stale"
assert payload["retryable"] is True
assert payload["restart_scheduled"] is False
assert payload["agent_update_state"] == diagnostic
assert "Restart Hermes WebUI manually" in payload["error"]
assert "success" not in payload["error"].lower()
def test_final_read_cannot_authorize_restart_without_atomic_handoff(
monkeypatch, changed_agent_checkout,
):
"""An updater can acquire its marker after a read; no worker may be queued."""
from api import agent_runtime
checkout = changed_agent_checkout
read_revision = agent_runtime._read_agent_revision
marker = checkout.home / ".hermes-update-in-progress"
revision_reads = []
def read_then_start_update(*args, **kwargs):
revision = read_revision(*args, **kwargs)
revision_reads.append(revision)
marker.write_text(f"{os.getpid()}\n{time.time()}\n", encoding="utf-8")
return revision
monkeypatch.setattr(agent_runtime, "_read_agent_revision", read_then_start_update)
response = checkout.request()
assert checkout.workers == [], "an unprotected read authorized automatic restart"
assert revision_reads == [_git(checkout.root, "rev-parse", "HEAD")]
assert marker.is_file()
assert response["status"] == 409
assert response["payload"]["restart_scheduled"] is False
@pytest.mark.parametrize("response_kind", ["http-error", "raised-error"])
def test_async_compression_preserves_manual_restart_diagnostics(
monkeypatch, changed_agent_checkout, response_kind,
):
"""The real guard reaches the job's terminal error and repeated status reads."""
from api import agent_runtime, helpers, routes
checkout = changed_agent_checkout
sid = "stale-compression"
(checkout.root / ".update-incomplete").touch()
monkeypatch.setattr(routes, "get_session", lambda _sid: None)
monkeypatch.setattr(routes, "_MANUAL_COMPRESSION_JOBS", {
sid: {"session_id": sid, "status": "running", "updated_at": time.time()},
})
def compress(handler, _body):
try:
agent_runtime.ensure_agent_runtime_current()
except agent_runtime.AgentRuntimeChangedError as exc:
if response_kind == "raised-error":
raise
helpers.j(handler, agent_runtime.agent_runtime_stale_payload(exc), status=409)
monkeypatch.setattr(routes, "_handle_session_compress", compress)
routes._run_manual_compression_job(sid, {"session_id": sid})
for _ in range(2):
response = routes._handle_session_compress_status(object(), sid)
payload = response["payload"]
assert payload["status"] == "error"
assert payload["error_status"] == 409
assert payload["type"] == "agent_runtime_stale"
assert payload["retryable"] is True
assert payload["restart_scheduled"] is False
assert payload["agent_update_state"] == "incomplete"
assert "Restart Hermes WebUI manually" in payload["error"]
assert checkout.workers == []
def test_loaded_agent_runtime_fails_closed_after_source_revision_changes(tmp_path: Path):
agent_dir = tmp_path / "hermes-agent"
agent_dir.mkdir()
@@ -85,14 +275,14 @@ try:
except RuntimeError as exc:
message = str(exc)
assert "Hermes Agent was updated" in message
assert "Restart Hermes WebUI" in message
assert "Restart Hermes WebUI manually" in message
else:
raise AssertionError("stale in-process AIAgent was reused after its source revision changed")
try:
agent_runtime.require_ai_agent_class()
except agent_runtime.AgentRuntimeChangedError as exc:
assert "Restart Hermes WebUI" in str(exc)
assert "Restart Hermes WebUI manually" in str(exc)
else:
raise AssertionError("unguarded AIAgent import was allowed after its source revision changed")
""".strip()
@@ -121,6 +311,105 @@ else:
assert result.returncode == 0, result.stdout + result.stderr
def test_read_live_agent_update_rejects_fifo_marker_without_hanging(tmp_path: Path):
"""A FIFO in the marker path must classify ``unknown`` fast, not block.
Regression guard: the marker read once used ``Path.read_text()``, which
blocks forever on a FIFO (no writer) and would wedge the stale-runtime
request path. The hardened read opens O_NONBLOCK|O_NOFOLLOW and fstat-checks
for a small regular file, so a FIFO is rejected immediately.
"""
from api import agent_runtime
fifo = tmp_path / "fifo-marker"
os.mkfifo(fifo)
start = time.time()
result = agent_runtime._read_live_agent_update(fifo)
elapsed = time.time() - start
assert result == "unknown"
assert elapsed < 1.0, f"marker read blocked on FIFO for {elapsed:.2f}s"
def test_read_live_agent_update_rejects_oversized_marker(tmp_path: Path):
"""An oversized regular marker must classify ``unknown``, never OOM-read.
Regression guard against unbounded ``read_text()``: even though the file
starts with a valid PID/timestamp, its size exceeds the cap so it is
rejected rather than read whole.
"""
from api import agent_runtime
big = tmp_path / "big-marker"
big.write_bytes(
f"{os.getpid()}\n{time.time()}\n".encode("utf-8")
+ b"x" * (agent_runtime._AGENT_UPDATE_MARKER_MAX_BYTES + 1024)
)
assert agent_runtime._read_live_agent_update(big) == "unknown"
def test_read_live_agent_update_does_not_follow_symlink_marker(tmp_path: Path):
"""A symlinked marker must classify ``unknown`` (O_NOFOLLOW), not be read
through to its target."""
from api import agent_runtime
if not getattr(os, "O_NOFOLLOW", 0):
pytest.skip("O_NOFOLLOW unavailable on this platform")
target = tmp_path / "real-marker"
target.write_text(f"{os.getpid()}\n{time.time()}\n", encoding="utf-8")
link = tmp_path / "link-marker"
link.symlink_to(target)
assert agent_runtime._read_live_agent_update(link) == "unknown"
def test_read_live_agent_update_still_classifies_valid_markers(tmp_path: Path):
"""The hardening must not regress the happy path: a small regular marker
with a live PID still reads as ``active`` and a stale one as ``stale``."""
from api import agent_runtime
active = tmp_path / "active-marker"
active.write_text(f"{os.getpid()}\n{time.time()}\n", encoding="utf-8")
assert agent_runtime._read_live_agent_update(active) == "active"
stale = tmp_path / "stale-marker"
stale.write_text(
f"{os.getpid()}\n{time.time() - agent_runtime._AGENT_UPDATE_MAX_AGE_SECONDS - 60}\n",
encoding="utf-8",
)
assert agent_runtime._read_live_agent_update(stale) == "stale"
assert agent_runtime._read_live_agent_update(tmp_path / "absent") == "absent"
def test_read_live_agent_update_windows_fallback_never_opens_marker(
tmp_path: Path, monkeypatch
):
"""When the atomic open flags are unavailable (e.g. native Windows), the
read must not call os.open (which would raise on a missing O_NONBLOCK) and
must fail closed: missing -> absent, anything present -> unknown.
Regression guard for the portability CORE: os.O_NONBLOCK is Unix-only, so
the fast path must be gated on both flags being present.
"""
from api import agent_runtime
monkeypatch.setattr(agent_runtime, "_MARKER_SAFE_OPEN_AVAILABLE", False)
def _boom(*_a, **_k): # os.open must never be reached on the fallback path
raise AssertionError("os.open called despite unavailable atomic flags")
monkeypatch.setattr(agent_runtime.os, "open", _boom)
# A present, otherwise-valid marker is unverifiable without the safe open →
# classified unknown (never active/absent), and does not raise.
present = tmp_path / "present-marker"
present.write_text(f"{os.getpid()}\n{time.time()}\n", encoding="utf-8")
assert agent_runtime._read_live_agent_update(present) == "unknown"
# A genuinely missing marker is still absent.
assert agent_runtime._read_live_agent_update(tmp_path / "missing") == "absent"
def test_initial_non_git_source_preserves_supported_runtime(monkeypatch):
"""Non-Git installs cannot be compared, so they preserve existing behavior."""
from api import agent_runtime
@@ -247,6 +536,83 @@ def test_known_revision_becoming_unreadable_fails_closed(monkeypatch):
agent_runtime.ensure_agent_runtime_current()
def test_live_agent_update_marker_reports_active(monkeypatch, tmp_path):
"""A fresh marker owned by a live PID means the Agent update is active."""
from api import agent_runtime
hermes_home = tmp_path / "hermes-home"
agent_root = tmp_path / "hermes-agent"
hermes_home.mkdir()
agent_root.mkdir()
(hermes_home / ".hermes-update-in-progress").write_text(
f"{os.getpid()}\n{time.time()}\n",
encoding="utf-8",
)
monkeypatch.setattr(agent_runtime, "_HERMES_HOME", hermes_home)
monkeypatch.setattr(agent_runtime, "_AGENT_SOURCE_DIR", agent_root)
monkeypatch.setattr(agent_runtime, "_AGENT_PYTHON", None)
assert agent_runtime._agent_update_transaction_state() == "active"
def test_dead_or_over_age_agent_update_marker_is_stale(monkeypatch, tmp_path):
"""A dead or over-age owner does not prove update completion."""
from api import agent_runtime
hermes_home = tmp_path / "hermes-home"
agent_root = tmp_path / "hermes-agent"
hermes_home.mkdir()
agent_root.mkdir()
marker = hermes_home / ".hermes-update-in-progress"
monkeypatch.setattr(agent_runtime, "_HERMES_HOME", hermes_home)
monkeypatch.setattr(agent_runtime, "_AGENT_SOURCE_DIR", agent_root)
monkeypatch.setattr(agent_runtime, "_AGENT_PYTHON", None)
marker.write_text(f"99999999\n{time.time()}\n", encoding="utf-8")
monkeypatch.setattr(agent_runtime, "_pid_is_alive", lambda _pid: False)
assert agent_runtime._agent_update_transaction_state() == "stale"
marker.write_text(
f"{os.getpid()}\n{time.time() - agent_runtime._AGENT_UPDATE_MAX_AGE_SECONDS - 1}\n",
encoding="utf-8",
)
monkeypatch.setattr(agent_runtime, "_pid_is_alive", lambda _pid: True)
assert agent_runtime._agent_update_transaction_state() == "stale"
@pytest.mark.parametrize(
"marker_name",
[".update-incomplete", ".lazy-refresh-incomplete"],
)
def test_agent_recovery_markers_block_automatic_restart(
monkeypatch, tmp_path, marker_name
):
"""Agent recovery markers report an incomplete environment."""
from api import agent_runtime
hermes_home = tmp_path / "hermes-home"
agent_root = tmp_path / "hermes-agent"
hermes_home.mkdir()
agent_root.mkdir()
(agent_root / marker_name).write_text("incomplete\n", encoding="utf-8")
monkeypatch.setattr(agent_runtime, "_HERMES_HOME", hermes_home)
monkeypatch.setattr(agent_runtime, "_AGENT_SOURCE_DIR", agent_root)
monkeypatch.setattr(agent_runtime, "_AGENT_PYTHON", None)
assert agent_runtime._agent_update_transaction_state() == "incomplete"
def test_unreadable_agent_update_state_fails_closed(monkeypatch):
"""An unreadable marker is unknown, never proof that restart is safe."""
from api import agent_runtime
class UnreadableMarker:
def read_text(self, **_kwargs):
raise PermissionError("denied")
assert agent_runtime._read_live_agent_update(UnreadableMarker()) == "unknown"
def test_import_recapture_cannot_downgrade_known_revision(monkeypatch, tmp_path: Path):
"""A second unreadable revision read must not erase a known identity."""
from api import agent_runtime
@@ -352,6 +718,7 @@ def test_runner_flag_does_not_bypass_webui_owned_hidden_turns(monkeypatch):
"error": "restart required",
"type": "agent_runtime_stale",
"retryable": True,
"restart_scheduled": False,
}
@@ -385,6 +752,7 @@ def test_chat_start_rejects_stale_runtime_before_session_materialization(monkeyp
"error": "restart required",
"type": "agent_runtime_stale",
"retryable": True,
"restart_scheduled": False,
},
}
@@ -486,7 +854,7 @@ def test_stream_admission_uses_one_gateway_ownership_snapshot(monkeypatch, gatew
monkeypatch.setattr(routes, "_active_run_stream_for_session", lambda _sid: None)
monkeypatch.setattr(routes, "_is_hidden_empty_session", lambda _session: False)
monkeypatch.setattr(routes, "_prepare_chat_start_session_for_stream", prepare)
monkeypatch.setattr(routes, "set_last_workspace", lambda _workspace: None)
monkeypatch.setattr(routes, "set_last_workspace", lambda _workspace, **_kw: None)
monkeypatch.setattr(routes.threading, "Thread", FakeThread)
monkeypatch.setattr(turn_journal, "append_turn_journal_event", lambda *_args, **_kwargs: {})
@@ -583,6 +951,7 @@ def test_git_commit_message_stale_runtime_returns_typed_409(
"error": "restart required",
"type": "agent_runtime_stale",
"retryable": True,
"restart_scheduled": False,
},
}
@@ -688,6 +1057,7 @@ def test_server_side_turn_rejects_stale_runtime_before_session_acceptance(monkey
"error": "restart required",
"type": "agent_runtime_stale",
"retryable": True,
"restart_scheduled": False,
},
)
monkeypatch.setattr(
@@ -704,5 +1074,6 @@ def test_server_side_turn_rejects_stale_runtime_before_session_acceptance(monkey
"error": "restart required",
"type": "agent_runtime_stale",
"retryable": True,
"restart_scheduled": False,
"_status": 409,
}
@@ -0,0 +1,210 @@
"""agent v0.21 transport-API port tests for out-of-band LLM completions.
hermes-agent v0.21 removed ``agent.anthropic_adapter.normalize_anthropic_response``
and ``AIAgent._normalize_codex_response``. The sanctioned out-of-band contract is
``agent._get_transport(<mode>).normalize_response(...)`` returning a
``NormalizedResponse`` (``.content``).
These tests pin the two out-of-band consumers against a v0.21-shaped fake agent
(no legacy normalizer symbols exist; the transport normalizes):
- ``api.streaming.generate_title_raw_via_agent`` (anthropic_messages branch)
- the nested ``_agent_text_completion`` inside
``api.routes._handle_handoff_summary`` (codex_responses branch)
On the pre-port code both branches die (ImportError / AttributeError) and
silently fall back to empty titles / local fallback summaries; these tests
fail in that state.
"""
import sys
import types
TITLE_TEXT = 'Claude Transport Title'
SUMMARY_TEXT = '- The remaining work is the final review pass.'
class _FakeTransport:
"""v0.21-shaped transport: normalize_response returns a NormalizedResponse-like object."""
def __init__(self, content):
self._content = content
self.calls = []
def normalize_response(self, resp, **kwargs):
self.calls.append((resp, kwargs))
return types.SimpleNamespace(content=self._content)
def _make_anthropic_fake_agent(transport):
"""Fake agent exposing only the v0.21 surface the title branch needs."""
class _FakeAgent:
api_mode = 'anthropic_messages'
model = 'anthropic/claude-sonnet-4-5'
provider = 'anthropic'
base_url = ''
def __init__(self):
self.reasoning_config = None
self._is_anthropic_oauth = True
def _anthropic_preserve_dots(self):
return False
def _anthropic_messages_create(self, api_kwargs):
# Opaque response blob; normalizing is the transport's job on v0.21.
return object()
def _get_transport(self, mode=None):
assert mode is None
return transport
return _FakeAgent()
def test_anthropic_title_uses_transport_normalize_response(monkeypatch):
"""The anthropic_messages title branch must normalize via the agent transport.
v0.21 deleted ``normalize_anthropic_response``; the branch must not import it.
"""
from api.streaming import generate_title_raw_via_agent
# Simulate the v0.21 adapter module in-process (works in agent-less CI too):
# it still exports build_anthropic_kwargs but deliberately does NOT define
# normalize_anthropic_response — exactly the environment this port targets.
fake_adapter = types.ModuleType('agent.anthropic_adapter')
fake_adapter.build_anthropic_kwargs = lambda **kwargs: {}
fake_agent_pkg = sys.modules.get('agent') or types.ModuleType('agent')
fake_agent_pkg.anthropic_adapter = fake_adapter
monkeypatch.setitem(sys.modules, 'agent', fake_agent_pkg)
monkeypatch.setitem(sys.modules, 'agent.anthropic_adapter', fake_adapter)
transport = _FakeTransport(TITLE_TEXT)
agent = _make_anthropic_fake_agent(transport)
result, status = generate_title_raw_via_agent(
agent,
user_text='Hey nur ein kurzer Test',
assistant_text='Alles klar, ich helfe dir dabei.',
)
assert result == TITLE_TEXT
assert status == 'llm'
assert transport.calls, 'transport normalize_response must be called'
assert transport.calls[0][1].get('strip_tool_prefix') is True
def test_codex_handoff_summary_uses_transport_normalize_response(monkeypatch):
"""The codex_responses handoff branch must normalize via the agent transport.
v0.21 deleted ``AIAgent._normalize_codex_response``; the branch must use
``agent._get_transport('codex_responses').normalize_response(...)``.
"""
import api.config as cfg
import api.models as models
import api.routes as routes
monkeypatch.setattr(routes, 'require', lambda body, *keys: None)
monkeypatch.setattr(
routes,
'bad',
lambda _handler, msg, status=400: {'ok': False, 'error': msg, 'status': status},
)
monkeypatch.setattr(
routes,
'j',
lambda _handler, payload, status=200, extra_headers=None: payload,
)
monkeypatch.setattr(
routes,
'_persist_handoff_summary',
lambda sid, summary, channel, rounds, fallback=False: persisted.append(
{
'sid': sid,
'summary': summary,
'channel': channel,
'rounds': rounds,
'fallback': fallback,
}
)
or {'ok': True},
)
monkeypatch.setattr(
models,
'count_conversation_rounds',
lambda sid, since=None: models.CONVERSATION_ROUND_THRESHOLD,
)
monkeypatch.setattr(
models,
'get_cli_session_messages',
lambda sid: [
{'role': 'user', 'content': 'What remains to do?', 'timestamp': 1.0},
{'role': 'assistant', 'content': 'One review step remains.', 'timestamp': 2.0},
{'role': 'user', 'content': 'And after that?', 'timestamp': 3.0},
{'role': 'assistant', 'content': 'Then we ship.', 'timestamp': 4.0},
],
)
monkeypatch.setattr(
cfg,
'resolve_model_provider',
lambda resolved_model=None: (
'gpt-test',
'openai-codex',
'https://chatgpt.com/backend-api/codex',
),
)
persisted = []
transport = _FakeTransport(SUMMARY_TEXT)
class _CodexAgent:
api_mode = 'codex_responses'
def __init__(self, *args, **kwargs):
self.model = kwargs.get('model')
self.provider = kwargs.get('provider')
self.base_url = kwargs.get('base_url')
self.reasoning_config = None
def _build_api_kwargs(self, api_messages):
return {'model': self.model, 'instructions': 'summary', 'input': [], 'store': False}
def _run_codex_stream(self, kwargs):
# Opaque response blob; normalizing is the transport's job on v0.21.
return object()
def _get_transport(self, mode=None):
assert mode == 'codex_responses'
return transport
def release_clients(self):
return None
# v0.21: the legacy method is gone from the agent.
assert not hasattr(_CodexAgent, '_normalize_codex_response')
fake_run_agent = types.ModuleType('run_agent')
fake_run_agent.AIAgent = _CodexAgent
monkeypatch.setitem(sys.modules, 'run_agent', fake_run_agent)
fake_runtime_module = types.ModuleType('hermes_cli.runtime_provider')
fake_runtime_module.resolve_runtime_provider = lambda requested=None: {
'api_key': 'x',
'provider': 'openai-codex',
'base_url': 'https://chatgpt.com/backend-api/codex',
}
fake_hermes_cli = types.ModuleType('hermes_cli')
fake_hermes_cli.__path__ = []
fake_hermes_cli.runtime_provider = fake_runtime_module
monkeypatch.setitem(sys.modules, 'hermes_cli', fake_hermes_cli)
monkeypatch.setitem(sys.modules, 'hermes_cli.runtime_provider', fake_runtime_module)
response = routes._handle_handoff_summary(
object(), {'session_id': 'session-codex-v021-transport'}
)
assert response['ok'] is True
assert response['fallback'] is False
assert response['summary'] == SUMMARY_TEXT
assert persisted, 'handoff summary must be persisted'
assert persisted[0]['summary'] == SUMMARY_TEXT
assert persisted[0]['fallback'] is False
+7 -3
View File
@@ -9,6 +9,7 @@ import textwrap
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
BOOT_JS = ROOT / "static" / "boot.js"
WORKSPACE_JS = ROOT / "static" / "workspace.js"
SESSIONS_JS = ROOT / "static" / "sessions.js"
UI_JS = ROOT / "static" / "ui.js"
@@ -208,12 +209,15 @@ def test_api_has_default_timeout_and_per_call_override_contract():
def test_update_flows_keep_explicit_longer_timeouts():
"""Legitimately long update flows should not inherit the generic 30s guard."""
"""Network-bound update checks need a single-attempt 300s latency budget."""
boot = _source(BOOT_JS)
src = _source(UI_JS)
panels = _source(PANELS_JS)
# /api/updates/check builds its body in a _checkBody var (to optionally add
# an explicit channel), but must still carry the 60s timeout override.
assert "api('/api/updates/check',{method:'POST',body:JSON.stringify(_checkBody),timeoutMs:60000})" in panels
# an explicit channel). Both boot and manual callers need enough time for
# the endpoint's bounded fetches; api() does not retry ordinary timeouts.
assert "api(_checkUrl,{method:_testUpdates?'GET':'POST',body:_testUpdates?undefined:JSON.stringify({force:false}),timeoutMs:300000})" in boot
assert "api('/api/updates/check',{method:'POST',body:JSON.stringify(_checkBody),timeoutMs:300000})" in panels
assert "api('/api/updates/summary',{method:'POST',body:JSON.stringify({updates:scopedUpdates,target:target||null}),timeoutMs:60000})" in src
# apply/force now build their body inline to optionally carry the offered
# channel (Codex debounce-race fix), but MUST still carry the 120s override.
+4 -2
View File
@@ -113,8 +113,10 @@ class TestSSEStaticAnalysis:
cb_start = streaming_src.find("def _approval_notify_cb(approval_data):")
cb_end = streaming_src.find("_reg_notify(session_id, _approval_notify_cb)", cb_start)
cb_body = streaming_src[cb_start:cb_end]
assert "head, total = _submit_pending_for_polling(session_id, approval_data)" in cb_body, \
"_approval_notify_cb must mirror approval data into polling state before SSE"
assert "auto_resolved, head, total = _settle_pending_for_polling(" in cb_body, \
"_approval_notify_cb must settle local admission before SSE"
assert "if auto_resolved and head is None:" in cb_body, \
"_approval_notify_cb must suppress SSE for an auto-resolved local approval"
assert '"pending_count": total' in cb_body, \
"_approval_notify_cb must emit the reconciled pending count"
assert "put('approval', approval_data)" in cb_body, \
+23 -11
View File
@@ -232,8 +232,10 @@ class TestApprovalModuleExports:
assert cb_start != -1, "_approval_notify_cb must exist"
cb_end = STREAMING_SRC.find("_reg_notify(session_id, _approval_notify_cb)", cb_start)
cb_body = STREAMING_SRC[cb_start:cb_end]
assert "head, total = _submit_pending_for_polling(session_id, approval_data)" in cb_body, \
"approval notify callback must mirror approval data into polling state"
assert "auto_resolved, head, total = _settle_pending_for_polling(" in cb_body, \
"approval notify callback must settle admission at the YOLO handoff boundary"
assert "if auto_resolved and head is None:" in cb_body, \
"an auto-resolved local approval must not publish a stale card"
assert '"pending_count": total' in cb_body, \
"approval notify callback must publish the reconciled pending count"
assert "put('approval', approval_data)" in cb_body, \
@@ -418,13 +420,20 @@ class TestApprovalHTTPEndpoints:
with _lock:
r._pending.pop(sid, None)
r._gateway_queues.pop(sid, None)
r._gateway_queues[sid] = [_ApprovalEntry(approval_a)]
entry_a = _ApprovalEntry(approval_a)
r._gateway_queues[sid] = [entry_a]
try:
ra.submit_gateway_pending_mirror(sid, approval_a)
# Production notifies WebUI with a COPY of the entry payload
# (core: notify_cb(dict(entry.data))), which carries the entry's
# stamped request_id — pass that faithful copy, not the pre-stamp
# source dict, so the mirror matches its producer the way it does
# in production.
ra.submit_gateway_pending_mirror(sid, dict(entry_a.data))
with _lock:
r._gateway_queues.pop(sid, None)
r._gateway_queues[sid] = [_ApprovalEntry(approval_b)]
ra.submit_gateway_pending_mirror(sid, approval_b)
entry_b = _ApprovalEntry(approval_b)
r._gateway_queues[sid] = [entry_b]
ra.submit_gateway_pending_mirror(sid, dict(entry_b.data))
parsed = urllib.parse.urlparse(f"/api/approval/pending?session_id={urllib.parse.quote(sid)}")
r._handle_approval_pending(object(), parsed)
@@ -507,7 +516,8 @@ class TestApprovalHTTPEndpoints:
r._pending.pop(sid, None)
r._gateway_queues[sid] = [entry_a]
try:
ra.submit_gateway_pending_mirror(sid, approval_a)
# Faithful to production: submit the stamped copy (see note above).
ra.submit_gateway_pending_mirror(sid, dict(entry_a.data))
with _lock:
mirror_aid_a = r._pending[sid][0]["approval_id"]
@@ -516,7 +526,7 @@ class TestApprovalHTTPEndpoints:
entry_b = _ApprovalEntry(approval_b)
with _lock:
r._gateway_queues[sid] = [entry_b]
ra.submit_gateway_pending_mirror(sid, approval_b)
ra.submit_gateway_pending_mirror(sid, dict(entry_b.data))
resolved = r._resolve_approval_legacy(sid, mirror_aid_a, "once")
assert resolved is False, "stale approval_id must not resolve live B"
@@ -617,7 +627,8 @@ class TestApprovalHTTPEndpoints:
r._pending.pop(sid, None)
r._gateway_queues[sid] = [entry, sibling_entry]
try:
ra.submit_gateway_pending_mirror(sid, approval)
# Faithful to production: submit the stamped copy (see note above).
ra.submit_gateway_pending_mirror(sid, dict(entry.data))
with _lock:
approval_id = r._pending[sid][0]["approval_id"]
@@ -728,8 +739,9 @@ class TestApprovalHTTPEndpoints:
r._pending.pop(sid, None)
r._gateway_queues[sid] = [entry_a, entry_b]
try:
ra.submit_gateway_pending_mirror(sid, approval_a)
ra.submit_gateway_pending_mirror(sid, approval_b)
# Faithful to production: submit the stamped copies (see note above).
ra.submit_gateway_pending_mirror(sid, dict(entry_a.data))
ra.submit_gateway_pending_mirror(sid, dict(entry_b.data))
with _lock:
approval_b_id = next(
item["approval_id"] for item in r._pending[sid]
+160
View File
@@ -0,0 +1,160 @@
"""TOCTOU hardening for chat-attachment uploads.
handle_upload deduped filenames with an exists() check and then wrote with
plain write_bytes — two concurrent uploads of the same name could both pass
the check and the last writer silently won. These tests pin the #3398-style
anchored O_CREAT|O_EXCL|O_NOFOLLOW creation semantics (mirrored from the
workspace upload path): a raced duplicate must 409 instead of overwriting,
and the non-raced happy path / dedup behavior must stay unchanged.
Test-pattern cribbed from tests/test_raw_audio_upload.py (real multipart body
through parse_multipart + fake handler) and
tests/test_session_active_profile_authorization.py (monkeypatched
get_session / _get_active_profile_name).
"""
import io
import json
from pathlib import Path
from types import SimpleNamespace
import pytest
import api.upload as upload
from api.upload import handle_upload
def _multipart_body(fields=None, files=None, boundary=b"testboundary"):
fields = fields or {}
files = files or {}
body = b""
for name, value in fields.items():
body += b"--" + boundary + b"\r\n"
body += f'Content-Disposition: form-data; name="{name}"\r\n\r\n'.encode()
body += str(value).encode() + b"\r\n"
for name, (filename, data, content_type) in files.items():
body += b"--" + boundary + b"\r\n"
body += (
f'Content-Disposition: form-data; name="{name}"; filename="{filename}"\r\n'
f"Content-Type: {content_type}\r\n\r\n"
).encode()
body += data + b"\r\n"
body += b"--" + boundary + b"--\r\n"
return body, f"multipart/form-data; boundary={boundary.decode()}"
class _FakeHandler:
def __init__(self, body: bytes, content_type: str):
self.rfile = io.BytesIO(body)
self.wfile = io.BytesIO()
self.headers = {
"Content-Type": content_type,
"Content-Length": str(len(body)),
}
self.status = None
self.sent_headers = {}
def send_response(self, status):
self.status = status
def send_header(self, key, value):
self.sent_headers[key] = value
def end_headers(self):
pass
def payload(self):
return json.loads(self.wfile.getvalue().decode("utf-8"))
@pytest.fixture()
def attachment_env(tmp_path, monkeypatch):
"""Isolate the attachment inbox and stub the session lookup."""
root = tmp_path / "attachments"
monkeypatch.setenv("HERMES_WEBUI_ATTACHMENT_DIR", str(root))
monkeypatch.setattr(
upload,
"get_session",
lambda sid: SimpleNamespace(session_id=sid, profile=None),
)
monkeypatch.setattr(upload, "_get_active_profile_name", lambda: "default")
return root
def test_raced_duplicate_returns_409_and_preserves_existing_bytes(attachment_env, monkeypatch):
"""An attacker winning the exists()-check race must not get its bytes written.
Simulates the race deterministically: the destination file already exists
with known content, and _upload_destination returns that exact path as if
the dedup check had just passed. The anchored O_EXCL create must fail with
FileExistsError -> 409, and the pre-existing bytes must be untouched.
"""
session_id = "race-sess"
dest_dir = upload._session_attachment_dir(session_id)
dest_dir.mkdir(parents=True, exist_ok=True)
target = dest_dir / "report.txt"
target.write_bytes(b"ORIGINAL-CONTENT")
monkeypatch.setattr(
upload,
"_upload_destination",
lambda session_id, safe_name, dest_dir=None: target,
)
body, content_type = _multipart_body(
fields={"session_id": session_id},
files={"file": ("report.txt", b"ATTACKER-BYTES", "text/plain")},
)
handler = _FakeHandler(body, content_type)
handle_upload(handler)
assert handler.status == 409
assert handler.payload() == {
"error": "Upload destination already exists: report.txt"
}
assert target.read_bytes() == b"ORIGINAL-CONTENT"
def test_new_upload_happy_path_unchanged(attachment_env):
"""A normal upload of a fresh name keeps the exact response shape."""
body, content_type = _multipart_body(
fields={"session_id": "happy-sess"},
files={"file": ("notes.txt", b"hello world", "text/plain")},
)
handler = _FakeHandler(body, content_type)
handle_upload(handler)
assert handler.status == 200
payload = handler.payload()
assert set(payload) == {"filename", "path", "size", "mime", "is_image"}
assert payload["filename"] == "notes.txt"
assert payload["mime"].startswith("text/")
assert payload["is_image"] is False
assert payload["size"] == len(b"hello world")
dest = attachment_env / "happy-sess" / "notes.txt"
assert Path(payload["path"]) == dest.resolve()
assert dest.read_bytes() == b"hello world"
def test_non_raced_dedup_still_suffixed(attachment_env):
"""Uploading an existing name (no race) still picks the -1 suffixed name."""
body1, ctype1 = _multipart_body(
fields={"session_id": "dedup-sess"},
files={"file": ("dup.txt", b"first", "text/plain")},
)
h1 = _FakeHandler(body1, ctype1)
handle_upload(h1)
assert h1.status == 200
body2, ctype2 = _multipart_body(
fields={"session_id": "dedup-sess"},
files={"file": ("dup.txt", b"second", "text/plain")},
)
h2 = _FakeHandler(body2, ctype2)
handle_upload(h2)
assert h2.status == 200
assert h2.payload()["filename"] == "dup-1.txt"
dest_dir = upload._session_attachment_dir("dedup-sess")
assert (dest_dir / "dup.txt").read_bytes() == b"first"
assert (dest_dir / "dup-1.txt").read_bytes() == b"second"
@@ -0,0 +1,253 @@
"""Regression coverage for provider-qualified auxiliary model persistence.
``GET /api/models`` may expose WebUI-only ``@provider:model`` routing IDs, but
auxiliary configuration stores provider and model in separate fields. The
provider-native model value must be used by both the settings UI and the shared
backend persistence boundary.
"""
from __future__ import annotations
import json
import shutil
import subprocess
from pathlib import Path
import pytest
REPO = Path(__file__).resolve().parents[1]
PANELS_JS_PATH = REPO / "static" / "panels.js"
NODE = shutil.which("node")
@pytest.mark.skipif(NODE is None, reason="node not available")
def test_auxiliary_picker_uses_provider_native_model_values():
"""A custom-provider catalog must render and select provider-native values."""
script = r"""
const fs = require('fs');
const src = fs.readFileSync(process.argv[1], 'utf8');
function extract(name){
const re = new RegExp('function\\s+' + name + '\\s*\\(');
const start = src.search(re);
if(start < 0) return '';
let i = src.indexOf('{', start);
let depth = 0;
while(i < src.length){
const ch = src[i];
if(ch === '{') depth += 1;
else if(ch === '}'){
depth -= 1;
if(depth === 0) return src.slice(start, i + 1);
}
i += 1;
}
throw new Error(name + ' parse failed');
}
global.t = () => '';
global.document = {
createElement: () => ({value:'', textContent:'', selected:false}),
};
function makeSelect(){
return {
children: [],
_value: '',
set innerHTML(_value){ this.children = []; this._value = ''; },
get innerHTML(){ return ''; },
get options(){ return this.children; },
appendChild(opt){ this.children.push(opt); },
insertBefore(opt, before){
const idx = this.children.indexOf(before);
if(idx < 0) this.children.push(opt);
else this.children.splice(idx, 0, opt);
},
set value(value){
this._value = value;
for(const opt of this.children) opt.selected = opt.value === value;
},
get value(){
const selected = this.children.find((opt) => opt.selected);
return selected ? selected.value : this._value;
},
};
}
const sharedHelper = extract('_modelBareNameForProvider');
const buildOptions = extract('_buildAuxModelOptions');
if(!buildOptions) throw new Error('_buildAuxModelOptions not found');
eval(sharedHelper + '\n' + buildOptions);
const select = makeSelect();
const providers = [{
slug: 'my-local-ai-gateway',
name: 'My Local AI Gateway',
models: [
{
id: '@my-local-ai-gateway:example-chat-model',
label: 'Example Chat Model',
},
{
id: '@my-local-ai-gateway:example-side-model',
label: '@my-local-ai-gateway:example-side-model',
},
{id: 'vendor/example-model', label: 'Vendor Example Model'},
{id: 'example-chat-model', label: 'Duplicate Chat Model'},
],
}];
const canonicalCurrent = _buildAuxModelOptions(
select,
'my-local-ai-gateway',
providers,
'@my-local-ai-gateway:example-side-model',
);
const modelOptions = select.options.filter(
(opt) => opt.value && opt.value !== '__custom__',
);
const helperCases = sharedHelper ? [
_modelBareNameForProvider('@custom:router-alias:chat-model', 'custom'),
_modelBareNameForProvider('@custom:backup:model:free', 'custom:backup'),
_modelBareNameForProvider('vendor/example-model', 'my-local-ai-gateway'),
_modelBareNameForProvider('@other-gateway:other-model', 'my-local-ai-gateway'),
] : null;
console.log(JSON.stringify({
values: modelOptions.map((opt) => opt.value),
labels: modelOptions.map((opt) => opt.textContent),
selected: select.value,
canonicalCurrent,
helperCases,
}));
"""
proc = subprocess.run(
[NODE, "-e", script, str(PANELS_JS_PATH)],
capture_output=True,
text=True,
timeout=20,
)
assert proc.returncode == 0, f"node probe failed:\n{proc.stderr}"
result = json.loads(proc.stdout.strip().splitlines()[-1])
assert result == {
"values": [
"example-chat-model",
"example-side-model",
"vendor/example-model",
],
"labels": [
"Example Chat Model",
"example-side-model",
"Vendor Example Model",
],
"selected": "example-side-model",
"canonicalCurrent": "example-side-model",
"helperCases": [
"router-alias:chat-model",
"model:free",
"vendor/example-model",
"@other-gateway:other-model",
],
}
def test_matching_legacy_value_exposes_explicit_apply_repair():
"""Loading a safe legacy prefix should offer, but not force, a config write."""
source = PANELS_JS_PATH.read_text(encoding="utf-8")
assert "if(canonicalModel!==cfg.model) needsCanonicalSave=true;" in source
assert "applyBtn.style.display=needsCanonicalSave?'':'none'" in source
@pytest.mark.parametrize(
("provider", "requested_model", "persisted_model"),
[
(
"my-local-ai-gateway",
"@my-local-ai-gateway:example-side-model",
"example-side-model",
),
(
"custom",
"@custom:router-alias:chat-model",
"router-alias:chat-model",
),
("custom:backup", "@custom:backup:model:free", "model:free"),
(
"my-local-ai-gateway",
"vendor/example-model",
"vendor/example-model",
),
(
"my-local-ai-gateway",
"example-side-model",
"example-side-model",
),
],
)
def test_set_auxiliary_model_persists_provider_native_model(
monkeypatch,
tmp_path,
provider,
requested_model,
persisted_model,
):
from api import config
config_path = tmp_path / "config.yaml"
config_path.write_text(
"auxiliary:\n vision:\n provider: auto\n model: ''\n",
encoding="utf-8",
)
monkeypatch.setattr(config, "_get_config_path", lambda: config_path)
monkeypatch.setattr(config, "reload_config", lambda: None)
monkeypatch.setattr(
config,
"resolve_model_provider",
lambda model: (model, provider, None),
)
result = config.set_auxiliary_model("vision", provider, requested_model)
saved = config._load_yaml_config_file(config_path)["auxiliary"]["vision"]
assert saved["provider"] == provider
assert saved["model"] == persisted_model
assert result["provider"] == provider
assert result["model"] == persisted_model
@pytest.mark.parametrize(
("provider", "model"),
[
("my-local-ai-gateway", "@other-gateway:other-model"),
("my-local-ai-gateway", "@my-local-ai-gateway:"),
("auto", "@my-local-ai-gateway:example-side-model"),
],
)
def test_set_auxiliary_model_rejects_invalid_qualified_pair_without_write(
monkeypatch,
tmp_path,
provider,
model,
):
from api import config
config_path = tmp_path / "config.yaml"
original = (
"auxiliary:\n"
" vision:\n"
" provider: openai\n"
" model: gpt-5.5\n"
)
config_path.write_text(original, encoding="utf-8")
monkeypatch.setattr(config, "_get_config_path", lambda: config_path)
monkeypatch.setattr(config, "reload_config", lambda: None)
with pytest.raises(ValueError, match="provider-qualified auxiliary model"):
config.set_auxiliary_model("vision", provider, model)
assert config_path.read_text(encoding="utf-8") == original
+464
View File
@@ -312,6 +312,280 @@ console.log(JSON.stringify({
"Missing _buildAuxModelOptions() for model dropdown rebuild"
)
@pytest.mark.skipif(NODE is None, reason="node not on PATH")
def test_configured_aux_provider_absent_from_catalog_is_preserved(self):
"""#7486: a configured provider missing from /api/models must survive Apply.
The aux provider <select> is populated with 'auto' plus the /api/models
catalog. When the configured provider is absent from that catalog (e.g.
its group is filtered out because it exposes no models), no option
matched, so the select fell back to its first entry ('auto') and the
next Apply persisted 'auto' — silently discarding the configured value.
"""
script = r"""
const fs = require('fs');
const src = fs.readFileSync(process.argv[1], 'utf8');
function extract(name){
const re = new RegExp('function\\s+' + name + '\\s*\\(');
const start = src.search(re);
if(start < 0) throw new Error(name + ' not found');
let i = src.indexOf('{', start);
let depth = 0;
while(i < src.length){
const ch = src[i];
if(ch === '{') depth += 1;
else if(ch === '}') {
depth -= 1;
if(depth === 0){
break;
}
}
i += 1;
}
if(depth !== 0) throw new Error(name + ' parse failed');
return src.slice(start, i + 1);
}
global.t = (key) => key;
// Minimal <select>/<option> model. The value getter mirrors real browser
// behavior: with no option explicitly selected a single-select reads back its
// first option, which is what made the downgrade silent.
function makeSelect(){
const sel = { options: [] };
Object.defineProperty(sel, 'innerHTML', {
get(){ return ''; },
set(v){ if(v === '') sel.options = []; },
});
sel.appendChild = (node) => { sel.options.push(node); return node; };
sel.insertBefore = (node, ref) => {
const idx = sel.options.indexOf(ref);
if(idx < 0) sel.options.push(node); else sel.options.splice(idx, 0, node);
return node;
};
Object.defineProperty(sel, 'value', {
get(){
const picked = sel.options.find((opt) => opt.selected);
if(picked) return picked.value;
return sel.options.length ? sel.options[0].value : '';
},
set(v){ sel.options.forEach((opt) => { opt.selected = opt.value === v; }); },
});
return sel;
}
global.document = {
createElement(tag){
if(String(tag).toLowerCase() === 'option'){
return { tagName: 'OPTION', value: '', textContent: '', selected: false };
}
return { tagName: String(tag).toUpperCase(), value: '', textContent: '' };
},
};
eval(extract('_buildAuxProviderOptions'));
const catalog = [
{slug: 'openai', name: 'OpenAI'},
{slug: 'anthropic', name: 'Anthropic'},
];
function build(providers, currentProvider){
const sel = makeSelect();
_buildAuxProviderOptions(sel, providers, currentProvider);
return {
values: sel.options.map((opt) => opt.value),
labels: sel.options.map((opt) => opt.textContent),
selected: sel.options.filter((opt) => opt.selected).map((opt) => opt.value),
value: sel.value,
};
}
console.log(JSON.stringify({
absent: build(catalog, 'custom-router'),
absentEmptyCatalog: build([], 'custom-router'),
present: build(catalog, 'anthropic'),
auto: build(catalog, 'auto'),
empty: build(catalog, ''),
}));
"""
proc = subprocess.run(
[NODE, "-e", script, str(PANELS_JS_PATH)],
capture_output=True,
text=True,
timeout=20,
)
assert proc.returncode == 0, f"node probe failed:\n{proc.stderr}"
result = json.loads(proc.stdout.strip().splitlines()[-1])
absent = result["absent"]
assert "custom-router" in absent["values"], (
"configured provider absent from the catalog must keep a selectable option"
)
assert absent["selected"] == ["custom-router"], (
f"configured option must be selected, got {absent['selected']}"
)
assert absent["value"] == "custom-router", (
"Apply must read back the configured provider, not the fallback first option"
)
assert "custom-router" in absent["labels"][-1]
assert result["absentEmptyCatalog"]["value"] == "custom-router"
assert result["present"]["value"] == "anthropic"
assert result["present"]["values"] == ["auto", "openai", "anthropic"]
assert result["present"]["selected"] == ["anthropic"]
assert result["auto"]["value"] == "auto"
assert result["empty"]["value"] == "auto"
assert result["empty"]["values"] == ["auto", "openai", "anthropic"]
def test_apply_does_not_downgrade_untouched_aux_row(self):
"""#7486: applying one row must not re-save another row's provider as 'auto'.
End-to-end over the Apply path: two task rows are built with the real
option builder, the user edits only the second row's model, and the
first row keeps a provider that is missing from the catalog. The first
row must not produce any POST at all — the destructive symptom was it
being re-saved with provider 'auto'.
"""
script = r"""
const fs = require('fs');
const src = fs.readFileSync(process.argv[1], 'utf8');
function extract(name){
const re = new RegExp('function\\s+' + name + '\\s*\\(');
const start = src.search(re);
if(start < 0) throw new Error(name + ' not found');
let i = src.indexOf('{', start);
let depth = 0;
while(i < src.length){
const ch = src[i];
if(ch === '{') depth += 1;
else if(ch === '}') {
depth -= 1;
if(depth === 0){
break;
}
}
i += 1;
}
if(depth !== 0) throw new Error(name + ' parse failed');
return src.slice(start, i + 1);
}
// Minimal <select>/<option> shim. The value getter mirrors the browser: a
// single select with nothing explicitly selected reads back its first option,
// which is why the downgrade was silent rather than an explicit error.
function makeOption(value, label){
return {
tagName: 'OPTION',
value: String(value),
textContent: label === undefined ? String(value) : label,
selected: false,
};
}
function makeSelect(){
const sel = {options: []};
Object.defineProperty(sel, 'innerHTML', {
get(){ return ''; },
set(v){ if(v === '') sel.options = []; },
});
sel.appendChild = (node) => { sel.options.push(node); return node; };
sel.insertBefore = (node, ref) => {
const idx = sel.options.indexOf(ref);
if(idx < 0) sel.options.push(node); else sel.options.splice(idx, 0, node);
return node;
};
Object.defineProperty(sel, 'value', {
get(){
const picked = sel.options.find(o => o.selected);
if(picked) return picked.value;
return sel.options.length ? sel.options[0].value : '';
},
set(v){ sel.options.forEach(o => { o.selected = (o.value === String(v)); }); },
});
return sel;
}
function setSelect(sel, values, chosen){
sel.innerHTML = '';
for(const v of values) sel.appendChild(makeOption(v));
sel.value = chosen;
return sel;
}
global.document = {
createElement(tag){
return String(tag).toLowerCase() === 'option'
? makeOption('', '')
: {tagName: String(tag).toUpperCase(), style: {}, appendChild(){}};
},
};
const els = {};
global.$ = (id) => els[id] || null;
global.t = (key) => key;
global.showToast = () => {};
global._loadAuxiliaryModels = () => {};
eval(extract('_buildAuxProviderOptions'));
const catalog = [{slug: 'openai', name: 'OpenAI'}];
// Persisted config: 'simple' is configured with a provider that is no longer
// present in /api/models; 'complex' uses a catalog provider.
global._auxTasks = [{task: 'simple'}, {task: 'complex'}];
global._auxOriginalConfig = {
simple: {provider: 'custom-router', model: 'gpt-x'},
complex: {provider: 'openai', model: 'gpt-5'},
};
els['aux-prov-simple'] = makeSelect();
els['aux-model-simple'] = setSelect(makeSelect(), ['gpt-x', '__custom__'], 'gpt-x');
els['aux-prov-complex'] = makeSelect();
els['aux-model-complex'] = setSelect(makeSelect(), ['gpt-5', 'gpt-5-mini', '__custom__'], 'gpt-5');
// The settings panel rebuilds every row's selects on load
_buildAuxProviderOptions(els['aux-prov-simple'], catalog, 'custom-router');
_buildAuxProviderOptions(els['aux-prov-complex'], catalog, 'openai');
// The user edits only the 'complex' row, then clicks Apply
els['aux-model-complex'].value = 'gpt-5-mini';
const posted = [];
global.api = async (path, opts) => { posted.push(JSON.parse(opts.body)); return {}; };
// _applyAuxModels is async: extract() slices from the `function` keyword, so
// re-add the modifier before evaluating the declaration.
eval('async ' + extract('_applyAuxModels'));
_applyAuxModels().then(() => {
console.log(JSON.stringify({posted, untouchedProvider: els['aux-prov-simple'].value}));
});
"""
proc = subprocess.run(
[NODE, "-e", script, str(PANELS_JS_PATH)],
capture_output=True,
text=True,
timeout=20,
)
assert proc.returncode == 0, f"node probe failed:\n{proc.stderr}"
result = json.loads(proc.stdout.strip().splitlines()[-1])
assert result["untouchedProvider"] == "custom-router", (
"the untouched row must still read back its configured provider"
)
assert result["posted"] == [
{
"scope": "auxiliary",
"task": "complex",
"provider": "openai",
"model": "gpt-5-mini",
}
], f"only the edited row may be saved, got {result['posted']}"
assert not any(p.get("provider") == "auto" for p in result["posted"]), (
"an untouched row must never be persisted as 'auto' (data loss)"
)
def test_custom_model_prompt(self):
"""Selecting 'Custom model…' must prompt for model ID."""
assert "__custom__" in PANELS_JS, (
@@ -937,3 +1211,193 @@ class TestAuxiliaryModelsBackend:
assert result["status"] == 400
assert "Unknown auxiliary task slot" in result["error"]
def test_aux_slot_base_url_uses_selected_provider_not_active_endpoint(
self, monkeypatch, tmp_path
):
"""Overlapping-id sibling fix: persisting an auxiliary slot for a named
custom provider must record THAT provider's own base_url, not the active
main provider's endpoint.
Repro: main provider is custom:dogapi (base_url dogapi), and both dogapi
and packyapi list 'shared-model'. Selecting custom:packyapi for the vision
slot previously persisted {provider: custom:packyapi, base_url: dogapi's}
because the base_url was resolved with a bare resolve_model_provider(model)
that ignores the selected provider. It must persist packyapi's base_url.
"""
from api import config
shared_cfg = {
"model": {
"default": "shared-model",
"provider": "custom",
"base_url": "https://www.dogapi.cc/v1",
},
"custom_providers": [
{"name": "dogapi", "base_url": "https://www.dogapi.cc/v1",
"models": ["shared-model"]},
{"name": "packyapi", "base_url": "https://www.packyapi.ai/v1",
"models": ["shared-model"]},
],
}
config_path = tmp_path / "config.yaml"
import yaml
config_path.write_text(yaml.safe_dump(shared_cfg), encoding="utf-8")
monkeypatch.setattr(config, "_get_config_path", lambda: config_path)
monkeypatch.setattr(config, "reload_config", lambda: None)
# base_url resolution reads the in-memory cfg / get_config snapshot.
monkeypatch.setattr(config, "cfg", dict(shared_cfg))
monkeypatch.setattr(config, "get_config", lambda: dict(shared_cfg))
config.set_auxiliary_model("vision", "custom:packyapi", "shared-model")
saved = config._load_yaml_config_file(config_path)["auxiliary"]["vision"]
assert saved["provider"] == "custom:packyapi"
assert saved["base_url"] == "https://www.packyapi.ai/v1", (
f"aux slot must persist the SELECTED provider's base_url, got "
f"{saved.get('base_url')!r}"
)
def test_aux_slot_custom_provider_base_url_no_deadlock(self, monkeypatch, tmp_path):
"""Regression: saving a named custom auxiliary model must not self-deadlock.
set_auxiliary_model() holds the non-reentrant _cfg_lock while resolving
the selected custom provider's base_url. The pre-fix code called
resolve_custom_provider_connection() -> get_config() ->
reload_config_if_stale(), which re-acquires _cfg_lock and hangs forever
whenever the config cache is stale or the profile path changed. This test
uses the REAL get_config (only reload_config, which runs AFTER the lock is
released, is stubbed) and forces get_config()'s reload branch, then runs
the call on a watchdog thread that fails the test if it hangs.
"""
import threading
import yaml
from api import config
shared_cfg = {
"model": {
"default": "shared-model",
"provider": "custom",
"base_url": "https://www.dogapi.cc/v1",
},
"custom_providers": [
{"name": "dogapi", "base_url": "https://www.dogapi.cc/v1",
"models": ["shared-model"]},
{"name": "packyapi", "base_url": "https://www.packyapi.ai/v1",
"models": ["shared-model"]},
],
}
config_path = tmp_path / "config.yaml"
config_path.write_text(yaml.safe_dump(shared_cfg), encoding="utf-8")
monkeypatch.setattr(config, "_get_config_path", lambda: config_path)
# reload_config runs AFTER _cfg_lock is released; stub it so the test
# doesn't mutate global module state. get_config stays REAL — that is the
# path the pre-fix code re-entered while holding the lock.
monkeypatch.setattr(config, "reload_config", lambda: None)
# Force the reload branch inside the real get_config the pre-fix code
# called: a mismatched cached path makes path_changed True.
monkeypatch.setattr(config, "_cfg_path", None, raising=False)
result_box: dict = {}
error_box: dict = {}
def _run():
try:
result_box["r"] = config.set_auxiliary_model(
"vision", "custom:packyapi", "shared-model"
)
except Exception as exc: # pragma: no cover - surfaced via join
error_box["e"] = exc
worker = threading.Thread(target=_run, daemon=True)
worker.start()
worker.join(timeout=10)
assert not worker.is_alive(), (
"set_auxiliary_model deadlocked while resolving a custom provider "
"base_url under _cfg_lock"
)
if "e" in error_box:
raise error_box["e"]
saved = config._load_yaml_config_file(config_path)["auxiliary"]["vision"]
assert saved["provider"] == "custom:packyapi"
assert saved["base_url"] == "https://www.packyapi.ai/v1", (
f"aux slot must persist the SELECTED provider's base_url, got "
f"{saved.get('base_url')!r}"
)
def test_aux_slot_custom_provider_slug_collision_fails_closed(self, monkeypatch, tmp_path):
"""Saving a named custom auxiliary model whose slug collides with another
config entry must fail closed, not silently persist the wrong endpoint.
The inline base_url resolution added for the deadlock fix shares the same
all-entry uniqueness helper, so a custom:foo-bar save with colliding
'Foo Bar' + 'foo-bar' entries raises AmbiguousCustomProviderError and
writes nothing (the slot stays 'auto'). Operates on the in-scope
config_data, so it remains lock-safe.
"""
import yaml
from api import config
shared_cfg = {
"auxiliary": {"vision": {"provider": "auto", "model": ""}},
"custom_providers": [
{"name": "Foo Bar", "base_url": "https://a.example/v1"},
{"name": "foo-bar", "base_url": "https://b.example/v1",
"models": ["shared-model"]},
],
}
config_path = tmp_path / "config.yaml"
config_path.write_text(yaml.safe_dump(shared_cfg), encoding="utf-8")
monkeypatch.setattr(config, "_get_config_path", lambda: config_path)
monkeypatch.setattr(config, "reload_config", lambda: None)
with pytest.raises(config.AmbiguousCustomProviderError):
config.set_auxiliary_model("vision", "custom:foo-bar", "shared-model")
# The ambiguous save must not have persisted: the slot stays 'auto'.
saved = config._load_yaml_config_file(config_path)["auxiliary"]["vision"]
assert saved.get("provider") == "auto", (
f"ambiguous aux save must not persist, got {saved!r}"
)
def test_aux_slot_parenthesized_name_collision_fails_closed(self, monkeypatch, tmp_path):
"""Finding #1 on the aux persistence path: a parenthesized-name collision
that a looser slug key MISSED ('Foo (Bar)' vs 'foo-bar', both producing
custom:foo-bar) must also fail closed on save.
This is the case the earlier collision key got wrong: it normalized
'Foo (Bar)' to 'foo-(bar)' and never saw the collision, so the aux save
could persist endpoint A while the credential lookup later returned B.
"""
import yaml
from api import config
shared_cfg = {
"auxiliary": {"vision": {"provider": "auto", "model": ""}},
"custom_providers": [
{"name": "Foo (Bar)", "base_url": "https://a.example/v1"},
{"name": "foo-bar", "base_url": "https://b.example/v1",
"models": ["shared-model"]},
],
}
config_path = tmp_path / "config.yaml"
config_path.write_text(yaml.safe_dump(shared_cfg), encoding="utf-8")
monkeypatch.setattr(config, "_get_config_path", lambda: config_path)
monkeypatch.setattr(config, "reload_config", lambda: None)
with pytest.raises(config.AmbiguousCustomProviderError):
config.set_auxiliary_model("vision", "custom:foo-bar", "shared-model")
saved = config._load_yaml_config_file(config_path)["auxiliary"]["vision"]
assert saved.get("provider") == "auto", (
f"ambiguous aux save must not persist, got {saved!r}"
)
+189 -39
View File
@@ -1,5 +1,5 @@
"""Regression: blank assistant turn (对话消失) — dead empty live-turn shell
survives a session-updated swap re-render and hides the settled answer.
"""Regression: blank assistant turn (对话消失) AND duplicate settled render (#6948)
— the live-turn preserve guard must require a PROVABLE live owner.
Root cause (reproduced + fixed on an isolated debug instance, 2026-07-01)
------------------------------------------------------------------------
@@ -15,20 +15,47 @@ whenever `INFLIGHT[sid]` existed:
}
When a turn's SSE dropped (S.activeStreamId cleared to null) but its
`INFLIGHT[sid]` entry was NOT cleaned, the live turn was a DEAD empty shell —
`INFLIGHT[sid]` entry was NOT cleaned, the live turn was a DEAD EMPTY shell —
avatar + an empty worklog group ("Processed Ns", no body/tool rows). On the
next `session-updated` self-heal swap (loadSession force + keepStaleUntilLoaded,
common under repeated self-wake restarts), the guard re-attached that empty
shell OVER the freshly-wiped transcript, pinning an avatar-only blank turn on
top of the already-persisted answer. That is the reported "对话消失".
Fix
---
Preserve the live turn ONLY when it is genuinely live: an active stream is
still running (`S.activeStreamId`) — the #3877 case — OR it already holds real
rendered content (`.msg-body`, `.tool-card-row`, or `.wl-reason`). A dead empty
shell (no content, no active stream) is no longer preserved, so the swap wipe
drops it and the settled transcript renders normally.
First fix (#5390)
-----------------
Gate preservation on "real rendered content OR an active stream":
const _hasRealLiveContent=!!_lt.querySelector(
'.msg-body, .tool-card-row, .wl-reason'
);
if(_hasRealLiveContent || S.activeStreamId){ _preservedLiveTurn=_lt; }
That stopped the EMPTY shell but made rendered content itself act as authority.
Second bug (#6948)
------------------
After an assistant turn COMPLETES, the same message can render twice in the
feed (first copy without model label, second with it). Data was always clean —
state.db, the sidecar, and /api/session each hold one row; the duplicate is a
rendering artifact: a stale live-turn DOM node survives the settled-transcript
swap and is re-attached on top of it. When the stream has ended (S.activeStreamId
nulled) but `INFLIGHT[sid]` has not been cleaned yet, `_hasRealLiveContent` is
still true (the completed body is in the DOM), so the DEAD live turn is
preserved and re-attached OVER the settled transcript — two copies.
Final contract (this file)
--------------------------
Preservation requires a PROVABLE live owner: current-session ownership, an
`INFLIGHT` owner, and EITHER an active stream (`S.activeStreamId`) OR explicit
live-assistant evidence in the current message projection (`S.messages` — a
client-side `_live` / `_activityBurstId` / `_liveSegmentSeq` marker merged in
from the INFLIGHT tail or a server journal snapshot, i.e. the reconnect /
terminal-projection case). Bare `.msg-body` presence never proves liveness: a
settled transcript has no live projection, so the durable settled transcript
wins once no live owner remains. Both reported regressions are covered: the
dead EMPTY shell (no projection → dropped) and the dead CONTENTFUL turn
(no projection → dropped, settled transcript renders exactly once).
"""
import pathlib
import re
@@ -54,24 +81,39 @@ def _preserve_guard_src():
class TestBlankLiveTurnPreserveGuard:
def test_guard_requires_real_content_or_active_stream(self):
def test_guard_requires_live_owner_not_dom_content(self):
guard = _preserve_guard_src()
# Must gate the preserve on real content OR an active stream — not merely
# on INFLIGHT existence.
assert "_hasRealLiveContent" in guard, (
"preserve guard must compute whether the live turn has real content"
)
assert ".msg-body" in guard and ".tool-card-row" in guard and ".wl-reason" in guard, (
"real-content check must look for a visible body / tool card / reason row"
)
# Must gate the preserve on a PROVABLE live owner — an active stream or
# explicit live-assistant evidence in the message projection — never on
# bare DOM content presence. (#6948)
assert "S.activeStreamId" in guard, (
"preserve guard must still preserve a genuinely-streaming turn (#3877)"
)
assert "_hasLiveAssistantProjection" in guard, (
"preserve guard must consult the current message projection for "
"live-assistant evidence"
)
assert "S.messages" in guard, (
"live-assistant evidence must come from the message projection"
)
# The projection check must look at assistant-role messages carrying a
# client-side live marker (_live / _activityBurstId / _liveSegmentSeq).
assert "role==='assistant'" in guard, (
"live-assistant evidence must be scoped to assistant messages"
)
assert "m._live" in guard and "_activityBurstId" in guard and "_liveSegmentSeq" in guard, (
"live-assistant evidence must accept the client-side live markers"
)
# The assignment must be inside the new conditional.
assert re.search(
r"if\(_hasRealLiveContent\s*\|\|\s*S\.activeStreamId\)\{\s*_preservedLiveTurn=_lt;",
r"if\(S\.activeStreamId\s*\|\|\s*_hasLiveAssistantProjection\)\{\s*_preservedLiveTurn=_lt;",
guard,
), "preserve assignment must be gated by (hasRealContent || activeStreamId)"
), "preserve assignment must be gated by (activeStreamId || live projection)"
# DOM-content presence must NOT act as authority (#6948 regression guard).
assert "_hasRealLiveContent" not in guard, (
"bare .msg-body presence must not prove liveness — a contentful dead "
"live turn re-attached over the settled transcript is the #6948 duplicate"
)
def test_runtime_rejects_dead_shell_preserves_live(self):
node = shutil.which("node")
@@ -81,30 +123,138 @@ class TestBlankLiveTurnPreserveGuard:
script = textwrap.dedent(
"""
const assert=require('assert');
// Mirror the guard's decision predicate exactly.
function guardWouldPreserve(lt, activeStreamId){
// Mirror the guard's decision predicate exactly: current-session
// ownership + INFLIGHT owner + (active stream || live projection).
function guardWouldPreserve(lt, activeStreamId, messages, inflightOwner, sessionId){
if(!inflightOwner) return false; // no-owner rejected
if(!lt) return false;
const hasReal=!!lt.querySelector('.msg-body, .tool-card-row, .wl-reason');
return hasReal || !!activeStreamId;
if(lt.dataset&&lt.dataset.sessionId&&lt.dataset.sessionId!==sessionId){
return false; // wrong-session rejected
}
const hasLiveProjection=Array.isArray(messages)&&messages.some(m=>
m&&m.role==='assistant'&&(m._live||m._activityBurstId!==undefined||m._liveSegmentSeq!==undefined)
);
return !!activeStreamId || hasLiveProjection;
}
// Minimal DOM element stub with querySelector over a class set.
function el(classes){
function el(classes, dataset){
const set=new Set(classes||[]);
return { querySelector(sel){
// sel is a comma list of .class tokens
return sel.split(',').map(s=>s.trim().replace(/^\\./,''))
.some(c=>set.has(c)) ? {} : null;
}};
return {
dataset: dataset||null,
querySelector(sel){
// sel is a comma list of .class tokens
return sel.split(',').map(s=>s.trim().replace(/^\\\\./,''))
.some(c=>set.has(c)) ? {} : null;
}
};
}
const settledMessages=[]; // no live markers
const liveMessages=[{role:'assistant',content:'answer',_live:true}];
const burstLiveMessages=[{role:'assistant',content:'x',_activityBurstId:3}];
const segLiveMessages=[{role:'assistant',content:'x',_liveSegmentSeq:1}];
const deadShell = el([]); // empty worklog shell, no content
const withBody = el(['msg-body']);
const withTool = el(['tool-card-row']);
const withReason= el(['wl-reason']);
assert.strictEqual(guardWouldPreserve(deadShell, null), false, 'dead shell must NOT be preserved');
assert.strictEqual(guardWouldPreserve(deadShell, 'sid'), true, 'streaming empty shell preserved (#3877)');
assert.strictEqual(guardWouldPreserve(withBody, null), true, 'body content preserved');
assert.strictEqual(guardWouldPreserve(withTool, null), true, 'tool card preserved');
assert.strictEqual(guardWouldPreserve(withReason, null), true, 'reason row preserved');
const withBody = el(['msg-body']); // contentful (completed) turn
const wrongSess = el(['msg-body'], {sessionId:'other-sid'});
// #5390: dead empty shell must NOT be preserved (settled or live projection).
assert.strictEqual(guardWouldPreserve(deadShell, null, settledMessages, true, 's1'), false, 'dead shell must NOT be preserved');
// #3877: streaming shell IS preserved regardless of projection.
assert.strictEqual(guardWouldPreserve(deadShell, 'sid', settledMessages, true, 's1'), true, 'streaming empty shell preserved (#3877)');
// #6948: a CONTENTFUL dead live node (stream ended, stale INFLIGHT,
// settled projection) must NOT be preserved — this is the duplicate.
assert.strictEqual(guardWouldPreserve(withBody, null, settledMessages, true, 's1'), false, 'contentful dead turn must NOT be preserved (#6948)');
// Reconnect: contentful DOM + explicit live projection IS preserved.
assert.strictEqual(guardWouldPreserve(withBody, null, liveMessages, true, 's1'), true, 'reconnect live projection preserved');
assert.strictEqual(guardWouldPreserve(withBody, 'sid', settledMessages, true, 's1'), true, 'active stream preserved');
// The projection markers are accepted as live-assistant evidence.
assert.strictEqual(guardWouldPreserve(withBody, null, burstLiveMessages, true, 's1'), true, '_activityBurstId projection preserved');
assert.strictEqual(guardWouldPreserve(withBody, null, segLiveMessages, true, 's1'), true, '_liveSegmentSeq projection preserved');
// Ownership gates: wrong-session DOM and no-owner are always rejected.
assert.strictEqual(guardWouldPreserve(wrongSess, 'sid', liveMessages, true, 's1'), false, 'wrong-session DOM rejected');
assert.strictEqual(guardWouldPreserve(withBody, 'sid', liveMessages, false, 's1'), false, 'no-owner (no INFLIGHT) rejected');
console.log('OK');
"""
)
out = subprocess.run([node, "-e", script], capture_output=True, text=True)
assert out.returncode == 0, f"node harness failed: {out.stderr}\n{out.stdout}"
assert "OK" in out.stdout
def test_settled_transcript_with_stale_inflight_renders_once(self):
"""#6948 browser-lifecycle equivalent in the node harness: the full
renderMessages decision — capture guard + wipe + rebuild — must yield
exactly ONE assistant row when a contentful live DOM node exists but
the projection is settled (stream ended, INFLIGHT[sid] stale)."""
node = shutil.which("node")
if not node:
import pytest
pytest.skip("node not available")
script = textwrap.dedent(
"""
const assert=require('assert');
// Faithful mirror of the renderMessages preserve decision + the
// re-attach step. Semantics of the real code (static/ui.js):
// - the rebuilt DOM renders settled assistant rows from the
// projection; a live-projection assistant renders as the
// rebuilt live turn;
// - when the guard captured the live node it REPLACES the rebuilt
// live row (segment/whole-turn swap) — never appends on top of a
// row that already exists in the projection;
// - mid-stream the in-progress turn is NOT yet in the projection,
// so the preserved node appends as the live tail (one live row);
// - when the guard did NOT capture (settled + stale INFLIGHT), the
// rebuilt settled rows stand alone — exactly one copy.
function renderMessagesSim(lt, activeStreamId, messages, inflightOwner){
let preservedLiveTurn=null;
const sid='s1';
if(sid&&inflightOwner){
if(lt&&(!lt.dataset||!lt.dataset.sessionId||lt.dataset.sessionId===sid)){
const hasLiveProjection=Array.isArray(messages)&&messages.some(m=>
m&&m.role==='assistant'&&(m._live||m._activityBurstId!==undefined||m._liveSegmentSeq!==undefined)
);
if(activeStreamId || hasLiveProjection){ preservedLiveTurn=lt; }
}
}
const liveMark=m=>m&&m.role==='assistant'&&(m._live||m._activityBurstId!==undefined||m._liveSegmentSeq!==undefined);
const hasLiveRow=Array.isArray(messages)&&messages.some(liveMark);
const finalRows=(messages||[]).filter(m=>m&&m.role==='assistant'&&!liveMark(m));
if(preservedLiveTurn){
// Swap-in replaces the rebuilt live row when one exists; the
// mid-stream case appends the live tail (turn not yet in the
// projection). Either way: one row for the turn.
finalRows.push({role:'assistant',_preserved:true});
}else if(hasLiveRow){
finalRows.push({role:'assistant',_rebuiltLive:true});
}
return finalRows;
}
const settled=[{role:'user',content:'hi'},{role:'assistant',content:'final answer'}];
const midStreamMessages=[{role:'user',content:'hi'}]; // turn not yet persisted
const liveTail=[{role:'user',content:'hi'},{role:'assistant',content:'partial',_live:true}];
const contentfulLiveNode={dataset:{sessionId:'s1'},querySelector:()=>({})};
// The bug (#6948): stale INFLIGHT + settled projection + contentful
// dead node → TWO rows before the fix; exactly ONE after.
assert.strictEqual(
renderMessagesSim(contentfulLiveNode, null, settled, true).length, 1,
'settled transcript must render exactly once despite stale INFLIGHT + contentful DOM (#6948)'
);
// Mid-stream: active stream keeps the live node attached (one live row).
assert.strictEqual(
renderMessagesSim(contentfulLiveNode, 'sid', midStreamMessages, true).length, 1,
'mid-stream render must keep exactly one live row'
);
// Reconnect: explicit live projection keeps the live node (one row).
assert.strictEqual(
renderMessagesSim(contentfulLiveNode, null, liveTail, true).length, 1,
'reconnect with live projection must keep exactly one live row'
);
// Rapid turn completion: a second settled turn cannot be duplicated.
const twoTurns=[
{role:'user',content:'a'},{role:'assistant',content:'answer one'},
{role:'user',content:'b'},{role:'assistant',content:'answer two'},
];
assert.strictEqual(
renderMessagesSim(contentfulLiveNode, null, twoTurns, true).length, 2,
'rapid turn completion must not duplicate settled assistant rows'
);
console.log('OK');
"""
)
+84
View File
@@ -197,3 +197,87 @@ class TestBootstrapStructure:
assert "_load_repo_dotenv" in bootstrap_src, (
"bootstrap.py must load .env so direct invocation matches start.sh behaviour"
)
class TestLeakedWebuiPasswordIsolation:
"""#7168 review: a local repo .env containing HERMES_WEBUI_PASSWORD leaks
into os.environ when any test imports bootstrap (import-time
_load_repo_dotenv() runs OUTSIDE monkeypatch's undo scope). The conftest
autouse guard strips the leaked var around every test so later tests
don't see is_auth_enabled()==True (the #5588 failure shape)."""
def _load_guard(self):
import importlib.util
cpath = Path(__file__).resolve().parent / "conftest.py"
spec = importlib.util.spec_from_file_location("_test_conftest", cpath)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod._strip_leaked_webui_password_env
def test_conftest_guard_strips_leaked_password(self):
strip = self._load_guard()
os.environ["HERMES_WEBUI_PASSWORD"] = "leaked-from-repo-dotenv"
try:
strip()
assert "HERMES_WEBUI_PASSWORD" not in os.environ
finally:
os.environ.pop("HERMES_WEBUI_PASSWORD", None)
def test_conftest_guard_preserves_intentional_empty_override(self):
"""ctl.sh-style override semantics: an explicitly empty value means
'keep auth off' and must survive the strip."""
strip = self._load_guard()
sentinel = object()
os.environ["HERMES_WEBUI_PASSWORD"] = ""
try:
strip()
assert os.environ.get("HERMES_WEBUI_PASSWORD", sentinel) == ""
finally:
os.environ.pop("HERMES_WEBUI_PASSWORD", None)
class TestLeakedHermesCommandIsolation:
"""#7168 re-gate round 7: a local repo .env carrying HERMES_COMMAND leaks
into os.environ via bootstrap import-time _load_repo_dotenv() and
redirects gateway_restart._resolve_hermes_command() away from its mocked
shutil.which result — every later active-profile-restart test then fails
on a machine-specific CLI path. The autouse conftest guard strips the
leaked var around every test; upstream code never reads HERMES_COMMAND,
so stripping restores exact upstream semantics."""
def _load_guard(self):
import importlib.util
cpath = Path(__file__).resolve().parent / "conftest.py"
spec = importlib.util.spec_from_file_location("_test_conftest_hc", cpath)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod._strip_leaked_webui_password_env
def test_conftest_guard_strips_leaked_hermes_command(self):
strip = self._load_guard()
os.environ["HERMES_COMMAND"] = "/machine/local/hermes-gateway-wrapper"
try:
strip()
assert "HERMES_COMMAND" not in os.environ
finally:
os.environ.pop("HERMES_COMMAND", None)
def test_conftest_guard_still_strips_password_when_command_leaks(self):
"""Both leaked vars are stripped in one pass (password branch must
not be short-circuited by the HERMES_COMMAND handling)."""
strip = self._load_guard()
os.environ["HERMES_WEBUI_PASSWORD"] = "leaked-from-repo-dotenv"
os.environ["HERMES_COMMAND"] = "/machine/local/hermes-gateway-wrapper"
try:
strip()
assert "HERMES_WEBUI_PASSWORD" not in os.environ
assert "HERMES_COMMAND" not in os.environ
finally:
os.environ.pop("HERMES_WEBUI_PASSWORD", None)
os.environ.pop("HERMES_COMMAND", None)
+178
View File
@@ -0,0 +1,178 @@
"""A cancelling run is lifecycle-busy but NOT attachable live UI work.
``ACTIVE_RUNS`` tracks worker lifecycle. ``cancel_stream()`` deliberately keeps
the row as ``phase="cancelling"`` while the worker unwinds so a successor turn
cannot start on top of it. The client, however, has already reached a terminal
state for that stream: its run journal ends in a terminal event.
Before this fix the recovery lookups treated every same-session ``ACTIVE_RUNS``
row as attachable, so an idle session whose sidecar had already cleared
``active_stream_id`` still received a recovered ``server_turn_started`` for the
cancelled stream on EVERY ``/api/session/stream`` subscription. The client
attached, replayed the terminal event, tore the renderer down, resubscribed, and
the server replayed the same frame again — an endless attach/replay loop.
These tests pin both directions of the resulting contract:
* browser recovery excludes cancelling rows, while busy checks still keep a
FRESH cancellation busy (a successor must not overlap the unwinding worker);
* a cancelling row past the bounded unwind window with no live ``STREAMS``
channel is reclaimed, so a wedged worker cannot suppress wakeups forever;
* a cancelling row that still owns a live ``STREAMS`` channel is NOT reclaimed
on age alone.
"""
import sys
import time
from pathlib import Path
from types import SimpleNamespace
REPO_ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(REPO_ROOT))
from api import background_process as bp # noqa: E402
from api import config as cfg # noqa: E402
from api.session_ops import _live_active_stream_id # noqa: E402
def _clear(stream_id: str) -> None:
cfg.unregister_active_run(stream_id)
with cfg.STREAMS_LOCK:
cfg.STREAMS.pop(stream_id, None)
def test_running_run_is_still_attachable_and_busy():
"""Control: an ordinary running turn keeps both semantics unchanged."""
sid = "sess-running-control"
stream_id = "stream-running-control"
cfg.register_active_run(stream_id, session_id=sid, phase="running")
try:
assert bp.active_stream_id_for_session(sid) == stream_id
assert bp._session_has_active_turn(sid) is True
assert _live_active_stream_id(
SimpleNamespace(session_id=sid, active_stream_id=stream_id)
) == stream_id
finally:
_clear(stream_id)
def test_fresh_cancelling_run_is_not_attachable_but_stays_busy():
"""The loop-producing case: recovery must not replay a cancelled run.
The same row must still count as busy so a successor cannot start while the
cancelled worker is unwinding.
"""
sid = "sess-cancelling-fresh"
stream_id = "stream-cancelling-fresh"
cfg.register_active_run(
stream_id,
session_id=sid,
phase="cancelling",
cancelled_at=time.time() - 5.0,
)
try:
assert bp.active_stream_id_for_session(sid) is None
assert bp._session_has_active_turn(sid) is True
assert stream_id in cfg.ACTIVE_RUNS
finally:
_clear(stream_id)
def test_hidden_tab_status_does_not_resurrect_a_cancelling_run():
"""/api/session/status must not hand a cancelling stream to the poller."""
sid = "sess-cancelling-status"
stream_id = "stream-cancelling-status"
cfg.register_active_run(stream_id, session_id=sid, phase="cancelling")
try:
assert _live_active_stream_id(
SimpleNamespace(session_id=sid, active_stream_id=stream_id)
) is None
finally:
_clear(stream_id)
def test_hidden_tab_status_ignores_cancelling_run_with_live_stream_channel():
"""A cancelling row is terminal for the client even if STREAMS still holds it."""
sid = "sess-cancelling-status-live-channel"
stream_id = "stream-cancelling-status-live-channel"
cfg.register_active_run(stream_id, session_id=sid, phase="cancelling")
with cfg.STREAMS_LOCK:
cfg.STREAMS[stream_id] = object()
try:
assert _live_active_stream_id(
SimpleNamespace(session_id=sid, active_stream_id=stream_id)
) is None
finally:
_clear(stream_id)
def test_stale_cancelling_orphan_stops_blocking_and_is_reclaimed():
"""Past the unwind window with no live channel, the row is an orphan."""
sid = "sess-cancelling-stale"
stream_id = "stream-cancelling-stale"
cfg.register_active_run(
stream_id,
session_id=sid,
phase="cancelling",
cancelled_at=time.time() - 240.0,
)
cfg.register_stream_owner(stream_id, sid)
try:
assert bp._session_has_active_turn(sid) is False
assert stream_id not in cfg.ACTIVE_RUNS
assert cfg.stream_owner_session_id(stream_id) is None
finally:
_clear(stream_id)
def test_stale_cancelling_run_with_live_channel_is_not_reclaimed():
"""Reverse control: age alone must not reap a row that still owns a channel."""
sid = "sess-cancelling-stale-live-channel"
stream_id = "stream-cancelling-stale-live-channel"
cfg.register_active_run(
stream_id,
session_id=sid,
phase="cancelling",
cancelled_at=time.time() - 240.0,
)
with cfg.STREAMS_LOCK:
cfg.STREAMS[stream_id] = object()
try:
assert bp._session_has_active_turn(sid) is True
assert stream_id in cfg.ACTIVE_RUNS
finally:
_clear(stream_id)
def test_cancel_staleness_anchors_on_cancel_time_not_run_start():
"""A long-running turn cancelled just now must not read as stale."""
entry = {
"session_id": "sess-anchor",
"phase": "cancelling",
"started_at": time.time() - 4000.0,
"cancelled_at": time.time() - 5.0,
}
assert cfg.active_run_cancel_is_stale(entry, grace_seconds=180.0) is False
legacy = {
"session_id": "sess-anchor-legacy",
"phase": "cancelling",
"started_at": time.time() - 4000.0,
}
assert cfg.active_run_cancel_is_stale(legacy, grace_seconds=180.0) is True
running = {
"session_id": "sess-anchor-running",
"phase": "running",
"started_at": time.time() - 4000.0,
}
assert cfg.active_run_cancel_is_stale(running, grace_seconds=180.0) is False
def test_attachability_predicate_edges():
"""Non-dict/opaque entries stay attachable; only cancelling is excluded."""
assert cfg.active_run_is_attachable({"phase": "running"}) is True
assert cfg.active_run_is_attachable({"phase": "cancelling"}) is False
assert cfg.active_run_is_attachable({"phase": " cancelling "}) is False
assert cfg.active_run_is_attachable({}) is True
assert cfg.active_run_is_attachable(object()) is True
+3 -3
View File
@@ -151,9 +151,9 @@ def test_chat_start_no_longer_bare_404_on_keyerror():
def test_get_session_route_uses_shared_synthesiser():
"""The GET KeyError path must also delegate to the same helper."""
src = ROUTES_PY.read_text(encoding="utf-8")
# Find the /api/session GET block (not /api/sessions).
# Find the /api/session GET handler (extracted to _handle_session_get).
block = re.search(
r'if parsed\.path == "/api/session":.*?return j\(handler, \{"session": public_session_projection\(sess\)\}\)',
r'def _handle_session_get\(.*?return j\(handler, \{"session": public_session_projection\(sess\)\}\)',
src,
re.DOTALL,
)
@@ -185,7 +185,7 @@ def test_get_session_preserves_cli_read_only_flag():
sets to True for BOTH the explicit AND the source-refused
refusal paths."""
block = re.search(
r'if parsed\.path == "/api/session":.*?return j\(handler, \{"session": public_session_projection\(sess\)\}\)\)?',
r'def _handle_session_get\(.*?return j\(handler, \{"session": public_session_projection\(sess\)\}\)\)?',
ROUTES_PY.read_text(encoding="utf-8"),
re.DOTALL,
)
@@ -0,0 +1,138 @@
"""Regression coverage: the periodic streaming checkpoint must not rewrite
byte-identical sidecars.
The checkpoint thread (#765) fires whenever a tool call completes. A completed
tool call is *not* proof that anything the checkpoint persists has changed:
``run_conversation()`` mutates an internal copy of the transcript, and the turn
bookkeeping fields are written once before the run starts.
Each rewrite re-serializes the whole session while holding the GIL. Because the
HTTP server and the agent workers share one interpreter, that cost is added
latency for every concurrent request — ~100 ms for a 1.8 MB sidecar and ~300 ms
for a 36 MB one, repeated every 15 s for the whole turn.
These tests pin the behavior at the decision level (the fingerprint) so they do
not depend on wall-clock timing of the background thread.
"""
from __future__ import annotations
def _session(tmp_path, **overrides):
from api.models import Session
s = Session(session_id="ckpt_redundant_1")
s.messages = [{"role": "assistant", "content": "hello"}]
s.context_messages = []
s.pending_user_message = "do the thing"
s.pending_started_at = 1787000000.0
s.active_stream_id = "streamabc"
for key, value in overrides.items():
setattr(s, key, value)
return s
def test_fingerprint_is_stable_when_nothing_changes(tmp_path):
"""Two reads of an untouched session produce the same fingerprint.
This is the redundant case: the checkpoint would rewrite identical bytes.
"""
from api.streaming import _streaming_checkpoint_fingerprint
s = _session(tmp_path)
assert _streaming_checkpoint_fingerprint(s) == _streaming_checkpoint_fingerprint(s)
def test_fingerprint_tracks_every_field_the_checkpoint_persists(tmp_path):
"""Any state a mid-run checkpoint can legitimately advance must be detected.
Missing one of these would make the skip lose real data, so each field is
asserted individually rather than as a group.
"""
from api.streaming import _streaming_checkpoint_fingerprint
base = _session(tmp_path)
reference = _streaming_checkpoint_fingerprint(base)
mutations = {
# compression rotates the session id mid-turn
"session_id": "rotated_sid",
"parent_session_id": "parent_sid",
"profile": "other-profile",
# turn bookkeeping cleared at the end of the turn
"active_stream_id": None,
"pending_user_message": None,
"pending_started_at": None,
"pending_user_source": "telegram",
# transcript growth (error rows, continuation tail, compression)
"messages": [{"role": "assistant", "content": "hello"}, {"role": "user", "content": "more"}],
"context_messages": [{"role": "user", "content": "ctx"}],
"compression_anchor_visible_idx": 3,
"pre_compression_snapshot": {"messages": []},
"pending_attachments": [{"name": "f.png"}],
}
for field, value in mutations.items():
mutated = _session(tmp_path, **{field: value})
assert _streaming_checkpoint_fingerprint(mutated) != reference, (
f"a change to {field!r} must force a checkpoint write"
)
def test_fingerprint_fails_closed_on_unreadable_session():
"""An unreadable session yields None so the caller keeps writing.
Fail-closed: never skip a checkpoint on uncertainty.
"""
from api.streaming import _streaming_checkpoint_fingerprint
class Hostile:
def __getattr__(self, name):
raise RuntimeError("boom")
assert _streaming_checkpoint_fingerprint(Hostile()) is None
def test_checkpoint_loop_skips_redundant_writes_but_persists_changes(monkeypatch, tmp_path):
"""End-to-end on the real loop body: identical state writes once.
Drives the actual decision sequence used by the checkpoint thread: an
unchanged session must be persisted once, not on every tool completion,
while a real mutation must be persisted promptly.
"""
import time
from api.streaming import (
_streaming_checkpoint_fingerprint,
_CHECKPOINT_IDLE_REFRESH_SECONDS,
)
s = _session(tmp_path)
writes = []
# Mirror of the loop body in _periodic_checkpoint(), kept in sync with it.
last_fingerprint = None
last_write_at = 0.0
now = time.time()
def tick(session, at):
nonlocal last_fingerprint, last_write_at
fingerprint = _streaming_checkpoint_fingerprint(session)
stale = (at - last_write_at) >= _CHECKPOINT_IDLE_REFRESH_SECONDS
if fingerprint is None or fingerprint != last_fingerprint or stale:
writes.append(at)
last_fingerprint = fingerprint
last_write_at = at
# 10 tool completions, nothing else changed -> exactly one write.
for i in range(10):
tick(s, now + i)
assert len(writes) == 1, f"redundant rewrites: {len(writes)} writes for an unchanged session"
# A real transcript change must be persisted on the next tick.
s.messages = list(s.messages) + [{"role": "assistant", "content": "new answer"}]
tick(s, now + 10)
assert len(writes) == 2, "a real mutation must force a checkpoint write"
# updated_at must not go stale forever on a very long quiet turn.
tick(s, now + 10 + _CHECKPOINT_IDLE_REFRESH_SECONDS)
assert len(writes) == 3, "a periodic refresh must still happen on long turns"
+8 -8
View File
@@ -164,15 +164,15 @@ def test_get_cli_sessions_cache_invalidates_when_sqlite_wal_changes(monkeypatch,
Path(f"{db_path}-wal").write_text("new wal contents", encoding="utf-8")
second = models.get_cli_sessions()
# Two calls to get_cli_sessions() × 3 invocations each (first pass +
# cron-only pass + webhook-only pass) = 6 total calls to the mock.
assert calls == 6
# Two calls to get_cli_sessions() × 4 invocations each (first pass +
# cron-only + webhook-only + kanban-only passes) = 8 total mock calls.
assert calls == 8
# First pass of first call returned message_count=1 (calls was 1).
assert first[0]["message_count"] == 1
# First pass of second call returned message_count=4 (calls was 4;
# source-specific passes incremented calls to 2, 3, 5, and 6 but excluded
# the cli-source session from those pass results).
assert second[0]["message_count"] == 4
# First pass of second call returned message_count=5 (calls was 5;
# source-specific passes incremented the other counters but excluded the
# cli-source session from those pass results).
assert second[0]["message_count"] == 5
def test_session_import_cli_returns_read_only_claude_code_payload(monkeypatch, tmp_path):
@@ -200,7 +200,7 @@ def test_session_import_cli_returns_read_only_claude_code_payload(monkeypatch, t
monkeypatch.setattr(routes, "j", lambda _handler, payload, status=200, extra_headers=None: payload)
monkeypatch.setattr(routes, "get_cli_session_messages", lambda _sid, profile=None: messages if _sid == sid else [])
monkeypatch.setattr(routes, "get_cli_sessions", lambda source_filter=None, all_profiles=False: [meta])
monkeypatch.setattr(routes, "get_last_workspace", lambda: tmp_path / "workspace")
monkeypatch.setattr(routes, "get_last_workspace", lambda profile=None: tmp_path / "workspace")
monkeypatch.setattr(routes, "import_cli_session", lambda *args, **kwargs: (_ for _ in ()).throw(AssertionError("read-only import must not persist")))
response = routes._handle_session_import_cli(object(), {"session_id": sid})
+64
View File
@@ -0,0 +1,64 @@
"""Smoke tests for the packaged ``hermes-webui`` CLI entry point (#6739).
The console script is declared in ``pyproject.toml`` under
``[project.scripts]`` and must keep resolving to ``bootstrap.main`` — the
bootstrap entry point runs launcher-python discovery and dep-ensure before
launching the server, matching the effective behavior of
``python server.py`` / ``python -m server``. These tests
guard the wiring without booting a real server:
1. The console script is declared in the packaging metadata.
2. The declared target (``bootstrap:main``) actually resolves to a callable.
3. When the package is installed in the test environment (e.g. CI editable
install), the real ``importlib.metadata`` entry point resolves end-to-end.
"""
from __future__ import annotations
import importlib
import importlib.metadata
import os
import tomllib
import pytest
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
EXPECTED_SCRIPT = "hermes-webui"
EXPECTED_TARGET = "bootstrap:main"
def _pyproject_scripts() -> dict:
"""Parse the [project.scripts] table from pyproject.toml."""
with open(os.path.join(REPO_ROOT, "pyproject.toml"), "rb") as fh:
pyproject = tomllib.load(fh)
return dict(pyproject.get("project", {}).get("scripts", {}))
def test_console_script_declared_in_pyproject():
"""The packaged CLI entry point is declared and maps to bootstrap.main."""
scripts = _pyproject_scripts()
assert scripts.get(EXPECTED_SCRIPT) == EXPECTED_TARGET
def test_console_script_target_resolves_to_callable():
"""The declared module:attr target imports and is callable."""
module_name, _, attr = EXPECTED_TARGET.partition(":")
module = importlib.import_module(module_name)
target = getattr(module, attr)
assert callable(target)
def _installed_entry_point():
"""Return the installed hermes-webui console-script EntryPoint, or None."""
selected = importlib.metadata.entry_points().select(
group="console_scripts", name=EXPECTED_SCRIPT
)
return selected[0] if selected else None
def test_installed_entry_point_wiring():
"""When installed, the console script must resolve to bootstrap.main."""
ep = _installed_entry_point()
if ep is None:
pytest.skip("hermes-webui not installed in this test environment")
assert ep.value == EXPECTED_TARGET
assert callable(ep.load())
@@ -0,0 +1,196 @@
"""``_strip_compact_echo_suffix`` must find the cut point in one pass.
The reasoning/visible echo stripper is called up to four times per interim
assistant message. The original implementation probed every candidate cut
index and re-folded the whole remaining tail for each probe, so the cost grew
with ``window x suffix`` instead of with the data actually inspected. On a
6000-character conclusion block that is seconds of pure CPU, held under the
GIL, which stalls every other stream in the process.
These tests pin both halves of the contract:
* the fast path must return byte-identical results to a straightforward
reference implementation, including the leftmost-cut tie-break; and
* it must not re-fold the search window once per candidate cut index.
"""
import random
import re
import string
import pytest
def _reference_strip_compact_echo_suffix(value, suffix, *, search_window: int = 4096):
"""Deliberately naive oracle: probe every cut index, fold the whole tail.
This mirrors the original behaviour exactly and exists only so the
optimised implementation can be proven equivalent to it.
"""
def compact(v):
return re.sub(r'\s+', '', str(v or ''))
raw = str(value or '')
candidate = compact(suffix)
if not raw or not candidate:
return raw, False
tail = raw[-max(len(str(suffix or '')) * 3, search_window):]
offset = len(raw) - len(tail)
for idx in range(len(tail) + 1):
if compact(tail[idx:]) == candidate:
return raw[: offset + idx].rstrip(), True
return raw, False
def _cases():
cases = [
# Exact echo at the end of the buffer.
("bla bla bla une conclusion", "une conclusion"),
# Echo that differs only by whitespace shape.
("bla bla une conclusion", "une conclusion"),
("bla bla\nune\tconclusion", "une conclusion"),
# No echo at all.
("texte quelconque sans rapport", "une conclusion"),
# Partial echo must NOT match.
("bla une conclu", "une conclusion"),
# The buffer is exactly the echo.
("une conclusion", "une conclusion"),
# Empty / None inputs.
("", "une conclusion"),
("bla bla", ""),
("", ""),
(None, "x"),
("x", None),
# Trailing and interior whitespace around the cut point.
("bla bla une conclusion ", "une conclusion"),
("bla bla \n une conclusion", "une conclusion"),
# Unicode, accents, emoji.
("préambule ✅ conclusion émise", "conclusion émise"),
("texte 🟢 Réponse", "🟢 Réponse"),
# Whitespace-only suffix folds to nothing.
("bla bla", " \n "),
# Repetitions: the leftmost cut point is the contractual one.
("abc abc abc", "abc"),
("abc abc abc", "abc abc"),
("aaaa", "aa"),
# Long buffer with the echo at the very end.
("remplissage " * 500 + "la vraie conclusion", "la vraie conclusion"),
# Echo whose whitespace-padded span is wider than the search window:
# the window only exposes ``conclusion``, so nothing may be stripped.
("bla une" + " " * 5000 + "conclusion", "une conclusion"),
# Exotic whitespace that ``\s`` and ``str.isspace`` must treat alike.
("bla\u00a0bla une conclusion", "une\u00a0conclusion"),
("bla\u2028une conclusion", "une conclusion"),
]
rnd = random.Random(20260825)
alphabet = string.ascii_letters + " \n\t" + "éà✅\u00a0"
for _ in range(2000):
base = "".join(rnd.choice(alphabet) for _ in range(rnd.randint(0, 120)))
if rnd.random() < 0.5 and len(base) > 4:
k = rnd.randint(1, max(1, len(base) // 2))
suffix = base[-k:]
else:
suffix = "".join(rnd.choice(alphabet) for _ in range(rnd.randint(0, 30)))
cases.append((base, suffix))
return cases
def test_strip_compact_echo_suffix_matches_reference_oracle():
"""The optimised scan must be indistinguishable from the naive probe."""
import api.streaming as streaming
divergences = []
for value, suffix in _cases():
expected = _reference_strip_compact_echo_suffix(value, suffix)
actual = streaming._strip_compact_echo_suffix(value, suffix)
if actual != expected:
divergences.append((value, suffix, expected, actual))
assert not divergences[:5], (
f"{len(divergences)} divergence(s) from the reference implementation; "
f"first: {divergences[:1]}"
)
def test_strip_compact_echo_suffix_honours_a_custom_search_window():
"""The bounded window must keep its meaning on the fast path."""
import api.streaming as streaming
buffer_text = "z" * 400 + " la conclusion"
for window in (16, 64, 512, 4096):
assert streaming._strip_compact_echo_suffix(
buffer_text, "la conclusion", search_window=window
) == _reference_strip_compact_echo_suffix(
buffer_text, "la conclusion", search_window=window
), f"window={window}"
def test_strip_compact_echo_suffix_rejects_an_echo_wider_than_the_window():
"""An echo that only fits inside a wider window must not be found.
The buffer ends with the echo, but interior whitespace stretches its raw
span past the default window, so the folded window only exposes the last
word. Widening the window to cover the whole span makes the same echo
visible again, which proves the non-match is caused by the window and
not by the text.
"""
import api.streaming as streaming
buffer_text = "bla une" + " " * 5000 + "conclusion"
suffix = "une conclusion"
for impl in (streaming._strip_compact_echo_suffix, _reference_strip_compact_echo_suffix):
assert impl(buffer_text, suffix) == (buffer_text, False), impl.__name__
assert impl(buffer_text, suffix, search_window=len(buffer_text)) == ("bla", True), impl.__name__
def test_strip_compact_echo_suffix_does_not_refold_the_window_per_cut(monkeypatch):
"""One pass, not one fold per candidate cut index.
The original implementation called the whitespace-folding helper once per
possible cut position — roughly 4097 times for a default window — and each
of those calls re-scanned the remaining tail. Counting the calls pins the
algorithmic property directly, without depending on wall-clock timing.
"""
import api.streaming as streaming
calls = {'n': 0}
original = streaming._compact_for_echo_compare
def counting(value):
calls['n'] += 1
return original(value)
monkeypatch.setattr(streaming, '_compact_for_echo_compare', counting)
reasoning_buffer = "Analyse de la situation en cours. " * 1500
conclusion = "x" * 6000
streaming._strip_compact_echo_suffix(reasoning_buffer, conclusion)
assert calls['n'] <= 8, (
f"whitespace folding ran {calls['n']} times for a single call; the scan "
"is re-folding the window once per candidate cut index"
)
@pytest.mark.timeout(60)
def test_strip_compact_echo_suffix_stays_cheap_on_a_long_conclusion():
"""A long final message must not cost seconds of GIL-held CPU."""
import time
import api.streaming as streaming
reasoning_buffer = "Analyse de la situation en cours. " * 1500
conclusion = "x" * 6000
# Warm up so the first-call import/regex cache is not measured.
streaming._strip_compact_echo_suffix(reasoning_buffer, conclusion)
start = time.perf_counter()
for _ in range(5):
streaming._strip_compact_echo_suffix(reasoning_buffer, conclusion)
elapsed = (time.perf_counter() - start) / 5
# The original implementation needs ~7s here. A very loose budget keeps the
# test meaningful on a loaded CI box while still failing the quadratic scan
# by three orders of magnitude.
assert elapsed < 0.5, f"single call took {elapsed:.3f}s"
+380
View File
@@ -0,0 +1,380 @@
"""Behavioural DOM test for the composer-footer fit freeze (PR #7275).
`_fitComposerFooter()` resolves the compact stage of `.composer-footer` by
stripping the `.cf-icons`/`.cf-burger` classes, measuring the left cluster's
overflow, then re-adding the classes it needs. Between strip and restore the
footer is laid out at full width: the composer grows a few px and `#messages`
loses the same amount of `clientHeight`, then gets it back. Because the fit
pass runs on every context-indicator update during SSE streaming, a pinned
reader sees that as a vertical jitter of the whole transcript.
The fix pins the footer's border box (inline `height` + `visibility:hidden`)
for the duration of the probe and releases it in the same task, after the
resolved stage classes are back. This test drives the ACTUAL function from
static/ui.js via node against a small layout model in which the stage classes
dictate the footer's natural height and the left cluster's content width.
Every class or style mutation and every overflow measurement commits a layout
sample, so the recorded sequence of footer heights / messages client heights
is exactly what a browser would have painted.
Covered from each starting stage (full, icons, burger) and for each forced
overflow outcome (full, icons, burger), with empty and caller-owned prior
inline styles:
* the resolved stage classes are correct;
* the footer border-box height and the messages client height never move
while the probe runs (no intermediate geometry, no resize notification);
* a fit pass that lands on the stage it started from does not move the
footer at all (the streaming steady state);
* the prior inline `height`/`visibility` are restored verbatim;
* a zero-height footer skips the freeze without touching inline styles;
* an exception during measurement still releases the frozen box.
A control run with the freeze made ineffective proves the harness reports the
original jitter, so the test cannot pass vacuously.
"""
import json
import shutil
import subprocess
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).parent.parent.resolve()
UI_JS_PATH = REPO_ROOT / "static" / "ui.js"
NODE = shutil.which("node")
pytestmark = pytest.mark.skipif(NODE is None, reason="node not on PATH")
STAGES = ("full", "icons", "burger")
STAGE_CLASSES = {"full": "", "icons": "cf-icons", "burger": "cf-burger cf-icons"}
_DRIVER_SRC = r"""
const fs = require('fs');
const src = fs.readFileSync(process.argv[2], 'utf8');
function extractFunc(name) {
const re = new RegExp('function\\s+' + name + '\\s*\\(');
const start = src.search(re);
if (start < 0) throw new Error(name + ' not found');
let i = src.indexOf('{', start); let depth = 1; i++;
while (depth > 0 && i < src.length) {
if (src[i] === '{') depth++; else if (src[i] === '}') depth--; i++;
}
return src.slice(start, i);
}
// ── Layout model ───────────────────────────────────────────────────────────
// Stage classes decide the footer's natural border-box height and the left
// cluster's content width. The left cluster's clientWidth is the width the
// scenario makes available; scrollWidth is content width clamped to it, so
// overflow (scrollWidth > clientWidth + 1) depends on the CURRENT stage.
const STAGE_HEIGHT = { full: 56, icons: 48, burger: 40 };
const STAGE_LEFT_WIDTH = { full: 900, icons: 600, burger: 300 };
const OUTCOME_WIDTH = { full: 1000, icons: 700, burger: 400 };
const STAGE_CLASSES = { full: [], icons: ['cf-icons'], burger: ['cf-icons', 'cf-burger'] };
const VIEWPORT_HEIGHT = 800;
function stageOf(classes) {
return classes.has('cf-burger') ? 'burger' : classes.has('cf-icons') ? 'icons' : 'full';
}
function makeFooterDom(opts) {
const classes = new Set(STAGE_CLASSES[opts.start]);
const store = { height: opts.prevHeight, visibility: opts.prevVisibility };
const samples = [];
const styleWrites = [];
// Border box: an inline height wins (box-sizing:border-box), otherwise the
// stage's natural height. `ignoreInlineStyles` is the control mode where
// the freeze has no layout effect, i.e. the pre-fix behaviour.
function borderBoxHeight() {
if (opts.zeroHeight) return 0;
const inline = parseFloat(store.height);
if (!opts.ignoreInlineStyles && Number.isFinite(inline)) return inline;
return STAGE_HEIGHT[stageOf(classes)];
}
function hidden() {
return !opts.ignoreInlineStyles && store.visibility === 'hidden';
}
// One layout sample = what the screen would commit for the current state.
function layout(reason) {
const h = borderBoxHeight();
samples.push({
reason, stage: stageOf(classes), hidden: hidden(),
footerHeight: h, messagesClientHeight: VIEWPORT_HEIGHT - h,
});
}
const style = {};
for (const prop of ['height', 'visibility']) {
Object.defineProperty(style, prop, {
enumerable: true,
get() { return store[prop]; },
set(v) {
store[prop] = String(v);
styleWrites.push(prop + '=' + JSON.stringify(String(v)));
layout('style.' + prop);
},
});
}
const classList = {
add() { for (const n of arguments) classes.add(n); layout('classList.add'); },
remove() { for (const n of arguments) classes.delete(n); layout('classList.remove'); },
toggle(name, force) {
const on = force === undefined ? !classes.has(name) : !!force;
if (on) classes.add(name); else classes.delete(name);
layout('classList.toggle');
return on;
},
contains(name) { return classes.has(name); },
};
let throwOnMeasure = !!opts.throwOnMeasure;
const left = {
get clientWidth() { return opts.availableWidth; },
get scrollWidth() {
layout('measure');
if (throwOnMeasure) { throwOnMeasure = false; throw new Error('synthetic measurement failure'); }
return Math.max(opts.availableWidth, STAGE_LEFT_WIDTH[stageOf(classes)]);
},
};
const footer = {
style, classList,
querySelector(sel) { return sel === '.composer-left' ? left : null; },
getBoundingClientRect() { return { left: 0, top: 0, width: 0, height: borderBoxHeight() }; },
};
const document = {
querySelector(sel) { return sel === '.composer-footer' ? footer : null; },
};
return {
document, samples, styleWrites, store,
snapshot() { layout('snapshot'); return samples[samples.length - 1]; },
classes() { return Array.from(classes).sort().join(' '); },
stage() { return stageOf(classes); },
};
}
function dedupe(arr) { return arr.filter((v, i) => i === 0 || v !== arr[i - 1]); }
function runFit(fit, opts) {
const dom = makeFooterDom(opts);
const before = dom.snapshot();
global.document = dom.document;
let error = null;
try { fit(); } catch (e) { error = String((e && e.message) || e); }
const after = dom.snapshot();
// Samples committed by class mutations and overflow measurements: every one
// of them happens inside the probe window and must be frozen + hidden.
const probe = dom.samples.filter(s => s.reason.startsWith('classList') || s.reason === 'measure');
return {
start: opts.start, outcome: opts.outcome, availableWidth: opts.availableWidth,
prevHeight: opts.prevHeight, prevVisibility: opts.prevVisibility,
startHeight: before.footerHeight, finalHeight: after.footerHeight,
startMessagesHeight: before.messagesClientHeight, finalMessagesHeight: after.messagesClientHeight,
finalStage: dom.stage(), finalClasses: dom.classes(),
styleHeight: dom.store.height, styleVisibility: dom.store.visibility,
heights: dedupe(dom.samples.map(s => s.footerHeight)),
messagesHeights: dedupe(dom.samples.map(s => s.messagesClientHeight)),
paintedStages: dedupe(dom.samples.filter(s => !s.hidden).map(s => s.stage)),
hiddenAtEnd: after.hidden,
probeSamples: probe.length,
probeHeights: dedupe(probe.map(s => s.footerHeight)),
probeAllHidden: probe.length > 0 && probe.every(s => s.hidden),
styleWrites: dom.styleWrites,
error,
};
}
eval(extractFunc('_fitComposerFooter'));
const PREV_STYLES = [
{ prevHeight: '', prevVisibility: '' },
{ prevHeight: '52px', prevVisibility: 'visible' },
];
const result = { runs: [], control_unfrozen: [] };
for (const start of Object.keys(STAGE_CLASSES)) {
for (const outcome of Object.keys(OUTCOME_WIDTH)) {
for (const prev of PREV_STYLES) {
const opts = Object.assign({ start, outcome, availableWidth: OUTCOME_WIDTH[outcome] }, prev);
result.runs.push(runFit(_fitComposerFooter, opts));
result.control_unfrozen.push(runFit(_fitComposerFooter, Object.assign({ ignoreInlineStyles: true }, opts)));
}
}
}
// A footer that currently measures 0px (e.g. hidden ancestor) must still
// resolve its stage but must not be pinned or hidden by the fit pass.
result.zero_height = runFit(_fitComposerFooter, {
start: 'icons', outcome: 'burger', availableWidth: OUTCOME_WIDTH.burger,
prevHeight: '', prevVisibility: '', zeroHeight: true,
});
// An exception inside an overflow measurement must not leave the footer
// hidden or height-pinned.
result.throw_case = runFit(_fitComposerFooter, {
start: 'icons', outcome: 'icons', availableWidth: OUTCOME_WIDTH.icons,
prevHeight: '', prevVisibility: '', throwOnMeasure: true,
});
process.stdout.write(JSON.stringify(result));
"""
@pytest.fixture(scope="module")
def driver_path(tmp_path_factory):
p = tmp_path_factory.mktemp("composer_footer_fit_driver") / "driver.js"
p.write_text(_DRIVER_SRC, encoding="utf-8")
return str(p)
@pytest.fixture(scope="module")
def outcome(driver_path):
result = subprocess.run(
[NODE, driver_path, str(UI_JS_PATH)],
capture_output=True, text=True, timeout=30,
)
if result.returncode != 0:
raise RuntimeError(f"node driver failed: {result.stderr}")
return json.loads(result.stdout)
def _label(run):
return (
f"start={run['start']} outcome={run['outcome']} "
f"prev=(height={run['prevHeight']!r}, visibility={run['prevVisibility']!r})"
)
def test_matrix_covers_every_start_stage_and_outcome(outcome):
seen = {(r["start"], r["outcome"], r["prevHeight"]) for r in outcome["runs"]}
assert len(outcome["runs"]) == len(STAGES) * len(STAGES) * 2
for start in STAGES:
for final in STAGES:
for prev in ("", "52px"):
assert (start, final, prev) in seen
def test_fit_resolves_expected_stage_from_every_start(outcome):
"""The adaptive ladder is unchanged: the available width alone decides the
final stage, whatever stage the footer started from."""
for run in outcome["runs"]:
assert run["error"] is None, f"{_label(run)}: {run['error']}"
assert run["finalClasses"] == STAGE_CLASSES[run["outcome"]], (
f"{_label(run)}: expected {STAGE_CLASSES[run['outcome']]!r}, "
f"got {run['finalClasses']!r}"
)
def test_footer_border_box_frozen_through_probe(outcome):
"""Every class mutation and every overflow measurement of the ladder must
happen while the footer is hidden and its border-box height is pinned at
the pre-probe value: the expanded intermediate geometry is never committed
to the screen."""
for run in outcome["runs"]:
assert run["probeSamples"] >= 3, f"{_label(run)}: probe produced too few layout samples"
assert run["probeHeights"] == [run["startHeight"]], (
f"{_label(run)}: footer height moved during the probe: "
f"{run['probeHeights']} (start {run['startHeight']})"
)
assert run["probeAllHidden"], f"{_label(run)}: probe ran with a visible footer"
def test_messages_client_height_never_jitters(outcome):
"""The messages viewport may change at most once — directly from the start
geometry to the resolved stage's geometry — and must not change at all when
the fit pass lands on the stage it started from (the SSE steady state)."""
for run in outcome["runs"]:
heights = run["messagesHeights"]
assert heights[0] == run["startMessagesHeight"]
assert heights[-1] == run["finalMessagesHeight"]
assert len(heights) <= 2, (
f"{_label(run)}: messages clientHeight oscillated: {heights}"
)
if run["start"] == run["outcome"]:
assert heights == [run["startMessagesHeight"]], (
f"{_label(run)}: a fit pass that keeps the current stage must not "
f"resize the messages viewport: {heights}"
)
assert len(run["heights"]) <= 2, f"{_label(run)}: footer height oscillated: {run['heights']}"
def test_intermediate_stage_never_painted(outcome):
"""Only the start stage and the resolved stage may ever be visible."""
for run in outcome["runs"]:
allowed = {run["start"], run["outcome"]}
painted = run["paintedStages"]
assert set(painted) <= allowed, (
f"{_label(run)}: intermediate stage painted: {painted}"
)
assert len(painted) <= 2, f"{_label(run)}: painted stages oscillated: {painted}"
assert not run["hiddenAtEnd"], f"{_label(run)}: footer left hidden"
def test_prior_inline_styles_restored_verbatim(outcome):
"""Both empty and caller-owned inline height/visibility must come back
exactly as they were, and a caller-owned inline height must keep pinning
the border box after the pass."""
for run in outcome["runs"]:
assert run["styleHeight"] == run["prevHeight"], (
f"{_label(run)}: inline height not restored: {run['styleHeight']!r}"
)
assert run["styleVisibility"] == run["prevVisibility"], (
f"{_label(run)}: inline visibility not restored: {run['styleVisibility']!r}"
)
if run["prevHeight"] == "52px":
assert run["heights"] == [52], (
f"{_label(run)}: caller-owned inline height must pin the box throughout: {run['heights']}"
)
else:
# Freeze + release: the box is pinned and hidden during the probe
# and both styles are written back, in that order.
assert run["styleWrites"][:2] == [
f'height="{run["startHeight"]}px"', 'visibility="hidden"',
], f"{_label(run)}: unexpected freeze writes {run['styleWrites']}"
assert run["styleWrites"][-2:] == ['height=""', 'visibility=""'], (
f"{_label(run)}: unexpected release writes {run['styleWrites']}"
)
def test_zero_height_footer_skips_freeze(outcome):
run = outcome["zero_height"]
assert run["error"] is None
assert run["finalClasses"] == STAGE_CLASSES["burger"]
assert run["styleWrites"] == [], (
f"a 0px footer must not be pinned or hidden: {run['styleWrites']}"
)
assert run["styleHeight"] == "" and run["styleVisibility"] == ""
def test_measurement_exception_releases_frozen_box(outcome):
run = outcome["throw_case"]
assert run["error"] == "synthetic measurement failure", run["error"]
assert run["styleHeight"] == "" and run["styleVisibility"] == "", (
f"an exception during measurement left inline styles behind: "
f"height={run['styleHeight']!r} visibility={run['styleVisibility']!r}"
)
assert not run["hiddenAtEnd"]
assert run["heights"][-1] == run["finalHeight"]
def test_harness_detects_unfrozen_probe(outcome):
"""Control: with the inline freeze made ineffective (the pre-fix layout
behaviour), the same harness must report the transcript jitter — the
footer and messages heights bounce through the full-width stage whenever
the pass starts from a compact stage. Guarantees the assertions above are
not vacuous."""
jitter = [
r for r in outcome["control_unfrozen"]
if r["prevHeight"] == "" and r["start"] != "full"
and (len(r["heights"]) > 2 or len(r["messagesHeights"]) > 2
or r["probeHeights"] != [r["startHeight"]] or not r["probeAllHidden"])
]
steady = [
r for r in outcome["control_unfrozen"]
if r["prevHeight"] == "" and r["start"] != "full" and r["start"] == r["outcome"]
]
assert steady and all(len(r["messagesHeights"]) > 2 for r in steady), (
"control: an unfrozen steady-state pass from a compact stage must oscillate "
f"the messages viewport: {[r['messagesHeights'] for r in steady]}"
)
assert len(jitter) >= len(steady), [r["heights"] for r in outcome["control_unfrozen"]]
+147
View File
@@ -0,0 +1,147 @@
import json
import re
import subprocess
import textwrap
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
UI_JS = (ROOT / "static" / "ui.js").read_text(encoding="utf-8")
MESSAGES_JS = (ROOT / "static" / "messages.js").read_text(encoding="utf-8")
def _function_source(source: str, name: str) -> str:
start = source.index(f"function {name}(")
brace = source.index("{", start)
depth = 0
for index in range(brace, len(source)):
if source[index] == "{":
depth += 1
elif source[index] == "}":
depth -= 1
if depth == 0:
return source[start : index + 1]
raise AssertionError(f"could not extract {name}")
def _composer_status_source(source: str) -> str:
owner = "let _composerStatusTimer=null;"
owner_start = source.index(owner)
function_start = source.index("function setComposerStatus(", owner_start)
function_source = _function_source(source, "setComposerStatus")
assert owner_start < function_start
return source[owner_start : function_start + len(function_source)]
def test_composer_status_timeout_is_race_safe():
composer_status_source = _composer_status_source(UI_JS)
script = textwrap.dedent(
f"""
const statusElement = {{
style: {{display: 'none'}},
textContent: '',
classList: {{remove() {{}}}},
removeAttribute() {{}},
}};
globalThis.window = {{_composerControlVisibility: {{}}}};
globalThis.$ = (id) => id === 'composerStatus' ? statusElement : null;
let now = 0;
let nextTimerId = 1;
const timers = new Map();
globalThis.setTimeout = (fn, delay) => {{
const id = nextTimerId++;
timers.set(id, {{fn, due: now + delay}});
return id;
}};
globalThis.clearTimeout = (id) => timers.delete(id);
function advance(ms) {{
const target = now + ms;
while (true) {{
let next = null;
for (const [id, timer] of timers) {{
if (timer.due <= target && (!next || timer.due < next.timer.due)) {{
next = {{id, timer}};
}}
}}
if (!next) break;
timers.delete(next.id);
now = next.timer.due;
next.timer.fn();
}}
now = target;
}}
function visible() {{
return statusElement.style.display !== 'none' && statusElement.textContent;
}}
function expect(condition, message) {{
if (!condition) throw new Error(message);
}}
{composer_status_source}
setComposerStatus('Reconnected', 1000);
expect(visible() === 'Reconnected', 'timed status should be visible immediately');
advance(999);
expect(visible() === 'Reconnected', 'timed status should remain visible before expiry');
advance(1);
expect(!visible(), 'timed status should hide at expiry');
setComposerStatus('Old status', 1000);
advance(500);
setComposerStatus('New status');
advance(500);
expect(visible() === 'New status', 'old timer must not clear a newer untimed status');
setComposerStatus('Same status', 1000);
advance(500);
setComposerStatus('Same status', 1000);
advance(500);
expect(visible() === 'Same status', 'repeated timed status needs a fresh timeout');
advance(500);
expect(!visible(), 'repeated timed status should hide after its fresh timeout');
setComposerStatus('Untimed status');
advance(5000);
expect(visible() === 'Untimed status', 'untimed status must remain visible');
setComposerStatus('Clear me', 1000);
setComposerStatus('');
expect(!visible(), 'explicit clear must hide immediately');
advance(1000);
expect(!visible(), 'explicit clear must cancel the owned timer');
setComposerStatus('Hide me', 1000);
window._composerControlVisibility.hide_composer_status = true;
setComposerStatus('Still hidden');
expect(!visible(), 'hidden-status preference must hide the status');
advance(1000);
expect(!visible(), 'hidden-status preference must cancel timed expiry');
window._composerControlVisibility.hide_composer_status = false;
process.stdout.write(JSON.stringify({{ok: true}}));
"""
)
completed = subprocess.run(
["node", "-e", script],
check=True,
capture_output=True,
text=True,
)
assert json.loads(completed.stdout) == {"ok": True}
def test_messages_wires_all_timed_composer_statuses_through_shared_owner():
reconnected_calls = re.findall(
r"setComposerStatus\('Reconnected'(?:,[^)]*)?\);", MESSAGES_JS
)
assert reconnected_calls == [
"setComposerStatus('Reconnected',1000);",
"setComposerStatus('Reconnected',1000);",
]
assert "setComposerStatus('Reconnected');" not in MESSAGES_JS
assert (
"setComposerStatus(`${d.message||'Warning'}`,d.type==='fallback'?4000:undefined);"
in MESSAGES_JS
)
assert "setTimeout(()=>setComposerStatus(''),4000)" not in MESSAGES_JS
+2 -2
View File
@@ -12,7 +12,6 @@ import re
from pathlib import Path
import pytest
from playwright.sync_api import sync_playwright
ROOT = Path(__file__).resolve().parents[1]
STATIC = ROOT / "static"
@@ -200,7 +199,8 @@ async ({kind, doneSid}) => {
@pytest.fixture(scope="module")
def browser():
with sync_playwright() as playwright:
playwright_api = pytest.importorskip("playwright.sync_api")
with playwright_api.sync_playwright() as playwright:
instance = playwright.chromium.launch(
headless=True,
args=["--no-sandbox", "--disable-dev-shm-usage"],
File diff suppressed because it is too large Load Diff
+40
View File
@@ -65,3 +65,43 @@ def test_delivery_options_local_label():
local_entry = next(p for p in result["platforms"] if p["value"] == "local")
# Label should contain "Local" or be an i18n key — just verify it's non-empty
assert local_entry["label"], "Local platform label is empty"
def test_delivery_options_survives_the_authority_module_move():
"""The platform list must not silently empty when the Agent relocates it.
``_KNOWN_DELIVERY_PLATFORMS`` used to live in ``cron.scheduler`` and now
lives in ``cron.scheduler_delivery``. The endpoint previously imported the
old path inside a bare ``except`` and fell back to an EMPTY frozenset, so
the move silently degraded the cron delivery picker to local/origin only —
every messaging platform (telegram, discord, slack, feishu, ...) vanished
from the UI with no error anywhere. Pin the resolution order so a future
relocation fails loudly instead of silently dropping platforms.
Registered in ``_AGENT_DEPENDENT_TESTS`` (tests/conftest.py) alongside the
other delivery-options tests: CI runs agent-free, so there is no ``cron``
package and no authority to resolve there.
"""
import importlib
resolved = frozenset()
for module_name in ("cron.scheduler_delivery", "cron.scheduler"):
try:
mod = importlib.import_module(module_name)
except Exception:
continue
known = getattr(mod, "_KNOWN_DELIVERY_PLATFORMS", None)
if known:
resolved = frozenset(known)
break
assert resolved, (
"_KNOWN_DELIVERY_PLATFORMS resolved empty from every known module path "
"— the cron delivery picker would silently show only local/origin"
)
# The endpoint's own output must agree with the resolved authority.
result, status = get("/api/crons/delivery-options")
assert status == 200
values = {p["value"] for p in result["platforms"]}
missing = resolved - values
assert not missing, f"platforms resolved but absent from the endpoint: {sorted(missing)}"
+306 -10
View File
@@ -1,4 +1,5 @@
import json
import os
import pathlib
import shutil
import subprocess
@@ -98,6 +99,7 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
btn._classes = btn.classList._set;
if (id === 'dashboardRailBtn' || id === 'dashboardMobileBtn') {
btn.setAttribute('data-dashboard-link', '');
btn.setAttribute('data-tooltip', 'Dashboard');
btn.setAttribute('aria-label', 'Dashboard');
}
return btn;
@@ -109,6 +111,12 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
const uiSrc = fs.readFileSync(process.argv[5], 'utf8');
const panelsSrc = fs.readFileSync(process.argv[6], 'utf8');
const winHostname = process.env.DASH_WINDOW_HOSTNAME || '127.0.0.1';
const statusRunning = process.env.DASH_STATUS_RUNNING;
const statusBrowserUrl = process.env.DASH_STATUS_BROWSER_URL;
const statusUrl = process.env.DASH_STATUS_URL;
const statusPort = process.env.DASH_STATUS_PORT;
const modeEl = makeButton('settingsDashboardMode');
const urlEl = makeButton('settingsDashboardUrl');
const statusEl = makeButton('settingsDashboardStatus');
@@ -122,12 +130,21 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
const delayedConfigValue = { enabled: 'never', url: 'http://stale.local:1234' };
let delayedConfigUsed = false;
function dashboardStatusFromEnv() {
return {
running: statusRunning !== undefined ? statusRunning === '1' : modeEl.value !== 'never',
browser_url: statusBrowserUrl !== undefined ? statusBrowserUrl : (modeEl.value === 'never' ? '' : 'http://127.0.0.1:1234'),
...(statusUrl !== undefined ? { url: statusUrl } : {}),
...(statusPort !== undefined ? { port: statusPort } : {}),
};
}
global._dashboardLastNonNeverMode = 'auto';
global._dashboardStatusCache = null;
global._dashboardStatusFetchedAt = 0;
global._dashboardSettingsLoadSeq = 0;
global._dashboardSettingsWriteSeq = 0;
global.window = { location: { hostname: '127.0.0.1' } };
global.window = { location: { hostname: winHostname } };
global.document = {
createElement: () => makeEl(),
querySelectorAll: (sel) => {
@@ -149,7 +166,6 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
if (key === 'dashboard_loopback_warning') return 'Loopback';
return key;
};
global._dashboardIsBrowserLoopback = () => false;
global.api = (url, opts = {}) => {
result.calls.push({ url: String(url), method: (opts.method || 'GET').toUpperCase(), body: opts.body || '', timeoutToast: !!(opts.timeoutToast) });
@@ -177,10 +193,7 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
}
if (String(url) === '/api/dashboard/status') {
result.statusCalls += 1;
return Promise.resolve({
running: modeEl.value !== 'never',
browser_url: modeEl.value === 'never' ? '' : 'http://127.0.0.1:1234',
});
return Promise.resolve(dashboardStatusFromEnv());
}
return Promise.resolve({ running: false });
};
@@ -192,7 +205,7 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
for (const name of ['_normalizeDashboardEnabledMode','_setDashboardModeForChip','_getDashboardChipRestoreMode']) {
eval(extractFn(uiSrc, name));
}
for (const name of ['_dashboardBrowserUrl', '_applyDashboardStatus', 'refreshDashboardStatus', 'loadDashboardSettings', 'saveDashboardSettings']) {
for (const name of ['_dashboardBrowserUrl', '_dashboardHostIsLoopback', '_dashboardIsBrowserLoopback', '_dashboardUrlIsLoopback', '_applyDashboardStatus', 'refreshDashboardStatus', 'loadDashboardSettings', 'saveDashboardSettings']) {
let src = extractFn(uiSrc, name);
if(name === 'saveDashboardSettings'){
src = src.replace(
@@ -213,6 +226,7 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
display: btn.style.display || '',
dashboardUrl: btn._attrs['data-dashboard-url'] || '',
tooltip: btn._attrs['data-tooltip'] || '',
ariaLabel: btn._attrs['aria-label'] || '',
}));
}
@@ -277,6 +291,13 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
return;
}
if (action === 'status-apply') {
_applyDashboardStatus(dashboardStatusFromEnv());
recordButtons();
console.log(JSON.stringify({ buttonStates: result.buttonStates }));
return;
}
if (action === 'chip-toggle') {
_toggleDashboardVisibilityChip();
} else {
@@ -299,11 +320,14 @@ _DASHBOARD_LINK_DRIVER = textwrap.dedent(
)
def _run_dashboard_link_driver(action: str, mode: str = 'auto', url: str = '') -> dict:
def _run_dashboard_link_driver(action: str, mode: str = 'auto', url: str = '', env: dict | None = None) -> dict:
with tempfile.NamedTemporaryFile("w", suffix=".js", encoding="utf-8", delete=False) as f:
f.write(_DASHBOARD_LINK_DRIVER)
driver = f.name
try:
run_env = os.environ.copy()
if env:
run_env.update(env)
result = subprocess.run(
[
NODE,
@@ -319,6 +343,7 @@ def _run_dashboard_link_driver(action: str, mode: str = 'auto', url: str = '') -
capture_output=True,
timeout=2.0,
check=False,
env=run_env,
)
if result.returncode != 0:
raise RuntimeError(f"node harness failed: {result.stderr or result.stdout}")
@@ -462,7 +487,17 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
};
}
appendChild(child) { child.parentNode = this; this.children.push(child); return child; }
get nextElementSibling() {
if (!this.parentNode) return null;
const index = this.parentNode.children.indexOf(this);
return index >= 0 ? this.parentNode.children[index + 1] || null : null;
}
insertBefore(child, anchor) {
if (child.parentNode) this.insertBeforeMoves = (this.insertBeforeMoves || 0) + 1;
if (child.parentNode) {
const currentIdx = child.parentNode.children.indexOf(child);
if (currentIdx >= 0) child.parentNode.children.splice(currentIdx, 1);
}
const idx = anchor ? this.children.indexOf(anchor) : -1;
child.parentNode = this;
if (idx >= 0) this.children.splice(idx, 0, child);
@@ -507,6 +542,7 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
return [];
}
querySelector(sel) {
if (sel.startsWith('.nav-tab[data-panel="')) return this.children.find(el => el.classList.contains('nav-tab') && el.getAttribute('data-panel') === sel.slice(21, -2)) || null;
if (sel.startsWith('[data-nav-action-mirror="')) {
const id = sel.slice(25, -2);
return this.children.find(el => el.getAttribute('data-nav-action-mirror') === id) || null;
@@ -520,7 +556,10 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
const rail = new El('nav');
const sidebar = new El('div');
const source = new El('button');
const source2 = new El('button');
const dashboard = new El('button');
const chatPanel = new El('button');
const settingsPanel = new El('button');
let clicked = 0;
let sidebarClosed = 0;
source.id = 'hwxThemeCreatorRailBtn';
@@ -530,8 +569,21 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
source.onclick = () => { clicked += 100; };
source.innerHTML = '<svg></svg>';
source.click = () => { clicked += 1; };
source2.id = 'hwxTypographyRailBtn';
source2.classList.add('rail-btn', 'nav-tab', 'has-tooltip');
source2.setAttribute('data-tooltip', 'Typography');
source2.innerHTML = '<svg></svg>';
chatPanel.id = 'chatMobileBtn';
chatPanel.classList.add('nav-tab');
chatPanel.setAttribute('data-panel', 'chat');
settingsPanel.id = 'settingsMobileBtn';
settingsPanel.classList.add('nav-tab');
settingsPanel.setAttribute('data-panel', 'settings');
dashboard.classList.add('dashboard-link');
dashboard.id = 'dashboardMobileBtn';
sidebar.appendChild(chatPanel);
sidebar.appendChild(dashboard);
sidebar.appendChild(settingsPanel);
let observerCallback = null;
let observedRail = null;
let observedOptions = null;
@@ -561,17 +613,53 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
const helpers = Function(extractFn('_stripInlineEventHandlers') + '\\n' + extractFn('_syncNavActionMirrors') + '\\n' + extractFn('_initNavActionMirrors') + '; return { _initNavActionMirrors };')();
helpers._initNavActionMirrors();
rail.appendChild(source);
rail.appendChild(source2);
observerCallback();
const mirror = sidebar.children.find(el => el.getAttribute('data-nav-action-mirror') === 'hwxThemeCreatorRailBtn');
const mirror2 = sidebar.children.find(el => el.getAttribute('data-nav-action-mirror') === 'hwxTypographyRailBtn');
source.hidden = true;
observerCallback();
const mirrorHiddenWhenSourceHidden = !mirror.classList.contains('nav-action-visible');
source.hidden = false;
observerCallback();
mirror.click();
const mirrorBeforeDashboard = sidebar.children[0] === mirror;
const mirrorBeforeDashboard = sidebar.children[1] === mirror && sidebar.children[2] === mirror2 && sidebar.children[3] === dashboard;
const initialOrder = sidebar.children.map(el => el.id);
const tasksPanel = new El('button');
tasksPanel.id = 'tasksMobileBtn';
tasksPanel.classList.add('nav-tab');
tasksPanel.setAttribute('data-panel', 'tasks');
const logsPanel = new El('button');
logsPanel.id = 'logsMobileBtn';
logsPanel.classList.add('nav-tab');
logsPanel.setAttribute('data-panel', 'logs');
sidebar.appendChild(tasksPanel);
sidebar.appendChild(logsPanel);
['tasks', 'logs'].forEach(panel => {
const node = sidebar.querySelector('.nav-tab[data-panel="' + panel + '"]');
if (node) sidebar.insertBefore(node, dashboard);
});
const preSyncOrder = sidebar.children.map(el => el.id);
const movesBeforeReconcile = sidebar.insertBeforeMoves || 0;
observerCallback();
const reconciledOrder = sidebar.children.map(el => el.id);
const reconciliationMoves = (sidebar.insertBeforeMoves || 0) - movesBeforeReconcile;
const movesBeforeNoOpSync = sidebar.insertBeforeMoves || 0;
observerCallback();
const noOpSyncMoves = (sidebar.insertBeforeMoves || 0) - movesBeforeNoOpSync;
source.remove();
observerCallback();
const mirrorRemoved = !sidebar.children.some(el => el.getAttribute('data-nav-action-mirror') === 'hwxThemeCreatorRailBtn');
dashboard.remove();
logsPanel.remove();
const source3 = new El('button');
source3.id = 'hwxNoAnchorRailBtn';
source3.classList.add('rail-btn', 'nav-tab');
source3.setAttribute('aria-label', 'No anchor');
rail.appendChild(source3);
observerCallback();
const noAnchorMirror = sidebar.children.find(el => el.getAttribute('data-nav-action-mirror') === 'hwxNoAnchorRailBtn');
const noAnchorMirrorAppended = noAnchorMirror?.parentNode === sidebar && sidebar.children.at(-1) === noAnchorMirror;
console.log(JSON.stringify({
observerArmed: observedRail === rail,
observerAttributes: !!(observedOptions && observedOptions.attributes),
@@ -584,7 +672,13 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
mirrorOnclickAttribute: mirror.getAttribute('onclick'),
mirrorOnclickProperty: mirror.onclick === null,
mirrorBeforeDashboard,
mirrorRemoved: !sidebar.children.some(el => el.getAttribute('data-nav-action-mirror') === 'hwxThemeCreatorRailBtn'),
initialOrder,
preSyncOrder,
reconciledOrder,
reconciliationMoves,
noOpSyncMoves,
mirrorRemoved,
noAnchorMirrorAppended,
clicked,
sidebarClosed,
}));
@@ -610,7 +704,35 @@ def test_extension_rail_actions_are_mirrored_to_mobile_nav():
"mirrorOnclickAttribute": None,
"mirrorOnclickProperty": True,
"mirrorBeforeDashboard": True,
"initialOrder": [
"chatMobileBtn",
"hwxThemeCreatorRailBtnMobile",
"hwxTypographyRailBtnMobile",
"dashboardMobileBtn",
"settingsMobileBtn",
],
"preSyncOrder": [
"chatMobileBtn",
"hwxThemeCreatorRailBtnMobile",
"hwxTypographyRailBtnMobile",
"tasksMobileBtn",
"logsMobileBtn",
"dashboardMobileBtn",
"settingsMobileBtn",
],
"reconciledOrder": [
"chatMobileBtn",
"tasksMobileBtn",
"logsMobileBtn",
"hwxThemeCreatorRailBtnMobile",
"hwxTypographyRailBtnMobile",
"dashboardMobileBtn",
"settingsMobileBtn",
],
"reconciliationMoves": 2,
"noOpSyncMoves": 0,
"mirrorRemoved": True,
"noAnchorMirrorAppended": True,
"clicked": 1,
"sidebarClosed": 1,
}
@@ -703,3 +825,177 @@ def test_failed_dashboard_save_reloads_backend_after_stale_load_is_dropped():
("/api/dashboard/config", "GET"),
]
assert out["statusCalls"] == 0
@requires_node
def test_remote_webui_public_target_has_normal_tooltip_and_aria_label():
# Remote WebUI + public HTTPS target: resolved navigation URL host is
# non-loopback, so no loopback-only warning (normal tooltip + ARIA label).
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": "webui.example.test",
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": "https://dashboard.example.test",
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Dashboard"
assert state["ariaLabel"] == "Dashboard"
@requires_node
@pytest.mark.parametrize(
"target_url",
[
"http://127.0.0.1:1234",
"http://localhost:1234",
"http://[::1]:1234",
],
)
def test_remote_webui_loopback_target_keeps_loopback_warning(target_url):
# Remote WebUI + loopback target: the warning must remain even though
# browser_url is truthy — the resolved URL still points at loopback.
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": "webui.example.test",
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": target_url,
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Loopback"
assert state["ariaLabel"] == "Loopback"
@requires_node
def test_remote_webui_auto_probe_url_only_loopback_target_keeps_warning():
# Successful auto-probe arm: the probe returns a resolved URL with no
# browser_url field (enabled:always instead emits both url and browser_url).
# A loopback probe target must still warn for a remote WebUI — the warning
# keys off the resolved URL host, not browser_url presence.
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": "webui.example.test",
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": "",
"DASH_STATUS_URL": "http://127.0.0.1:1234",
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Loopback"
assert state["ariaLabel"] == "Loopback"
@requires_node
@pytest.mark.parametrize(
"target_url",
[
"http://127.8.9.10:1234", # other 127/8 address
"http://[::ffff:127.0.0.1]:1234", # IPv4-mapped IPv6 loopback (dotted)
"http://[::ffff:7f00:1]:1234", # IPv4-mapped IPv6 loopback (hex)
"http://myhost.localhost:1234", # .localhost name (RFC 6761)
"http://localhost.:1234", # terminal hostname dot
"http://[::1]:1234", # IPv6 loopback
],
)
def test_remote_webui_loopback_classifier_variants_keep_warning(target_url):
# enabled:always producer shape emits both url and browser_url. Every
# loopback spelling of the resolved target must keep the warning for a
# remote WebUI (tooltip + ARIA).
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": "webui.example.test",
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": target_url,
"DASH_STATUS_URL": target_url,
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Loopback"
assert state["ariaLabel"] == "Loopback"
@requires_node
@pytest.mark.parametrize(
"target_url",
[
"https://dashboard.example.test", # public HTTPS
"http://128.0.0.1:1234", # adjacent public IPv4
"http://[::2]:1234", # adjacent public IPv6
"http://127.0.0.1.example.com:1234", # DNS name containing loopback prefix
],
)
def test_remote_webui_public_boundary_targets_have_normal_tooltip(target_url):
# Public boundary controls: addresses adjacent to loopback and DNS names
# that merely contain a loopback-looking prefix must not warn.
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": "webui.example.test",
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": target_url,
"DASH_STATUS_URL": target_url,
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Dashboard"
assert state["ariaLabel"] == "Dashboard"
@requires_node
@pytest.mark.parametrize("bad_url", ["not a url", "/relative/dashboard"])
def test_remote_webui_invalid_or_relative_target_urls_do_not_warn(bad_url):
# Unparseable/relative targets fall back to the origin-derived URL; the
# warning must not fire spuriously and the link must not crash.
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": "webui.example.test",
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": bad_url,
"DASH_STATUS_URL": bad_url,
"DASH_STATUS_PORT": "1234",
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Dashboard"
assert state["ariaLabel"] == "Dashboard"
@requires_node
@pytest.mark.parametrize(
"origin_hostname",
["127.0.0.1", "127.0.0.2", "localhost", "myhost.localhost", "[::1]"],
)
def test_loopback_webui_origin_has_no_remote_browser_warning(origin_hostname):
# Loopback WebUI origin: the operator is on the same machine, so the
# loopback target is reachable — no remote-browser warning.
out = _run_dashboard_link_driver(
"status-apply",
mode="always",
env={
"DASH_WINDOW_HOSTNAME": origin_hostname,
"DASH_STATUS_RUNNING": "1",
"DASH_STATUS_BROWSER_URL": "http://127.0.0.1:1234",
},
)
assert out["buttonStates"]
for state in out["buttonStates"]:
assert state["tooltip"] == "Dashboard"
assert state["ariaLabel"] == "Dashboard"
+119
View File
@@ -0,0 +1,119 @@
"""The display merge for paginated loads must be memoized for inactive sessions.
Before the fix, GET /api/session re-ran merge_session_messages_append_only on
every request (~2-3s for multi-thousand-message transcripts). The cache must:
- return the same merged transcript as a direct merge (correctness),
- serve copies (caller mutation cannot corrupt the cache),
- invalidate when the sidecar file changes on disk,
- invalidate when the state.db rows change,
- never cache ACTIVE sessions (active_stream_id / pending_user_message).
"""
import time
from types import SimpleNamespace
import pytest
@pytest.fixture()
def routes_env(tmp_path, monkeypatch):
import api.config as config
import api.models as models
import api.routes as routes
home = tmp_path / "home"
state_dir = tmp_path / "state"
session_dir = state_dir / "sessions"
session_dir.mkdir(parents=True)
monkeypatch.setenv("HERMES_HOME", str(home))
monkeypatch.setenv("HERMES_WEBUI_STATE_DIR", str(state_dir))
monkeypatch.setattr(config, "STATE_DIR", state_dir)
monkeypatch.setattr(config, "SESSION_DIR", session_dir)
monkeypatch.setattr(models, "SESSION_DIR", session_dir)
monkeypatch.setattr(routes, "SESSION_DIR", session_dir)
with routes._display_merge_cache_lock:
routes._display_merge_cache.clear()
yield SimpleNamespace(config=config, models=models, routes=routes)
with routes._display_merge_cache_lock:
routes._display_merge_cache.clear()
def _make_session(routes_env, sid="20260101_000000_cache1", n=6):
models = routes_env.models
s = models.Session(session_id=sid)
now = time.time()
for i in range(n):
role = "user" if i % 2 == 0 else "assistant"
s.messages.append({"role": role, "content": f"turn {i}", "timestamp": now + i})
s.save()
return s
def _state_rows(base_ts, n=2):
return [
{"role": "assistant", "content": f"state row {i}", "timestamp": base_ts + 100 + i}
for i in range(n)
]
def test_merge_cached_and_equal_to_direct_merge(routes_env):
routes = routes_env.routes
s = _make_session(routes_env)
rows = _state_rows(s.messages[-1]["timestamp"])
first = routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
assert routes._display_merge_cache, "expected a cache entry for an inactive session"
second = routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
assert [m.get("content") for m in first] == [m.get("content") for m in second]
assert any(m.get("content") == "state row 0" for m in second)
def test_cache_hit_returns_copies(routes_env):
routes = routes_env.routes
s = _make_session(routes_env, sid="20260101_000000_cache2")
rows = _state_rows(s.messages[-1]["timestamp"])
first = routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
first[0]["content"] = "MUTATED"
second = routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
assert second[0]["content"] != "MUTATED"
def test_cache_invalidated_by_sidecar_write(routes_env):
routes = routes_env.routes
s = _make_session(routes_env, sid="20260101_000000_cache3")
rows = _state_rows(s.messages[-1]["timestamp"])
routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
s.messages.append({
"role": "assistant", "content": "new turn after write",
"timestamp": time.time() + 50,
})
s.save()
merged = routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
assert any(m.get("content") == "new turn after write" for m in merged)
def test_cache_invalidated_by_state_rows_change(routes_env):
routes = routes_env.routes
s = _make_session(routes_env, sid="20260101_000000_cache4")
rows = _state_rows(s.messages[-1]["timestamp"])
routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
rows2 = rows + [{
"role": "assistant", "content": "brand new state row",
"timestamp": rows[-1]["timestamp"] + 5,
}]
merged = routes._limited_webui_messages_for_display_with_sidecar(s, None, rows2)
assert any(m.get("content") == "brand new state row" for m in merged)
def test_active_session_never_cached(routes_env):
routes = routes_env.routes
s = _make_session(routes_env, sid="20260101_000000_cache5")
s.active_stream_id = "stream-live"
rows = _state_rows(s.messages[-1]["timestamp"])
routes._display_merge_cache.clear()
routes._limited_webui_messages_for_display_with_sidecar(s, None, rows)
assert s.session_id not in routes._display_merge_cache

Some files were not shown because too many files have changed in this diff Show More