The lazy-row enrichment runs on the first submit, so a row built through
`_ensure_session_db_row` is already enriched by the time hydration sees it —
one setup can no longer cover both halves. Build the row directly (cwd, no git
metadata), which is the state a pre-enrichment row is actually in, and assert
hydration heals it. Both paths stay independently covered.
Salvaged from #108789 (@fangliquanflq).
An agent told to make a worktree and work in it left the chat on the checkout it
started in — the sidebar kept the chat in the primary branch lane while every
later command ran in the worktree.
`_reconcile_session_cwd_from_terminal` refused whenever `explicit_cwd` was set,
but `session.create` sets that for ANY session whose cwd exists on disk, and the
desktop always sends a workspace. The guard that was meant to protect a
deliberately chosen workspace (#72776) therefore disabled the follow for the
whole desktop surface.
Key the pin on `cwd_pinned`, set only by the two moves the user drives:
`session.cwd.set` (the composer's folder picker) and `session.workspace.move`
(move-to-project). A cwd a chat merely STARTED in is not a pin, so the agent's
own worktree settle re-anchors the chat and the desktop moves it to that lane.
The cut printed before GitHub listed the dispatched run, so it almost always
said "the run is not listed yet". It now polls up to 30s for the run, and when
none lists it links the stable-release workflow page. The copy is shorter:
Release CI, the skipped-tests warning (bold on a tty), the draft link, and the
publish command on its own line.
git keeps an annotated tag's signature inside the message body, so
`git tag -l --format=%(contents)` returns the JSON record followed by an
armored block, and `json.loads` of the whole message raises "Extra data: line 2
column 1" at the armor's first byte. Claim tags, final receipts and build
receipts are created on a host whose git signs tags, so stable admission
(`python -m scripts.releases.stable admit`), ordered reconciliation
(`python -m scripts.releases.sequencer`) and tagless build receipts each
refused to read their own tag.
Add `versioning.tag_record()` as the one reader of a tag's record — the text
before the armor, unchanged for an unsigned tag or a detached signature — and
route every `%(contents)` read through it in stable.py, sequencer.py and
commit_build.py. commit_build imports it inside the function so its module
surface stays stdlib-only for the isolated checkout-admission step.
The reads inside the release tests use the same reader, and the one fixture
that creates a plain final tag says so explicitly, so the release suite passes
on a host whose git signs tags by default.
Main claims Ctrl/Cmd+R on every window (installPreviewShortcut's
before-input-event + preventDefault) so Chromium's default host
reload cannot fire, and sends preview-nav reload to the renderer.
The renderer's routing only knew browser panes, so a terminal-
focused chord fell through to window.location.reload() — and the
preventDefault meant xterm never saw the keystroke either, killing
readline reverse-i-search (#96482).
Route it the way ⌘W already routes: main keeps claiming blindly,
and the renderer adds the missing terminal rung. The terminal
registry keyed its handles by the inner xterm host while both
resolvers walk up to the [data-terminal] scope, so no handle was
ever found in production (only tests registered on the scope div
itself) — the registry now keys by that scope. The user terminal's
handle re-delivers the ^R byte to its PTY; the read-only agent
mirror swallows the chord (no PTY, and focus there must still stop
the app-level reload). The app-reload escape hatch elsewhere in
the app is unchanged.
electron-builder resolves every piece of packaging artwork from the workspace
— the app icon, the icon.ico extraResource, the MSIX logo set before-build
stages into build/appx, and the exe identity stamp — so a flavored canary or
commit build copies its rendered icon set over apps/desktop/assets before
packaging.
Those files stopped being gitignored in 4b7229d612, which turned a write that
was invisible into 15 dirty paths. The source custody check then refuses the
checkout twice: inside the same build, at the check that used to sit after the
copy, and on every later build from the first check on a now permanently dirty
tree.
Hold the render at the workspace path for the packaging window only and
restore the admitted files afterwards, whether packaging finished or failed.
The final custody check now asserts the build handed the checkout back.
stable-release.yml called desktop-bundled-release.yml, docker.yml and
termux-verify.yml without `secrets: inherit`, so every job inside them that
declares `environment: release-signing` (or container-publish) attached its
deployment and read the environment's `vars`, but resolved `secrets.*` to the
empty string. The signed-candidate legs died at "Archive every pinned input"
with `missing env CLOUDFLARE_R2_ACCESS_KEY_ID`; the docker publish legs and
the termux packaging checks would have failed the same way.
A called workflow receives no secrets by default, and a job-level
`environment:` inside it is required but not sufficient: a caller job that
`uses:` a reusable workflow cannot declare `environment:` at all, so
`secrets: inherit` is the only mechanism that can carry them
(actions/runner#1490).
`secrets: inherit` is added only where the callee reads a protected
environment's secrets. ci.yaml, nix, pm-bundle, windows-venv-e2e and the
install-e2e chain reference none, and ci.yaml must stay secret-free by its own
header (it runs PR-controlled code).
The spawn site (marker + pin) and the worker entry (PM's `activate_dependencies`) are one
contract across two files: dropping either half loses the worker's dependencies or its
generation lease. Name it in the area's hardening-invariants list so the next refactor of
`_launch_external_cron_worker` or `cron/worker_bootstrap.py` answers for it (#122222).
(cherry picked from commit eb00b48b1718fd7d2ac96503c8f9d2ea8443309b)
The restart-safe cron worker is spawned as `sys.executable -m cron.scheduler`, a Hermes
entry point that never goes through `hermes_bootstrap`. On a self-managed (PM) install its
interpreter is the store Python, which owns no dependencies and whose PYTHONPATH the
subprocess sanitizer has just stripped, so the worker died at its first dependency import
(`No module named 'ruamel'`, #122222) before it could publish its ownership
acknowledgement -- every scheduled job landed `last_status: failed`.
Salvaged from #122290 (thanks @123456hyu), whose two commits pinned the dependency
site-packages and leased the generation at the worker entry. Reworked onto PM's own
machinery:
* `pin_hermes_tree_on_pythonpath` asks `pm.environments.committed_venv` +
`site_packages()` for the generation, instead of deriving one from this process's
`sys.path` behind a `pyvenv.cfg` ancestor probe. The record is scoped to the active
home, so the child-spawn path no longer reads the real home -- that probe tripped
`tests/home_io_guard.py` when the suite runs from a self-managed install -- and the
uv-launcher, first-match-only and ordering edge cases of the old derivation go with it.
* `cron/worker_bootstrap.py` is now the PM call the worker actually needs: lease +
`site.addsitedir` + `sys.path`/`PYTHONPATH`/`PATH` are all
`pm.environments.activate_dependencies`, the same boot `hermes_bootstrap` runs, and it
holds the generation's lease for the process's life. It is gated on the
`HERMES_CRON_EXTERNAL_WORKER` marker the spawn site sets (consumed on read) rather than
sniffing `--external-worker-file` out of `sys.argv`, so the gateway keeps its launch
contract without argv identity inference.
Verified live on PM's store Python against a real generation (A/B, same child): tree-only
PYTHONPATH -> `ModuleNotFoundError: No module named 'ruamel'`, rc=1; pinned env -> imports,
`running_from_selected_environment` True, and the worker holds a real kernel lease while it
runs (released on exit). The real `-m cron.scheduler --external-worker-file ...` entry now
reaches its ownership handoff instead of dying at the import. `tests/cron/`: 1376 passed.
Tests: the two invariants that carry the contract -- the pin restores the committed
generation (and invents nothing without one), and the boot runs only in the marked worker,
once. The six construction tests that drove the old derivation are dropped with it.
Duplicate-cluster credit: #122238 (@webtecnica, earliest), #122382 (@Sakujakira), #122685
(@etrisnanto), #122719 (@aglowinthefield) and #122738 (@GabboAlarcon) diagnosed and fixed
this same worker-dependency failure in parallel while this branch was rebuilt. None of their
diffs is contained here -- the co-author trailers below record direction and duplicate-cluster
credit, not lifted code; their roles are listed in the PR body.
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Manuel S. <3703103+Sakujakira@users.noreply.github.com>
Co-authored-by: Eko Trisnanto <4469382+etrisnanto@users.noreply.github.com>
Co-authored-by: aglowinthefield <146008217+aglowinthefield@users.noreply.github.com>
Co-authored-by: GabboAlarcon <40816412+GabboAlarcon@users.noreply.github.com>
(cherry picked from commit 7c55715c0478280adf4e492b24edf331342470f7)
Models used to live at <profile home>/models (adb2fdbf5c) before the managed
runtime moved them to the machine-scoped <root>/models (43e67d872f). GGUFs
staged under a named profile's old dir silently stopped being served.
adopt_legacy_models() renames them (and their assets/) into the current dirs,
so everything downstream keeps reading one directory: listing, presets,
delete, the --models-dir fallback. It runs at boot, before the "anything
staged?" check, and in the Local Models status route so the pane lists them
while the runtime is off.
os.rename only: instant within a filesystem, while shutil.move would silently
copy tens of GB across devices at session start; a cross-device dir stays put
with a warning. An existing destination name is never replaced.
Every gate and every candidate now needs only admit, so the signed macOS
and Windows builds and the docker image stop waiting behind the full CI
run. acceptance stays the one join: it still needs ci, docker and all six
candidates, so what can be promoted is unchanged.
The four gates that skipped under --skip-tests only because ci skipped
(nix, termux-checks, windows-live, install-e2e) and pm-bundle needed their
own condition. SKIPPED_BY requires the gate to observe 'skipped', so
dropping the edge alone would have left them running and blocked the
release instead of failing it.
An install's in-tree venv whose editable record names another checkout
(what project_venv_dir used to cause) turns <install>/venv/bin/hermes
into that checkout's CLI. Every update Desktop hands to the launcher
then pulls the other tree: the install never moves, the hand-off
reports success, and posix.sh swaps the install's stale release/ build
back over the app. Pressing Update can never get out of that loop.
`hermes update` now notices it is running on another checkout's in-tree
venv and re-runs itself with that install's own code (PYTHONPATH=<install>,
cwd=<install>). The install's updater pulls the install and reinstalls it
into its venv, which points the editable record home again, so the next
relaunch runs a fresh build. A wrapper that put the running checkout on
PYTHONPATH on purpose chose that tree and is left alone.
project_venv_dir() fell back to the running interpreter's venv whenever
hermes_constants was loaded from the checkout. Where the code was loaded
from says nothing about who owns the interpreter: with
`PYTHONPATH=<dev checkout> <app install>/venv/bin/python -m hermes_cli.main`
(a shell wrapper around a dev tree), PROJECT_ROOT is the dev checkout but
the venv is the Desktop install's. base_venv() then returned the app's
venv, `hermes update` synced the dev tree into it, and the Desktop
install's venv became an editable install of the dev checkout. From then
on the Desktop shell tracked ~/.hermes/hermes-agent while its backend and
its handed-off `hermes update` ran and pulled the dev tree, so every
in-app update "succeeded" without moving the install. The same misread
let running_from_selected_environment() accept lazy extras into that
venv.
Only fall back to the running venv when its own hermes-agent install
records this checkout in direct_url.json, which every install of a
checkout into a venv writes (installers, uv sync). Otherwise the
checkout gets its own environment, as it did before 4f6c04cd07. The
#116148 out-of-tree layout (a venv installed from the checkout) still
resolves to the running interpreter.
findPythonForRoot ended in findSystemPython(), so a checkout with no
in-tree venv/.venv produced a PATH Python. Two designed refusals were
therefore unreachable: readSourceUpdate's `!managed && !probe.python`
guard (checkout-source.ts) and StateDbPreflight's `string | null` python
plus its "Python not found" throw (state-db-preflight.ts). Both were
written expecting null and could never see it.
A PATH Python can import a checkout while lacking its selected
dependencies, which is the failure the checkout-source comment already
names, so a failed read became a wrong answer instead of a refusal:
the update probe, the state.db pre-flight and the source backend all
ran under an interpreter nothing selected. PM deletes the in-tree
venv/.venv once a generation is committed, making null the ordinary
answer for a managed install — its callers resolve the installation
launcher instead, and the backend ladder falls through to its next rung.
resolveSourcePython owns the decision (override, then the checkout's own
venv, else null) and findPythonForRoot keeps its signature, so the four
call sites are unchanged. findSystemPython stays for the uninstaller,
whose interpreter choice is a separate, deliberate one for a locked venv.
#122234 gave Desktop update steps NUL stdin, but the hand-off script
that runs is the one from the checkout being updated FROM. Every update
that starts on an older commit still runs the old script, which gives
steps the hand-off console as stdin and captures their stdout until
they exit. Only the steps after `hermes update` run new code.
So `gateway start --all` from the new checkout still saw an interactive
console, asked "Install it now so the gateway starts on login?" into
the captured stdout, and waited forever. The update never relaunched.
start() now asks only when stdout is a terminal too. Nobody can answer
a question they cannot see.
The APT package does not work at the moment and a fix is in progress.
Warn at the top of the install guide so users do not burn time on steps
that will fail.
Sitting a plugin out after a fetch failure left the recorded stamp stale on
purpose, so every launch resynced, and the spawned-process guard in
prepare_launch raised "dependency sync left this install out of date".
One retry covers a blip; after that the plugin is disabled with the reason,
and hermes plugins enable restores it. Only requires_hermes still sits out,
which boot skips the same way.
enabled_member_dirs imported hermes_cli.plugins_manifest before its loop, so a
PM closure with no plugins selected needed the application's utils module
(tests/scripts/test_source_driver.py builds exactly that tree).
The requires_hermes sit-out test relied on the host's version identity. A
tagless CI checkout has no parseable version, which makes the gate permissive.
An update disabled any plugin that failed a trial build or its
requires_hermes check. Both can be about us, not the plugin: an untagged
source checkout reads as an older release (#122054), so requires_hermes
misjudges a fine plugin, and a download failure says nothing about the
plugin's code. Those now sit the plugin out of the build: config stays
untouched and it rejoins once the cause clears.
Disabling still happens on evidence about the plugin: requires-python vs
the pinned interpreter, manifest_version, an invalid declaration, a uv
resolution conflict, or its own build backend failing (new BuildFailure,
keyed on uv's 'The build backend returned an error').
enabled_member_dirs now skips a requires_hermes misfit instead of
raising. Boot's currency check raised on it before any sync could run, so
the launch path never reached the update sync. The loader skips such a
plugin anyway; admission still refuses enabling one.
Venv.apply refused the whole graph when any secondary profile's
config.yaml was unreadable, so one broken sibling config failed every
update. Update syncs now leave that profile's plugins out of the union,
report it (stderr + receipt warning), and build the rest; the profile's
plugins rejoin on the next sync once its config is fixed. Boot currency
already skips broken secondaries, so the result reads as current.
Ordinary syncs keep refusing to shrink the recorded graph.
An update resolves the enabled plugin union against the NEW core. A plugin
admitted against the old core can stop fitting when core moves (managed
Python 3.13 -> 3.14 vs a member's requires-python <3.14, a requires_hermes
upper bound, a bumped pin), and the whole update then died after the
source swap with a non-resolver InstallError whose 'retry' hint failed the
same way every time.
Update syncs now pass evict_incompatible_plugins=True (update completion,
historical takeover, launch-time completion, venv_sync, post-update
drift). PM screens statically first (requires-python vs the target
interpreter, manifest/requires_hermes), then, if the rest still fails,
builds core alone to prove the plugins are the cause and re-adds members
in config order, disabling each one that breaks the build. Misfits land in
plugins.disabled (memory.provider cleared) in every home that enables
them, published through the existing journaled change hook (the journal
now carries several configs), and are reported on stderr + receipt
warnings. Admission and ordinary syncs still refuse; only a core that
cannot build on its own fails an update.
gh v2.99 gained a native --attach flag for issues, PRs and comments, so
the custom trusted-publisher chain (gh-image extension + GH_IMAGE_SESSION_TOKEN
workflow_run job + attachment-upload script) is superseded.
Remove:
- .github/workflows/publish-e2e-evidence.yml (workflow_run publisher)
- scripts/ci/publish_e2e_evidence.py + its tests
- e2e-evidence-* artifact upload + evidence staging in e2e-desktop.yml /
e2e_screenshot_status.py, incl. the 'inline evidence is publishing...'
marker placeholder in the CI review comment status
Keep: the review-comment screenshot/diff counts and artifact links produced
by e2e_screenshot_status.py.
The setup stage installs the gateway service through
ensure_gateway_service. On Windows that asks the start-now, Scheduled
Task and UAC questions. The gateway stage then ran `hermes gateway
install`, which asked them all again.
`gateway install --if-missing` does nothing when a service is already
installed. Both installers' gateway stages use it, so they ask only
when setup did not install the service.
Steps inherited the hand-off console's stdin, so any step that asks a
question blocked forever. Its prompt went to the captured stdout, which
is shown only after the step exits. `gateway start --all` did exactly
this: it saw an interactive console and asked "Install it now so the
gateway starts on login?". The update stopped after `hermes update exit
code: 0` and never relaunched the app.
Steps now read NUL. Prompts see a non-interactive stdin and take their
defaults. The working-directory self-test also checks that a step's
stdin is not a console, and a new test runs it under a real console.
ensure_import synced the extra into a new generation and then always
raised "restart Hermes to activate", so choosing a provider whose SDK is an
opt-in extra (anthropic, bedrock, ...) failed the first message in every
fresh install or worktree even though the install had succeeded.
adopt_selected() already knows when swapping a live process onto a new
generation is safe (it was on the generation selected before the sync, and
nothing it imported changed version). Call it after the sync; raise only
when adoption was refused, naming restart_needed()'s reason.
A process that could not adopt a newly published dependency generation keeps
importing the old one. restart_needed() names that case (and stays silent for
dev venvs, Nix and anything not booted from a PM generation, so no false
restart prompts); adopt_selected() lets callers move onto the selection before
loading new code. The test helpers publish real generations and make the test
process run from one.
(cherry picked from commit 978abe8ec4b852778b8c63bcefdf141db51bc703)
(cherry picked from commit 28be27334adc531db59bf37a02a64acaaf47a2ee)
After a plugin's runtime request is published, the process that asked should be
able to import the result without a restart, but only when nothing it already
loaded would change underneath it. Adopt when this process was on the
generation that was selected when the build started, and no loaded module's
file appears in the RECORD of a distribution whose version changed or
disappeared. Adopting swaps the site-packages entry, processes only the new
generation's new .pth files, rewrites PYTHONPATH/PATH for child processes, and
leases the new generation while keeping the old lease, since loaded packages
still resolve submodules from it.
(cherry picked from commit 478f945d9f36e94229f58b251642b7c9f3b9e781)
PR #122161 stopped boot from activating the in-tree venv or .venv when
PM has committed no environment. Six Linux tests and four Windows tests
still used that tree to supply their probe modules.
- Launcher tests put the probe in a committed generation. A custom
HERMES_HOME is its own dependency root, so it gets its own commit.
- The legacy row of test_pre_pm_base_dependencies_activate_only_at_boot
asserted the removed behavior. test_boot_never_activates_the_pre_pm_venv
now covers the inverse. The payload row stays.
- The mint payload fixture writes manifest.json as real payloads do, so
boot selects the payload venv.
- The PowerShell activate test expects PYTHONPATH to be the checkout
alone. The bootstrap .venv packages do not leak in.
_get_anthropic_sdk() swallowed every ensure_import("anthropic") failure and
_require_sdk() then told the user to "Install it with: hermes pm install
--extra anthropic". PM often HAS installed it: sync_venv succeeds into a new
dependency environment that only activates at process boot, and
ensure_import raises "installed; restart Hermes". A lazy-install guard
("this process is not running from the install's dependency environment")
was flattened the same way. Users were told to install something that was
installed, or given a command that doesn't address the actual refusal.
Keep the import as the decider, but remember the InstallError and put its
text in the ImportError. bedrock_adapter._require_boto3 had the identical
shape; azure_identity_adapter already propagates str(exc) and is the model.
Under the updater's claim, sync whenever dependencies are not current
rather than only when nothing is committed. The tail's own children
and post-sync verification children are current and stay no-ops. A
stale generation still committed from the previous Python pin is the
same ABI trap as the pre-PM venv, and it now syncs too. If the sync
still leaves the tree out of date, raise instead of relaunching into
another sync.
prepare_launch returned early for any process running under the
updater's own claim, so it would not re-run the completion tail. That
also covered processes the updater spawns before PM commits a
generation (a restarted gateway), which then booted with no
environment: previously on the pre-PM venv, now refused.
Under the updater's claim with nothing committed, sync the dependency
generation (carrying the legacy venv's extras, as the first sync
always has), skip the tail since that belongs to the updater, and
relaunch on the store Python. The relaunched process sees the commit
and returns early as before, so the no-recursion guard still holds.
With nothing committed, activate_dependencies fell back to the in-tree
venv/.venv. After an update that venv was built for the old interpreter
(uv CPython 3.11) while the process ran PM's store Python 3.14, so every
compiled module in it was unloadable: the messaging gateway's Group Chat
worker died on `No module named 'pydantic_core._pydantic_core'` until PM
committed a generation ~40 minutes later and deleted the old venv.
committed_venv() returns the committed generation or a sealed payload's
environment, never the in-tree venv. Boot activation and child
activation environments use it. With nothing committed, a venv/Nix
interpreter keeps its own packages; PM's bare store Python refuses with
the repair remedy instead of running on inherited paths. selected_venv
keeps its contract because pre-PM updaters import it after the swap.
powershell -Command joins trailing argv into the script text, so $args
was never populated and the path became a stray token (ParserError).
An env var needs no quoting.
A hard-killed backend leaves its spawn-ledger record and its published
session token behind. Since the attach path started adopting published
tokens (f1247d2e01), such a record passed the token rung and then had its
refused port polled for the whole 45s readiness budget. The renderer's boot
timeout is the same 45s and starts first, so it gave up just before main
fell through to spawning, and every Retry/Repair killed the fresh backend
and produced another dead record: the app never started after an update.
A ledger record is written only after its backend binds, so a refused
connection on one means the process is gone. Readiness gains an opt-in
`alreadyBound` that fails at once on ECONNREFUSED; the attach probe and the
attached-backend liveness monitor set it. Remote/SSH readiness keeps
polling, since a tunnel still coming up refuses legitimately.
The extra removal left CI, the Docker image and the nix package still
requesting `hindsight`. Once the extra is gone, `--extra hindsight` and
extraDependencyGroups = [ "hindsight" ] ask for something that no longer
exists. Drop them the same way 73c598e319 originally did: remove it from
the CI extras lists, the Docker sealed-venv build and the nix default
groups, and point the nix examples/check at honcho. Also remove the
stray blank line left in the exclude-newer table.
4b7229d612 commits the default-brand icons, so build_source_web no
longer runs generate-icons.mjs. It updated test_source_build.py but
not test_web_ui_build.py, which uses the same source_products fixture.
The web build tests on main now expect a step that never runs.
record() looked up stamp['identity']['windowsExecutableName'], a key no
stamp writer emits, so every Windows record on a runner resolved no
executable and failed. The built package's Application/@Executable
names the exe; match it in the unpacked dir only, which also skips
before-pack's .bak rollback copy.
Importing a script from `node -e` leaves process.argv[1] undefined, and
pathToFileURL(undefined) throws on import. The bootstrap installer's
signing step imports batch-sign-binaries.mjs this way, and it crashed
when that pulled in sanitize-pe-signatures.mjs.
Hermes-Setup has never had a CI build; every published copy was built by
hand. This workflow_dispatch lane builds the Windows x64 exe (signed via
the desktop MSIX's Azure Trusted Signing path, batch-sign-binaries.mjs)
and the macOS arm64 dmg (Developer ID signed, app + dmg notarized and
stapled) from any ref, and uploads them as run artifacts. Nothing is
published. No caches: no actions/cache or setup-node cache, and fresh
npm/cargo/electron-builder cache dirs.
Every bootstrap downloads the script fresh, commit pins included; the
on-disk file is only the -File target for this run. Removes the reuse
path, ScriptSource::Cached and the in-place BOM upgrade of old caches.
A 429 from raw.githubusercontent.com made the installer fall back to a
cached pre-migration install-main.ps1. Its repository stage then
fast-forwarded the checkout to live main, whose lock only supports
Python >=3.14, so the old script's 3.11 venv got no core dependencies
and every install died at 'Baseline imports failed'.
The checkout follows the live branch, so the script must too: a failed
refresh of a mutable pin is now fatal (Retry re-downloads). Immutable
commit pins keep permanent cache reuse.
A source checkout without hermes_cli/source_check.py predates release
channels, so the only line it can be on is git. The desktop treated the
missing probe as unsupported and parked the user on a manual
`hermes update --help` card ("This checkout predates desktop
source-channel checks"), so an older non-bundled install could never
update itself from the app.
Report such a checkout as tracking main with an update available and let
apply take the normal git handoff. That update brings in the probe, so
later checks resolve normally. A probe that exists but fails still throws.
User installs failed with 'resvg-py is missing' because the web/desktop
source builds rendered icons on whatever python was on PATH. The default
brand outputs are now committed; source_build, apps/desktop build.mjs and
the npm/docusaurus pre-hooks consume them directly. Flavored release
bundles (canary/commit) still render into their own product dir.
icons-freshness-check now regenerates and fails on any byte diff.
A non-elevated Windows ARM64 install threw "run setup-hermes.ps1 in an
Administrator PowerShell" whenever Visual Studio ARM64 C++/Clang were
missing. Interactive runs now launch the signed VS installer through a
UAC prompt; CI, ssh and scheduled runs keep the explicit instruction.
A bare `pm install` fetched agent-browser (~200 MB with Chromium) before
the venv sync, so a Windows ARM64 machine without build tools downloaded
it and then failed preparing the native build. Install defaults only once
the venv (and its build tools) succeeded.
The e2e lane at 4 workers still starved the PTY turn and serve-SIGTERM
deadlines intermittently; drop it to 2. The upgrade lane had no cap and
ran cpu_count real-updater trees at once; cap it at 4.
activate() now tells check() whether to include the venv verdict; the
no-argument stand-in raised TypeError, which startup swallows, so the
check was never counted.
activate(allow_incomplete=True) discarded the venv verdict but still
computed it, and venv_is_current reads plugin selection through the
application config reader (ruamel). The update child runs the bare
bootstrap interpreter, so `hermes update` failed with "No module named
'ruamel'" once it published tools before the sync.
prepare_launch finishes an interrupted update, or a hand-run git pull, by
syncing the venv at startup. It skipped the required-tool step that
`hermes update` now runs first, so a bumped ripgrep/ffmpeg/python pin stayed
uninstalled and activation warned on every start.
A normal `hermes update` only re-synced the venv. The sync pulls uv/python
in through its own dependency, but a ripgrep/ffmpeg/node/npm pin bump in
pm/lock.json was never installed, so PATH activation warned and skipped the
managed tool dirs on every CLI start and gateway boot. Only the takeover
route for historical releases ensured the tool roots.
Move the takeover's loop into pm.client.ensure_tools_for_sync() and call it
from both routes before the sync: update_completion._prepare (the CLI and
Desktop route, running from the new tree so the new lockfile applies) and
_update_takeover.prepare. It uses explicit=True like the takeover (an update
is an explicit user action) and a failed download fails the update.
post_update.step_provision_runtimes / MACHINE_STEPS stay: `python -m
hermes_cli.post_update --scope machine` and tests still reference them.
The ffmpeg docstring no longer claims that step re-ensures it.
The driver passed --skip-browser whenever the installer's help listed it.
That was for historical installers' own Playwright step; on a PM installer
the flag is now a real opt-out and would hide the default users get.
The flag used to exit 1 as retired. Now that PM installs the browser tools
by default, it maps to `pm.cli install --without agent-browser`, which later
installs and `hermes update` honour.
A source update re-syncs only the venv, so installs created before
agent-browser became a default would never get it. Skipped for sealed
payloads and when lazy installs are disabled; honours the recorded
opt-out; a failed download warns and never fails the update.
Browser tools find agent-browser only in PM's store or on PATH and their
readiness check never installs it, so a fresh install silently had no
browser_* tools. Package.default marks an optional package that a bare
`pm install` also carries; a failed download of it warns instead of
failing the install. `pm install --without NAME` records the opt-out in
declined-packages.json beside PM's install state, and naming the package
explicitly clears it. agent-browser and chromium gain a Termux gap: Termux
owns its browser stack and there is no bionic Chromium.
The ubuntu-latest-32-* larger runners boot ubuntu-22.04 (glibc 2.35), and
the prebuilt llama-server in the native payload needs GLIBC_2.38, so the
staged-entry verification fails. Restore the two entries once the runner
images are updated.
The compaction scenario used an absolute threshold of 21K estimated
tokens. That value assumed the default tool surface. On the Linux CI
runner, main advertised 15 browser tools because npx and system Chrome
exist there. Those schemas added about 2.7K tokens, so the turn crossed
21K. The PM branch finds agent-browser only through PM or PATH (no npx
fallback), so the browser tools are absent and the turn peaked near 14K.
Compaction never ran and the summarizer assertion failed.
The turn now runs with --toolsets file, so the prompt size does not
depend on host tools. The threshold is 5.5K to match that prompt.
The ubuntu-latest-32-* larger runners boot ubuntu-22.04 images, whose
system python3 is 3.10. The bundle bootstrap imports tomllib (3.11+), so
linux-arm64 failed before doing any work. Pin the interpreter instead of
trusting the image.
The dashboard PTY test failed on this branch with "node: not installed and
lazy installs are disabled". On this branch _make_tui_argv asks PM for node
(hermes_cli/main_tui_launch.py _tui_node_bin). The E2E sandbox forbids lazy
installs (HERMES_DISABLE_LAZY_INSTALLS=1 in tests/e2e/core/dashboard
hermetic_env). It gets node only when the runner's PM store has it
(tests/e2e/core/_pm_dependencies.py). The e2e job got node from
actions/setup-node, so node was on PATH but not in the PM store. main passes
the test because main's TUI launch takes node from PATH.
Install the locked toolchain (setup-pm toolchain: all) before npm ci, in
place of setup-node. The locked Node 26 also satisfies the boot-contract
suite. When HERMES_E2E_REQUIRE_TUI=1, the fixture now fails at once if the
store has no node, not later with an opaque 1011 close.
KittenTTS 0.8.1 requires misaki[en]>=0.9.4. PyPI's misaki 0.9.4 caps
Python below 3.13. NousResearch/misaki f03fd2be73 is upstream main with
the cap raised to <3.15 and no code change, so the extra declares
misaki[en] at that commit and uv resolves KittenTTS's requirement to it.
misaki[en] also lists spacy-curated-transformers. The newest release
that fits spacy 3.8, 0.3.1, caps Python below 3.14, so PM's `uv pip
check` rejected every 3.14 venv that carried it. Nothing imports it:
only transformer spaCy models use it, and KittenTTS never builds a
misaki G2P. A never-true override marker removes it from the lock.
torch, transformers, spaCy and num2words stay: misaki.en imports them
at module level and KittenTTS imports misaki.en.
Native bundles sync with --all-extras, so un-gating the extra would put
KittenTTS, torch and spaCy in every payload. [tool.hermes] opt-in-extras
names the extras that only an explicit selection installs. PM adds
--no-extra for each of them to an all-extras sync, so bundles and their
recorded feature list leave kittentts out, and setup still installs it
through sync_venv when the user picks it.
The extra is gated on platform instead of Python: onnxruntime has no
Intel macOS wheel and torch has no Windows ARM64 wheel. Nix gets the
hatchling build backend the misaki git source needs.
Verified on macOS arm64 with CPython 3.14.7 and uv 0.12.3:
uv sync --frozen --extra kittentts from this lock, then uv pip check
reports all 126 packages compatible. _generate_kittentts writes a
3.87 s 24 kHz WAV (RMS 3661), and no spaCy model is downloaded. A dry-run
all-extras sync with --no-extra kittentts selects none of kittentts,
misaki, torch, transformers or spaCy. The new build test fails without
the environment change and passes with it.