979 Commits
Author SHA1 Message Date
Saul Moro 488c3074cf fix(learnings): publish what an older import --from-mr left in the checkout (#823) (#838)
* fix(learnings): publish what an older import --from-mr left in the checkout (#823)

Item 7. import --from-mr in 0.25.0 to 0.26.0-beta.3 wrote
learnings/<date>-<title>.md, with source_mr in its frontmatter, into the
learnings checkout and never committed it. Nothing published it. In single-repo mode it also kept
`git worktree remove` from removing the checkout an older teamai left in
.teamai/, so every pull and contribute stopped on CheckoutRefusedError.
publishQueuedLearnings now takes the sync lock first, and under it, before
listing the queue, queues every untracked file of exactly that shape
(directly under learnings/, date name, source_mr), in the active namespace
and with contribute's name, then deletes the original. It finds the one
checkout this repo registers for the branch (git worktree list), so the
shared checkout and the old .teamai/learnings-wt are both covered and
another repository's never is. A file the branch or the queue already has,
by source_mr or by content, is deleted instead, and the warning names what
has it. A dry run touches nothing.

Item 21. The branch side of that duplicate check was the checkout's own
tracked files. In single-repo mode the checkout is often the old
.teamai/learnings-wt, which nothing syncs any more, so a teammate's later
import of the same MR was missed and the remnant went out as a duplicate.
When there are remnants, the check now also fetches origin/teamai-learnings
(best effort) and reads what origin has that the checkout's commit lacks.

Item 20. pull --dry-run published the queue: publishQueuedLearnings
honoured dryRun only for the remnants. It now stops after listing the queue,
and pull prints "[dry-run] Would publish N queued learning(s)" instead of
publishing or warning.

Maintenance sweep. publishLearningsMaintenance staged all of learnings/,
so a confidence write-back or a prune swept any uncommitted file into its
commit. confidence write-back, prune and promote now return the files they
wrote or removed, and only those are staged (a removed file git never
tracked is left out, since naming it would fail the add). That exposed a
second bug:
simple-git lists a staged rename under `renamed`, not `staged`, so a
`prune --archive` with nothing else to stage counted as nothing to commit
and was never published. commitAndPushAt now counts renames.

#814 follow-ups. drainCheckoutQueue is gone: the preAction migration moves a
checkout's queue before contribute and import --from-mr. Retire-only now
says "Retired <legacy> to <backup>: this project's data already lives in
<partition>"; a linked worktree lands there too, so "Finished an
interrupted migration" was wrong for it. config.yaml.*.tmp, the temp an
interrupted config save leaves (#831), is ignored in the single-repo and
project-scope .gitignore, and the single-repo self-heal adds it.

Item 15. After a failed refresh, readableReportsWorktree called ensure
without the reports lock, so it could create the checkout while a writer
that had just taken the lock created it too. It now refreshes once more
under the lock and throws the cause if that fails as well.

Item 17. init replaced the team clone before saving the new config, so an
init that stopped in between (an unknown --role, a busy queue lock) left
the old team's config.yaml beside the new team's clone. Just before it
clones another owner's repo, init now settles the old install as the final
save would (queue set aside, indexes dropped) and moves its config.yaml to
config.yaml.previous. A failed init then leaves no config, and commands ask
for teamai init.

* fix(learnings): address review — literal pathspecs, carry settings after a failed clone

Maintenance now stages exactly the files it names: commitAndPushAt and the
removed-file ls-files lookup pass --literal-pathspecs, so a learning named
with [ or * no longer stages the stray files it matches as a pattern.

init reads the config it set aside when the rerun finds none, so an init
whose replacement clone failed no longer drops enabledAgents,
disabledAgents, toolRoots and inheritUserScope on the next run.

* fix(learnings): address review — HTTP maintenance, agent lists on a plain rerun

An HTTP install's learnings dir is no git checkout, so the removed-file
ls-files lookup threw after a prune had already deleted the file. It now
returns the same non-fatal failed publish commitAndPush gives.

init without --agent keeps the carried enabledAgents and disabledAgents,
so a rerun after a failed replacement clone no longer reactivates tools
uninstall --agent excluded.

* fix(learnings): address review — retry maintenance a busy lock or failed push kept local, keep remnants while origin is unreachable

* fix(learnings): address review — queue remnants when origin has no learnings branch, never let one bad maintenance record or remnant block the rest

* fix(learnings): address review — dedup remnants against origin's tree, keep a maintenance record a read failed on

* fix(learnings): address review — commit only the published paths, not the whole index (#823)

* fix(learnings): address review — keep a staged file across the push-retry rebase (#823)

The path-limited commit leaves a file someone else staged in the checkout,
and git refuses to rebase with anything staged, so a non-fast-forward push
failed every retry. Snapshot it with git stash create around the rebase, as
syncWorktree does, and re-apply it with --index so it stays staged.

* fix(learnings): address review — read the queue for remnant dedup under the queue lock and ownership check; move a stale config aside when init reuses a clone (#823)

* ci: re-run checks (flaky dry-run-load-path test, unrelated to this PR)

* fix(learnings): address review — keep staged files staged when the snapshot restore conflicts; never publish a hand edit as a recorded maintenance run (#823)

* fix(learnings): address review — resolve snapshot conflicts from the snapshot without a reset, keep a conflicting staged file unstaged, no hand-edit warning for a merged maintenance commit (#823)

* fix(learnings): address review — point the unpublished-edit warning at git status (#823)
2026-09-28 10:43:32 +08:00
Ruifeng Xue 47438926fa fix: prevent stale Copilot rules from reverting team updates (#857)
* fix: prevent stale Copilot rules from reverting team updates

* fix: skip excluded agents during rule pre-push sync
2026-09-28 10:42:07 +08:00
Leo Camus 78477fa16a fix(review): print review reject/apply output in English (#836) (#859)
* fix(review): print review reject/apply output in English (#836)

teamai review --reject and --apply printed Chinese strings in their
console output: 已拒绝 for rejected, 应用失败 for apply failed, and
Chinese error reasons (缺失, 为空, 不存在, 不支持自动应用). Issue #836
established that CLI output must be English, and #840 fixed the same
class of bug in cache-cmd and import-local; review-cmd was missed.

Translate the six user-facing strings in applyOne's error reasons and
reviewCmd's reject/apply messages. Code comments are left as-is — they
are not user-facing. Add three tests that assert the English text and
no CJK characters in the output; all three fail before this fix.

* fix(review): update existing test to expect English output

The domain-drift apply test in review-cmd.test.ts still asserted the
Chinese string '不支持' which was translated in the previous commit.
Update the assertion to match the English output.
2026-09-28 10:41:32 +08:00
ydflowandydflow a8ab8e00f7 fix(dry-run): thread { dryRun } through the loaders contribute, session save and recall use (#853)
#837 threaded LoadOptions through the config loaders, and #850 fixes the
three commands that still reach them bare (pull, push, status). Three more
commands load their config before their own dry-run guard and pass nothing:

- contribute (--scope project loads, then --scope user / auto-detect)
- session save (same three branches)
- recall (detection, the inherited user scope, and the user branch)

On a config pending the legacy role migration each of them rewrote
~/.teamai/config.yaml under --dry-run, printing the migration line without
any [dry-run] marker; the auto-detect and project branches can also adopt a
pre-#546 partition and run the single-repo self-heal bootstrap.

loadLocalConfigForScope is the loader #837 missed: it now takes
LoadOptions and forwards them to detectProjectConfig and both
migrateLegacyRoleConfig calls. Callers that pass nothing behave as before —
a real run still migrates in place.

Co-authored-by: ydflow <ydflow@users.noreply.github.com>
2026-09-27 21:00:02 +08:00
3ab39d94ce fix(persistence): write state.json and the search index atomically (#854) (#855)
* fix(persistence): write state.json and the search index atomically (#854)

saveState, saveStateForScope and buildIndex wrote in place: open + truncate +
write. A crash, kill, ENOSPC or power loss mid-write leaves truncated JSON,
and the readers answer null for what they cannot parse:

- state.json: loadStateForScope falls back to StateSchema.parse({}) —
  lastPullRev and every per-checkout pushBaseRevs entry are gone without a
  word, so the next push compares against a stale base. That is the
  stale-base overwrite class the worktree fixes (#827) closed.
- search-index.json: loadIndex returns null, so recall silently loses the
  whole corpus until the next rebuild, and the shrink guard loses its
  baseline.

The repo already writes config.yaml (#831) and the votes file atomically for
exactly this reason, and writeJsonAtomic exists: temp file + rename,
preserving the target's permission bits. Move the three writers to it. No
reader changes; a successful write behaves as before.

* test: inject index-write failures at the write call, not the file mode

CI failed recall-rebuild-roots: the 'cannot be written' test chmod'd the
index file 0o444, which fails the in-place writer but not the atomic one —
writeJsonAtomic stages a temp sibling (the directory is writable) and renames
over the target, and rename needs only the directory's permission. Same for
the EISDIR test: renaming a file over a directory is EISDIR on POSIX but
EPERM on Windows.

Fail the fse.writeFile calls at the index path itself — the staged temp file
included — so the injection works for both writers on every platform. The
root-skip guard goes with the chmod it existed for.

* test: force the #812 e2e state-write failure at the write call, not the file mode

CI failed the 'cannot be recorded' case: it chmod'd state.json 0o444, which
stops the in-place writer but not the atomic one — writeJsonAtomic stages a
temp sibling and renames, and rename needs only the directory's permission,
so the push completed normally and exited 0. (The e2e config retries once;
the retry then failed teammatePublishes' git commit with 'nothing to
commit', the run's second error.)

The CLI runs as a subprocess, so the write call cannot be spied on the way
state-atomic-save does it. Inject the failure from outside instead: a
preload hook (NODE_OPTIONS --require) fails every fs.writeFile targeting
the staged temp of this state file inside the CLI process — the same seam
as the unit tests, on every platform.

---------

Co-authored-by: ydflow <ydflow@users.noreply.github.com>
Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-27 20:59:28 +08:00
ydflowandydflow d319816c6b fix(agents): carry the agent's model into the Cursor render (#830) (#856)
renderForCursor was the one renderer that dropped spec.model: Claude, the
Codex family and Copilot all write the agent's model into their native
file, and reverseFromCursor reads it back (COMMON_CURSOR_FIELDS whitelists
it), so a team agent with a concrete model ran on Cursor's default model
silently. #830's design notes name the gap ('Cursor drops model when it
renders agents today') and ask for it as a separate change — this is that
change: write the value verbatim, as the other renderers do. The
model[effort=...] form stays with the alias proposal.

Co-authored-by: ydflow <ydflow@users.noreply.github.com>
2026-09-27 20:58:12 +08:00
ydflowandydflow 84d8ba711f fix(dashboard): serialize events.jsonl append and compaction (#804) (#841)
* fix(dashboard): serialize events.jsonl append and compaction (#804)

events.jsonl was appended lock-free by every dashboard hook on the machine
and periodically rewritten by compactEvents (read -> filter -> temp file ->
rename). The atomic rename only prevented a torn file: an append that landed
between the rewrite's read and its rename was overwritten and lost, so the
dashboard's sessions, interventions and prompt counts under-counted until
the next rebuild. Two concurrent compactions also shared one fixed temp
name, so their rewrites could clobber each other's copy.

Both writers now take <events file>.lock (acquireLock from update.ts, the
pattern the usage file took for the same lost update in #790/#803): a hook
append waits up to ~250 ms, inside its foreground budget, and one that
gives up records its line in an events.pending-<uuid>.jsonl side file that
the next lock holder folds into the file before it writes. The line carries
a pendingId, so a fold never appends a side file twice, identical events in
separate side files are both kept, and no reader (readEventsRaw) and no
compacted file keeps the id. A compaction waits up to ~5 s for a peer's
rewrite and skips, leaving the file as it is for the next compaction, when
a live holder outlasts the wait; a lock whose owner is gone is reclaimed,
and a rewrite's temp copy a killed compaction left
(events.jsonl.<pid>.<hex>.tmp) is removed by the next one, which now also
preserves the file's mode.

The side files and the lock live in ~/.teamai/dashboard/ beside the log,
which no repository tracks.

* test(dashboard-report-scope): provide the events lock in the update.js mock

dashboard-collector now imports acquireLock and releaseLock from update.js
(#804), and this file's factory mock replaced that module without them, so
appendEvent lost every event to the missing export and 49 report tests
failed on CI. Follow the pull-* tests' mock shape: acquireLock resolves
true and releaseLock resolves undefined, keeping these tests lock-free, as
they ran before the lock.

* test(dashboard-collector): match the compaction read by file name

macOS resolves a file under os.tmpdir() through /private, so compaction's
realpath'd target never equals eventsPath() and the counterexample's
injection never fired: live-b was never appended and the test failed on
every macos job. Match the read by its file name instead, which the
realpath keeps.

* fix(dashboard): skip the events lock when compaction has nothing to do

The detached compaction after every append took the events lock even when
the log was below the threshold and no side files were pending — nearly
always. On a state directory a test teardown deletes concurrently, that
turned the no-op into lock-file churn racing the deletion, which CI caught
as an unhandled ENOENT on macOS. Read first, lock-free: below the
threshold with no side files the compaction returns without creating the
lock, the I/O profile of the pre-#804 compaction; a rewrite, or side files
to fold, still take it.

* fix(dashboard): keep a surviving side file recognizable, fold and classify in time order

Two review findings on #841:

A fold that appended its side file but could not remove it left the file
unrecognizable after the next compaction: the rewrite dropped the line's
pendingId, so the next holder could not tell the side file was already in
the log and appended the event a second time, double-counting it in the
dashboard's metrics. The id now stays in the raw file — a rewrite keeps
it, as the usage file's rewrite does (#788) — until the side file itself
is gone; readEvents drops it, so no reader ever sees it.

Side files also folded in readdir's arbitrary order, and the compaction
classified sessions from that raw order: a side file that outlived newer
appends could place an older event after them, re-mark a live session
stopped once its stopped-display window had passed, and the rewrite would
drop all its events. Side files now fold in their events' own time order,
and the compaction classifies in time order too — the order every reader
rebuilds from (dedupeEvents sorts); the rewritten file keeps the raw
order.

---------

Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-27 11:16:48 +08:00
JeffandCursor e0bf2e967d fix(lint): use String#endsWith in the Swift extractor (#846)
oxlint --deny-warnings flags /Error$/ as a warning, so every main CI run has failed since Swift support landed.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-26 21:59:02 +08:00
Saul Moro 46ffa96f2c feat(init): let a member choose the git provider with --provider (#789) (#844)
* feat(init): let a member choose the git provider with --provider (#789)

A member of a team on self-hosted GitLab had to configure GITLAB_TOKEN
even when they only sync and never need the CLI to open merge requests.
`teamai init <repo> --provider <name>` now uses the named provider
instead of detecting one, and records it in the member's local config.
PR/MR creation and doctor's provider checks prefer it over the team's
teamai.yaml, which stays unchanged, so other members keep detection.
With `git`, push pushes the branch and says the MR must be opened by
hand, as it already does for a provider: git team repo.

* fix(init): address review — guard --provider gitlab and keep --provider git out of teamai.yaml

--provider gitlab on a host with no configured GitLab instance would send
the token to gitlab.com (the API base defaults there); stop with a hint to
set GITLAB_URL or use --provider git. A teamai.yaml that init creates now
records the provider detected from the URL instead of a member's git
override, matching the docs.

* fix(init): address review — do not record git as the team provider on an unconfigured GitLab

With --provider git, a teamai.yaml that init creates (empty team repo or
first self-mode init) recorded detectProvider(url), which skips the
self-hosted GitLab probe. On an unconfigured instance that wrote
`provider: git` and cost every teammate automatic merge requests. Init now
resolves the team provider as it would without the flag, including the
probe, and stops with a GITLAB_URL hint when the probe finds GitLab.

* fix(gitlab): address review — refuse a TEAMAI_GITLAB_HOST that disagrees with GITLAB_URL

Repos on TEAMAI_GITLAB_HOST were detected as GitLab while the API base,
token included, came from GITLAB_URL. Stop before any request when the
two name different hosts, and let gitlabWhoami surface the configuration
error instead of reporting a failed login.
2026-09-26 21:43:22 +08:00
Saul Moro c7e9f909e5 test(dry-run): stop git background maintenance from racing the tree snapshot (#845) 2026-09-26 21:15:45 +08:00
Smilewithoutfalling ba04f18205 feat(code-knowledge): add Swift support to the AST and heuristic tracks (#842)
Registers tree-sitter-swift (already shipped inside the pinned
tree-sitter-wasms@0.1.13) with the captures walk.ts needs for types,
protocols, functions, imports, calls and conformance relations, and adds a
regex extractor so Swift facts still surface when the AST path is
unavailable. .swift is now collected, mapped to the swift language and
marked by the key-file patterns.

Swift imports are module-level, so they must not reach the tsconfig `paths`
mapping, which is TypeScript-only: doing so made `import Shared` resolve to
an unrelated .ts file and suppressed the EXTERNAL_IMPORT gap that records
the truth. Swift specifiers are reported as gaps instead, the way Go's
`import "fmt"` already is.

No new dependency; package.json is untouched.

Closes #712
2026-09-26 21:10:57 +08:00
Saul Moro f7da1bb8b6 ci(lint): add oxlint and fail CI on any warning (#828) (#839)
* fix(tags,roles): honor --dry-run in tags subscribe, tags unsubscribe and roles set (#836)

The three commands saved the local config and reset lastPullRev even
under the global --dry-run, which is documented as "Preview mode, no
changes made". Their siblings (tags add/remove, roles init/add/remove/
update) already return early with a [dry-run] message.

The roles set preview names the additional roles it would save,
including none, because a real run replaces the existing list.

oxlint reported the unused options parameter in tagsSubscribe and
tagsUnsubscribe; rolesSet has the same bug but reads options.add.

* chore(lint): add oxlint with its default rules

Pinned to an exact version so a new default rule arrives in its own PR,
not as a CI failure on an unrelated one. The no-unused-vars options
keep oxlint's _ ignore patterns and add ignoreRestSiblings, which the
rest-omit in dashboard.ts relies on to keep config and roots out of
/api/workspaces.

* style(lint): apply oxlint safe fixes

Drop redundant escapes in regex character classes and template literals,
empty-object fallbacks in object spreads (spreading undefined adds
nothing), and anchored regexes that are plain startsWith/endsWith
checks. No behavior change.

* refactor(lint): remove unused imports and an unused catch binding

Applied with oxlint --fix-suggestions and reviewed by hand. Every
removed whole import is a library module with no import-time side
effects.

* refactor(lint): remove dead code reported by no-unused-vars

Each hit was checked against its callers and git history; none is
missing wiring (the two that were, tags subscribe/unsubscribe, are fixed
in the preceding commit). Removed: unused locals and functions, the
options parameter of tagsList, rolesList and generateDigest (read-only
commands), the never-read interactive option of importFromRepo, the
empty test/e2e.mjs left over from the E2E migration, and a try/catch
that only rethrew. new Array(n) becomes Array.from. No behavior change.

* test(lint): fix lint hits in tests

- contribute dry-run test asserted nothing; it now checks that the run
  leaves the repo/HOME tree unchanged (verified to fail when the dry-run
  early return is removed).
- Drop a no-op expect(result).not.toThrow on a string.
- Keep undefined in two optional-chain casts so a regression fails the
  assertion instead of throwing a TypeError.
- Remove unused locals, helpers and imports; new Array(n) becomes
  Array.from.

* refactor(lint): write control-character classes as \p{Cc}

no-control-regex flags literal control ranges. \p{Cc} names the same
set (C0, DEL, C1) and reads as what it means. Checked against the old
classes on every code point from U+0000 to U+10FFFF: manifest-schema and
agent-format match exactly, and contribute-check's normalization
pipeline produces the same output. The test assertion is now stricter
and checks every control character the sanitizer removes.

* refactor(lint): remove disable directives for rules that are not enabled

Four eslint-disable comments named rules this repo never ran
(no-await-in-loop, @typescript-eslint/no-explicit-any), so they
suppressed nothing.

* ci(lint): fail CI on any oxlint warning (#828)

npm run lint runs oxlint --deny-warnings and runs before the type check
in both GitHub Actions and Coding CI. The repo is at zero warnings, so
new code must stay clean. --report-unused-disable-directives also fails
on a disable comment that suppresses nothing, so a suppression cannot
outlive the code it was written for. CLAUDE.md, AGENTS.md,
CONTRIBUTING.md and the PR template list the command so contributors
and agents run it before opening a PR.

Closes #828

* chore(lint): pin oxlint 1.16.0, the newest release that accepts Node 20.0

oxlint 1.17.0 and later declare engines.node ^20.19.0 || >=22.12.0,
while the repo supports Node >=20. 1.16.0 declares >=8, supports
--deny-warnings and --report-unused-disable-directives, and reports 0
warnings on this branch.

* Revert "fix(tags,roles): honor --dry-run in tags subscribe, tags unsubscribe and roles set (#836)"

This reverts commit 224d459c99.

* refactor(lint): mark the unused options parameter of tags subscribe and unsubscribe

With the #837 dry-run fix reverted out of this PR, both functions no
longer read options. The underscore prefix keeps the signature and call
sites unchanged, so #837 can rebase onto it by renaming the parameter
back.

* chore(lint): restore oxlint 1.85.0

This reverts commit e2347efa. oxlint is a devDependency, so its Node
requirement (^20.19.0 || >=22.12.0) never reaches users installing
teamai-cli, and CI's node-version 20 resolves to the latest 20.x.
Staying on 1.85.0 keeps the #836 warning counts and the planned
type-aware follow-up on the same version.

* docs(contributing): note the Node version npm run lint needs

* docs(agents): note the Node version npm run lint needs
2026-09-26 19:56:03 +08:00
Saul Moro a47bb7eabc fix(cache,import): print cache and import review output in English (#836) (#840)
* fix(cache,import): print cache and import review output in English (#836)

CLI output must be English, but teamai import --cache-status and
--cache-gc printed Chinese headings, totals and lists, gcCache recorded
Chinese skip reasons that --cache-gc prints (and --json returns), and
the interactive import review printed a Chinese card, edit prompt and
write results.

Translate those user-facing strings and add tests that assert the
English text and no CJK characters in that output. The AI
classification prompt, the slug regex (which keeps CJK titles) and
code comments are unchanged.

* fix(cache): address review — print verbose cache debug output in English
2026-09-26 19:26:27 +08:00
Saul Moro be87a57534 fix(tags,roles): honor --dry-run and count namespaced skills in tags list (#836) (#837)
* fix(tags,roles): honor --dry-run in tags subscribe, tags unsubscribe and roles set (#836)

The three commands saved the local config and reset lastPullRev even
under the global --dry-run, which is documented as "Preview mode, no
changes made". Their siblings (tags add/remove, roles init/add/remove/
update) already return early with a [dry-run] message.

The roles set preview names the additional roles it would save,
including none, because a real run replaces the existing list.

oxlint reported the unused options parameter in tagsSubscribe and
tagsUnsubscribe; rolesSet has the same bug but reads options.add.

* fix(tags): count namespaced skills in tags list (#836)

The "N skill(s) have no tags and are always synced" hint counted the
top-level directories under skills/, so a namespace counted as one
skill and its skills were not counted at all. It also subtracted the
number of tags.yaml entries instead of checking which skills have tags.

Count the untagged skills among those pull delivers, using pull's own
resolver (resolveDesiredSkills with the member's role context). Pull
delivers every skill in the member's namespaces, all of them without
roles, whatever its tags; tags only add skills from elsewhere. So roles,
exclusions and a name held by both the root and a namespace are counted
the way pull counts them.

* fix(tags): address review — warn instead of aborting tags list on a malformed manifest (#836)

* fix(tags,roles): address review — keep the legacy role migration in memory under --dry-run (#836)

* fix(tags,roles): address review — preview the self-mode bootstrap and partition adoption under --dry-run (#836)

detectProjectConfig now takes the { dryRun } load option that roles set,
tags subscribe and tags unsubscribe already pass. Under it, a fresh
self-mode clone gets the config bootstrapSelfRepo would write, kept in
memory (previewSelfBootstrap), and a pre-#546 partition is read where it
is instead of being renamed. One test hashes HOME, the project (with .git)
and state.json around each command on three fixtures.

* fix(bootstrap): address review — make no provider auth call in the --dry-run bootstrap preview (#836)
2026-09-26 19:25:05 +08:00
Saul Moro 49675a9787 fix(push): stop reverting a teammate's update from HOME or a stale .teamai copy (#823) (#835)
* fix(push): stop reverting a teammate's update from HOME or a stale .teamai copy (#823)

Item 4, user scope: push compared HOME's rules and skills with the shared
lastPullRev only, because the per-checkout push bases of #819 were keyed for
project scope alone. After a push synced HOME's unedited copy to a teammate's
R2, the next push compared it with R1 and offered it back over the teammate's
R3. A user-scope pull now records HOME under checkoutKey(HOME) in the user
state.json, and push reads and extends it like a project checkout's. The
user-scope fast path still reads the shared fields, and an install with no
record yet keeps comparing with lastPullRev without the unrecorded-checkout
refusal: HOME is the scope's only checkout, so that revision is its own.

Item 19, inherited user scope: a project pull with inheritUserScope rewrites
HOME's skills, rules and agents under lastInheritedPullRev without moving the
user scope's push bases, so the next user-scope push offered a teammate's
newer update back the same way. That pull now adds its revision to HOME's
pushBaseRevs, creating the record from lastPullRev if there is none, and
leaves the record's rev, lastPullRev and the fast paths alone.

Item 10, single-repo: the active tree's .teamai/rules and .teamai/skills are
push sources, and on a branch behind the default branch they hold older team
versions nobody edited, which push listed as modified. The isPastVersionOf
guard that held only placed rules now covers every .teamai/rules copy, and
.teamai/skills gets the same guard: a skill is skipped with a warning when
every team file whose copy differs is an older version of it. A team file
missing locally is a teammate's addition when the member's branch never added
it, and the member's deletion otherwise; member-only files are ignored, as the
equality check already ignores them.

The checkout-base resolution that push and the agents scan each repeated
(key, record, checkoutBaseRevs, lastPullRev fallback) is now one exported
helper, resolveCheckoutBases, next to checkoutBaseRevs in pull.ts; pull uses
the same key for its record.

* fix(push): address review — record HOME's push base in an upgraded install (#823)

An upgraded user-scope install with no HOME record synced HOME from R1 to R2
on its first push but saved no base, because push recorded one only when the
bases came from a record. A teammate's R3 then made the next push compare the
R2 copy with lastPullRev R1 and offer it back. Push now creates HOME's record
from lastPullRev (userScopeRecord, shared with the inherited pull) and adds
the revision its sync reached. An unrecorded project checkout still records
nothing, since its fallback base may be another checkout's.

skill-data: contribute-member explains the stale .teamai copy warning and
how to publish an edit of such a copy.

* fix(push): address review — keep HOME's inherited base and a partial pull's delivered base (#823)

* fix(push): keep the push bases of skills a pull held, and match a replaced root rule at every base (#823)
2026-09-26 19:24:28 +08:00
Saul Moro 81aa8ea63e fix(learnings): find learnings despite a broken manifest and in the dashboard; show the MR import prompt (#823) (#834)
* fix(learnings): find learnings despite a broken manifest and in the dashboard; show the MR import prompt (#823)

Three places that decide which learnings are found, and the MR import
prompt.

Recall index rebuild (item 12). When recall had to build a missing index,
one unreadable roles.yaml or projects.yaml failed the whole build:
deliveredIndexSources and resolveActiveLearningsNamespaces threw, the error
went to log.debug, and recall said "No learnings available. Run `teamai
pull` first", which pull does not fix. Each now runs in its own try.
Learnings do not depend on the manifests and are always indexed: a broken
projects.yaml leaves only the shared root, never every namespace. Docs,
rules and skills get empty lists (undefined would index the whole trees),
and one warning names the cause; a broken projects.yaml, which both read,
gives one warning for both. The partial index is saved like any
other, so later recalls stay quiet and the next pull rebuilds it whole. A
skills collision with no index to keep skills from indexed none silently;
IndexedSkills' keep-indexed now carries the conflict line and recall shows
it when there are no indexed skills to keep (no index, or an older one).
Any other build failure is shown with its cause, not as "No learnings
available".

import --from-mr duplicate check (item 9). The scan listed only each root's
top level, so learnings under learnings/<ns>/, where #825 files MR
learnings, were never compared. Its result fed only a "marking as
superseded" warning, and nothing stored or read LearningDraft.supersedes,
so the claim was false. The scan (now findOverlappingLearnings) also walks
the active project namespaces, as the index does (safe single segments,
first root wins per relative path), and the warning becomes a
possible-duplicate notice naming the files. `import` ran the extraction as a task,
with the logger silenced, so importFromMR's warning never reached the
terminal; the extraction now runs before the tasks (item 18), where the
logger prints it. The namespaces are resolved first: a broken projects.yaml narrows the check to the
shared root with a warning, as recall does, so --dry-run and --output keep
working. LearningDraft.supersedes and SUPERSEDE_THRESHOLD are removed. The
CI extractor still reads only the root (it has no LocalConfig).

Dashboard knowledge report (item 16). Its fallback index, built when no
search index exists, passed no learningsNamespaces, so project learnings
were missing in both scopes; in user scope it read learningsRoots().read,
which keeps another repository's learnings checkout in the write root
(#808). It now uses indexableLearningsRoots in user scope, as project scope
already did, and resolves the active namespaces with the paths. A broken
projects.yaml leaves the shared root, with a warning when the report builds
its own index.

import --from-mr prompt (item 18). importFromMR, which asks "Accept
learning? [Y/n]" on a readline, ran inside the first listr2 task. In a
terminal the default renderer holds stdout back while a task runs, so only a
spinner showed and the prompt appeared after it was answered. The
extraction now runs before the task list, which starts at "Publish
learning" with the extraction's result as its context.

* fix(learnings): address review — write the partial recall index past the shrink guard (#823)

With a manifest recall cannot read, the rebuild leaves docs, rules and
skills out on purpose. Against an older-format index of a full corpus the
result is under 20% of it, so buildIndex's shrink guard kept the old file
and recall searched the entries its warning said were left out. The
degraded rebuild now passes `partial`, which skips the guard; every other
build keeps it.

* fix(learnings): address review — skip the older recall index when the partial one cannot be written (#823)

* fix(learnings): address review — show queued learnings in the dashboard's fallback index (#823)

The dashboard's temporary index now reads the contribution queue, as recall's
does; promotion and prune candidates keep to the published roots. The recall
rule tells the agent to relay the skipped-older-index warning.
2026-09-26 19:23:46 +08:00
Saul Moro b136c9c654 fix(pull): do not deliver an env, hook or MCP entry with a mistyped key (#822) (#833)
* fix(pull): do not deliver an env, hook or MCP entry with a mistyped key (#822)

Item 1. Env, hook and MCP entry schemas are plain z.object, which strips
unknown keys, so a mistyped scoping key (`role:` for `roles:`) vanished and
the entry reached every member. Each reader now reports the keys an entry
was written with that its schema does not know (known keys come from the
schema's own shape), and keepScopedEntry does not deliver such an entry and
warns once, naming the file, the entry and the key, the same path the
removed `projects:` key takes. doctor's per-entry-key check is retitled to
cover it. `env add`/`env remove` and `remove mcp` keep such a key when they
rewrite the file; `remove mcp` edits the YAML document instead of
re-serializing the parsed servers.

Item 4. recall ended every result with a Chinese line; it is English now.

Item 2 is not a bug: tags reaching a tagged skill in an inactive namespace
is the behavior #337 added and roles-tags-pull tests. The design doc's
Known gaps entry now says so.

Item 3 (pull --dry-run warnings) is left to #832.

* fix(env): warn when env add updates a variable pull does not deliver (#822)

Updating a variable that carries an unknown key keeps the key, so the
variable stays undelivered; env add now says so instead of only reporting
'Updated env variable'.

* fix(pull): keep installed MCP servers and hooks when their file has no known top-level key (#822)

A hooks or MCP file with `server:` for `servers:` parsed as empty and
removed every installed team server or hook for every member, silently.
Such a file now fails like one that does not parse, naming the keys found
and the key expected. An extra key beside a known one is still ignored.
2026-09-26 19:17:22 +08:00
ydflowandydflow fe0378782f fix(pull): report hooks and MCP entry warnings on a dry run (#832)
* fix(pull): report hooks and MCP entry warnings on a dry run

`pull --dry-run` returned before the hooks and MCP reconcile stages, so the
warnings those stages raise never reached the maintainer who ran the dry run
to see exactly them: an unknown entry id, a per-entry `roles:` key, a
hooks.yaml that does not parse. A dry run must resolve and warn, then skip
the write (#822, item 3).

MCP already had the capability — `McpReconcileOptions.dryRun` gates every
write in mcp-reconcile.ts and `teamai mcp inject --dry-run` uses it — so
`reconcileMcpAllScopes` only forwards it. Hooks had no dry-run path at all,
so `reconcileTeamHooksForConfig` gained one: it resolves the entries, reports
what it found, and stops before the first write. `resolveTeamHooks` takes a
`preview` flag so the transparency line reads "Would apply N team hook(s)"
instead of claiming they were applied.

The tests drive `pull()` rather than the reconcile functions, the way
pull-env-shape-warning.test.ts does: the defect was in the orchestration
layer, so that is where it has to be pinned. They cover the dry-run warning,
the zero writes, the unchanged real-pull behavior, and the dryRun forwarding.

* fix(pull): preview-word the MCP dry-run summary and update the design doc

Review follow-up: with dryRun forwarded, the MCP summary still read as a
completed apply — "N change(s) ... Restart your AI tool session to load them"
after a run that wrote nothing. A dry run now reads "Would make N change(s)"
and omits the restart instruction; a real pull keeps the applied wording.

docs/designs/multi-project-management.md still said `pull --dry-run` prints
no hooks or MCP warnings; it now describes the resolve-and-warn behavior this
PR ships. The new tests pin the wording on each side of the dry-run boundary:
the dry-run case failed on the previous commit, the real-pull case passed.

* fix(pull): preview-word the hooks debug trail on a dry run

The [scope] "Reconciled N team hook(s)" debug line still claimed a reconcile
after a run that wrote nothing. It now reads "Would apply N team hook(s)" on a
dry run and keeps "Reconciled" on a real pull; both wordings are pinned by the
dry-run tests.

* fix(hooks): report the Pi skip on a dry run, like a real reconcile does

The per-entry `tools: [pi]` case: a real pull warns "Pi supports built-in
lifecycle hooks only; skipping N custom team hook(s)" during the per-tool
pass, but the dry run returned before that pass, so its "Would apply" line
promised hooks no tool will run.

Extract the Pi report from the loop so both paths print the same warning —
the dry run gated on the same toolPaths and agent filters the real pass uses
— and pin it with a test that is RED without the fix.

---------

Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-26 14:43:13 +08:00
Saul Moro 7c834ce428 fix(data-layout): let every self-mode worktree publish learnings and keep its queue (#808) (#814) 2026-09-26 00:29:00 +08:00
Saul Moro 87a606b727 fix(config): keep config.yaml readable while it is being saved (#823) (#831)
config.yaml was rewritten in place, so a command that read it mid-save saw
an empty file ("Invalid project config"), and a failed write left it
truncated. The three local config writers (saveLocalConfig,
saveLocalConfigForScope, the legacy role migration) now go through
writeFileAtomic: a sibling temp file renamed over the target, removed on
failure. The partition config.yaml already used it.

An existing config.yaml keeps its mode; a newly created one is 0600 (was
the umask default, usually 0644), as the partition config already is.

writeFileAtomic now writes a symlinked target at the end of its link chain
(temp file next to that file), so a symlinked config.yaml keeps its link
instead of becoming a regular file. A dangling link gets its missing target
(and directory) created, as the in-place write did; a link loop is refused
with an error and nothing is written.
2026-09-25 21:46:12 +08:00
Saul Moro e79db174c4 fix(push): stop offering a teammate's update back as a local edit (#823) (#827)
Three more ways push could list a copy the member never edited as modified,
ready to send a teammate's change back as the old version.

Single-repo mode (item 2). Push runs against a knowledge worktree whose team
root is <wt>/.teamai, a subdirectory of the git repo. The pre-push sync read
each base version with `git show <rev>:rules/x.md`, which git resolves from
the repo root, so it never found one, and every rule or skill a teammate
updated read as a local edit. The three reads now pass `./<path>`, which git
resolves from the working directory, as getFileContentWhenAdded and the agent
guard already did.

Placed agents (item 3). An agent placed with --role/--project is held when it
changed on the team since this machine's copy was current, and "current"
meant the version at the shared lastPullRev, which a pull in another checkout
moves past a copy a stale worktree still holds (the #812 revert, for agents).
The guard now reads this checkout's bases through checkoutBaseRevs, and falls
back to the shared lastPullRev for a checkout with no entry, as the pre-push
sync does. Push bases record where the sync moved rules and skills, not
agents, so the copy stays at the revision pull delivered: the guard holds an
agent that differs from its version at any base. Push records the team HEAD
as a base before the scan, and the file there is always the current one, so
the version the agent was added with is compared too whenever a base
predates it; otherwise a placement that landed after the last pull would go
back over a teammate's later edit. The hold message now says "this checkout".

Skill copy (item 5). The sync overwrote a local skill in place, so a copy
that failed partway left files from two revisions, matching no base, and the
next push listed the skill as modified. The update is now built in a hidden
sibling (the local copy, then the team version over it, so files only the
member has survive as before) and renamed into place; a failure leaves the
previous version whole. The stage carries the local modes, so cleanup makes
a read-only stage writable before removing it, and warns with the path if a
leftover cannot be removed; if the previous version cannot be renamed back,
the error names where it is.

Item 4 (user-scope push base) follows once #814 is merged.
2026-09-25 20:15:59 +08:00
Saul Moro 21cb76aa49 feat: one namespace model for every resource type (#707) (#816)
* chore: start one namespace model for every resource type (#707)

* refactor(pull): check agent and skill namespace collisions with one resolver (#707)

Add src/namespace-resolver.ts, the pure rule tickets 02-05 build on: an
active namespace item replaces the root item of the same name, and a name
twice in one place or in two active namespaces is a tagged conflict naming
both sources. The result depends only on the active order, not read order.

Agents and skills now run their duplicate checks through it and throw the
same messages. Agents still treat root + namespace as an error.

Add fast-check for the resolver's property tests.

* feat(env,hooks,mcp): scope env, hooks and MCP servers by namespace (#707)

env/<ns>/env.yaml, hooks/<ns>/hooks.yaml and mcp/<ns>/mcp.yaml are read where
<ns> is active in resources.env/hooks/mcp; a namespace entry replaces the root
entry of the same key, hook id or server name. A broken active file, a name
twice in one file or in two active namespaces stops that type for the run and
keeps what is installed. Per-entry projects: (and roles: on env) reach nobody;
roles: on hooks and MCP keeps filtering with a deprecation warning. Unknown
resources: keys warn instead of failing the manifest.

* feat(pull): let an active namespace item replace the root item for skills, agents, rules and claudemd (#707)

With a role or project configured, an item in an active namespace now
replaces the root item of the same name, whole:

- agents by stem: root + namespace is no longer a duplicate error, and a
  recorded (placed) agent replaces the root agent too
- skills by name, including a root skill received through a tag; an
  install removes the files of the version it replaces
- rules by first-level file name, in tool dirs and Hermes' SOUL.md block
- claudemd files by name in the managed block

Two active namespaces with one rule or claudemd name stop that type for
the run and keep what is installed. Push writes an edited overridden
skill, agent or rule back to its namespace, and the skills push scan
covers role and project namespaces. The placement record is withdrawn
by a same-name shared-root file only in legacy mode. Recall indexes the
skills pull delivers. doctor lists overrides, and in legacy mode repeated
names, as notes. Legacy mode is otherwise unchanged.

* fix(pull): deliver both namespace rules and claudemd files of one name (#707)

Rules and claudemd have no namespace-vs-namespace conflict: each
namespace rule keeps its own local path and each claudemd file its own
place in the block, so two active namespaces with one first-level name
are both delivered, as before. Only root suppression applies.

doctor override and legacy repeated-name notes now use the same wording
as the env, hooks and MCP ones. The usage guide and admin reference say
to keep overridable shared content at the root, with an example.

* fix(pull): stop only skills or agents on a namespace collision (#707)

Two active namespaces with one skill name or agent stem used to throw
and abort the whole scope, so rules, env, docs, cleanup and the search
index were skipped too. resolveDesiredSkills, resolveDesiredAgents,
scanRoleAwareSkills and filterAgentsByNamespaces now return a tagged
conflict. pull warns, leaves that type as installed (no install, no
inactive-namespace sweep) and syncs the rest. doctor reports the
collision as before; recall indexes no skills while it stands.

* feat(env,hooks,mcp): namespace flags, origins in status and doctor, docs (#707)

env add/remove take --role/--project; remove mcp searches every file and asks
for --role/--project when several define the name; push picks up
env/<ns>/env.yaml. status, list and doctor show where each entry comes from;
doctor lists overrides as notes and per-entry roles:/projects: as one
informational check. Usage guides and the admin skill reference describe the
namespace files; the #668 e2e moves onto them.

* test(env): show a broken env file stops env only (#707)

* fix(remove): remove an MCP server from the root file by default (#707)

remove mcp <name> follows push's convention: mcp/mcp.yaml when it defines the
name, else the one namespace file that does; --role/--project pick a namespace.
Only a name several namespace files (and not the root) define is refused.

* test(remove): expect namespace files in sorted order (#707)

* feat(docs): deliver a declared docs namespace only where it is active (#707)

A docs/<ns>/ that any role or project lists under resources.docs now
reaches only members with that namespace active; an undeclared
docs/<dir>/ stays shared. Leaving a namespace removes its local docs that
are byte-equal to the team copy and keeps edited ones with a line.
team-codebase is rejected as a docs namespace. The search index (pull,
recall, contribute) and doctor's "Team docs delivered" use the same set.

* feat(models): scope team model profiles by namespace and bind keys to their gateway (#707)

models/<ns>/models.yaml, declared under resources.models, replaces the root
profile with the same id while <ns> is active. A team API key is stored per
profile id and base_url origin, so pull never writes a key next to a gateway
on another origin; it prints the models switch line instead. Conflicts and
broken files stop model updates for the run. models list and doctor show
where each profile comes from; push validates every models file.

* docs(models): describe model profiles by namespace and key binding (#707)

* docs: list docs among the axes declared by hand (#707)

* docs: describe one namespace model for every resource type (#707)

Rewrite the multi-project design doc's precedence section for
namespace-over-root, add the per-type conflict, failure and legacy-mode
rules, and replace the per-entry key rows in the product overview. The
JSON doctor notes now also carry namespace notes.

* docs(changelog): replace per-entry scoping with namespace files (#707)

Drop the beta-only per-entry projects: entry, add the namespace axes,
the override, the per-type failure policy, the roles: deprecation on
hooks and MCP, and the upgrade-every-member-first note.

* docs(changelog): say the model key binding re-keys once and affects only betas (#707)

* refactor(namespaces): one warn-once registry instead of the quiet flag (#707)

Namespace fallback warnings, entry notices and unknown resources: keys
now go through utils/warn-once, reset once per pull, so the quiet option
threaded through eleven signatures is gone. The one-line wrappers
resolveTeamEnv, resolveTeamMcpServers and resolveTeamProfiles are removed;
every caller uses resolveEntries/resolveEntriesFor with the type's reader.

* refactor(namespaces): shared entry-file helpers, no unsafe casts in new code (#707)

- listEntryFiles/entryFileAbsolutePath replace the per-type file listers
  in env, mcp and models and the repeated path joins.
- gatewaySuffix replaces three spellings of the gateway suffix.
- LATER_RESOURCE_TYPES and friends are named for what they are:
  HAND_DECLARED_RESOURCE_TYPES, HandDeclaredNamespacesShape.
- Error messages use instanceof Error; manifest role/project ids are
  narrowed instead of cast; mapResources builds a typed object; pull
  writes env through an EnvHandler instance instead of a cast.
- status keys counts by entry type; rules localNameFor reuses
  deliversEveryNamespace, whose false answer is now documented.

* refactor(desired): move the desired-set resolvers out of pull.ts (#707)

Commands must stay thin, and recall, contribute and doctor imported
pull.js only to learn what a member receives. The skills, agents, rules
and claudemd resolvers, RolePullContext and the index sources now live in
src/resources/desired.ts; root suppression, the override note and the
repeated-name grouping live in namespace-resolver.

- A skill or agent conflict is a tagged DeliveryConflict carrying the
  resolver's NamespaceConflict, rendered once by describeDeliveryConflict
  (wording unchanged); DesiredItems names the result union.
- DesiredItems keeps each override, so doctor no longer rebuilds skill
  and agent overrides by hand.
- recall and contribute share deliveredIndexSources; pull indexes through
  the same indexedSkills instead of a second copy.
- doctor: one unresolvableCheck for skills, agents and docs; the docs
  check reports an unreadable manifest instead of returning nothing; the
  namespace notes catch only the team-repo reads.
- docs withdrawal reuses utils pruneEmptyDirs.

* fix(pull): name both files in a skill or agent conflict (#707)

Story 8 asks for a message naming both files. A skill or agent conflict
named only the namespaces, and an agent defined twice inside one
namespace (agents/a/x.md next to agents/a/x.yaml) read as 'found in
active namespaces "a" and "a"'. The duplicate case now names its one
place, and both cases list the two files.

* fix(rules): only a delivered namespace rule replaces the root rule, in every tool (#707)

- The rules override ran before the tag filter, so a namespace rule the
  member's tag subscription excludes still suppressed the root rule and
  the member received neither. The tag filter now runs first.
- JoyCode, OMP, Pi and Copilot share their rule directory with the
  member's own rules, so the stale sweep deletes nothing there unless a
  tombstone names it: the root rule a namespace rule replaced stayed
  installed and both versions loaded (story 6). pullAllRules now removes
  such a copy while it is byte-equal to its render, as agents do.

* fix(recall): keep the indexed skills while a skills conflict holds them (#707)

On a skills conflict pull keeps the installed skills, but the index was
rebuilt with none, so recall returned none of the skills the member still
has. The index now keeps the skills entries it already held, and pull
does the same when resolving the skills fails.

* fix(hooks,mcp): fail hooks inject and mcp inject when team entries do not resolve (#707)

reconcileTeamHooksForConfig returned [] when the team hooks could not be
resolved, the same value as a team without hooks, so hooks inject printed
'Hooks injected into all AI tool settings' and exited 0 over a broken
hooks/<ns>/hooks.yaml. It now returns { ok: false }, and hooks inject
exits 1 after the warning that names the file. mcp inject said 'Already
up to date.' in the same case; the MCP reconcile now marks the result
unresolved and mcp inject exits 1.

* fix(models): bind a beta API key to the gateway it was sent to, once (#707)

- A key a 0.26.0 beta stored under team:<id> counted for whatever origin
  the root profile had now, so a root profile moved to another host got
  the old key written next to it. The first pull or models command that
  reads such a key now binds it to the origin TeamAI last wrote into the
  agents switched to that profile (the root's current origin when it is
  among them), else to the root's current origin, and never re-reads the
  unbound key. Pull then leaves a moved agent alone and asks for the
  switch.
- The 'switch to set a key' and 'no longer active' lines are written to
  debug.log too: SessionStart pulls run silent.
- Legacy mode, which reads no namespace, says a profile 'was removed'.
- A models command whose profiles do not resolve reports it with
  log.error and exit code 1, as env list and mcp list do, instead of
  throwing.

* fix(entries): keep 0.25 files that repeat a name under different roles: working (#707)

0.25.0 let hooks.yaml and mcp.yaml repeat a hook or server name under
different roles:, delivering every copy that passed the role filter (MCP
kept the last). The namespace resolver treated that as a duplicate, so a
member holding both roles, or a role-less member in a team with
projects.yaml, stopped receiving hooks or MCP entirely. During the
roles: deprecation window such a repeat is delivered as in 0.25; a name
repeated without roles: on every copy is still a duplicate.

Also restores the test that the role filter runs before
requireTeamScripts, so the transparency print lists only what will run.

* feat(doctor): say where each entry type's entries come from (#707)

The spec asks doctor, like status and the list commands, to show each
entry's namespace; doctor listed overrides only. For env, hooks, MCP and
models, a namespace contributing any entry now adds a note counting the
entries by origin, 'env: 3 received here (2 root, 1 checkout)', from the
describeOrigins that status uses.

* test(pull): env and hooks conflicts between two namespaces, and builtin: in a namespace file (#707)

Seam 1 asks every type to show, through pull, that two active namespaces
defining one name keep the installed state and name both files. Env and
hooks were covered only through doctor and the handler; so was the
warning for builtin: in a namespace hooks file.

* docs(skill-data): hooks and MCP edits are published with git, not teamai push (#707)

manage-admin.md told admins to publish hooks/MCP file edits with
teamai push, which sweeps only rules/, env/ and .codebuddy-plugin/, so
an agent following it would push nothing. It now says to commit and
push the file with git, as the usage guide does.

* refactor(models): read switched agents without a cast (#707)

* docs(changelog): doctor counts entries per namespace; inject fails on unresolved entries (#707)

* refactor(pull): drop imports the resolver move left unused (#707)

* fix(hooks): install the built-in hooks when the team hooks do not resolve (#707)

A first init or bootstrap whose team hooks did not resolve (a broken
namespace file, a clash, a duplicate id) installed no built-in hook, so
the session-start pull that heals the member never ran. Installed team
hooks are still kept; the built-in hooks are now installed where missing,
with the root file's builtin: overrides whenever hooks/hooks.yaml parses,
and with their defaults only in a tool with no teamai hook when it does
not. init and bootstrap say that the team hooks were not installed.

* fix(manifest): keep an unknown resources: key when roles and projects save (#707)

zod stripped the key this CLI only warns about, so a projects or roles
command run on this version deleted a newer CLI's type from the team
repo for everyone.

* fix(docs): withdraw a copy the team edited after delivery, not only an unchanged one (#707)

Withdrawing an inactive docs namespace compared the local copy with the
current team file only. A doc the team changed after the member received
it was then kept forever with a false 'you edited it' line. A copy equal
to an earlier team commit is what the mirror delivered, so it goes too.

* fix(rules): withdraw a replaced root rule edited in the same push, name a kept copy (#707)

In the JoyCode, OMP, Pi and Copilot rule dirs, a replaced root rule's
copy was removed only while it matched the current root rule. When the
admin edited the root rule and added its namespace override in one push,
the member's unedited copy stayed loaded beside the override, silently.
It is now also compared with the render at the last pull, and a copy
that is kept is named with the fix.

* fix(doctor): fail a check when team hooks or model profiles do not resolve (#707)

teamai status counts such a type as 0 and says to run doctor, but doctor
had failing checks only for env and MCP, so a duplicate hook id or a
two-namespace clash showed nothing there.

* fix(env): warn when --role names a namespace nothing declares (#707)

env add/remove --role <ns> wrote env/<ns>/env.yaml for a namespace no
role or project lists under resources.env, so the variable reached
nobody and nothing said so. The same applies to remove mcp --role.

* fix(env): find a changed namespace env file whose name is not ASCII on push (#707)

git ls-files quotes such a path by default, so it never matched the name
on disk and push skipped the change.

* fix(entries): an active env, hooks, MCP or models file that cannot be read stops the type (#707)

The readers folded every read error into 'file does not exist', so an
unreadable namespace file silently delivered the root entry in place of
its override. Only ENOENT is absence now; any other error is a broken
file, like one that does not parse.

* fix(entries): match env, hooks, MCP and models namespace dirs case-folded, as docs does (#707)

A declared namespace was joined onto the path as written, so with
env: [checkout] and a directory env/Checkout/, macOS and Windows members
got the override and Linux members the root value.

* fix(doctor): split the legacy claudemd paths on '/', not path.sep (#707)

listFilesRecursive always joins with '/', so on Windows every path was
one segment and a claudemd/<ns>/x.md beside claudemd/x.md was never
reported.

* fix(recall): index the rules pull delivers, not the whole rules/ tree (#707)

A namespace rule replaces the root rule of its name, but recall, contribute
and pull indexed every file under rules/: the replaced root rule and the rules
of inactive namespaces came back from recall. Index the resolved rule set, as
docs and skills already do.

* fix(entries): write a namespace file into the directory pull reads it from (#707)

Pull matches a declared namespace to its directory case-folded, but --role and
--project returned the spelling typed. On a case-sensitive filesystem
`env add --project checkout` created env/checkout/, which shadowed
env/Checkout/ and dropped its variables from delivery.

* fix(remove): remove no MCP server by a bare name while an MCP file does not parse (#707)

The team scan skips a file that does not parse. With mcp/mcp.yaml broken,
`remove mcp db` took the one readable checkout/db as the target and removed
it. Refuse and name the file unless the readable root defines the name.

* docs: rules in recall, namespace writes and remove mcp on a broken file (#707)

* fix(remove): say the MCP file --role or --project names does not parse, not that the name is missing (#707)

The team scan skips a file that does not parse, so `remove mcp db --project
checkout` with a broken mcp/checkout/mcp.yaml reported "Not found". Name the
file and remove nothing.

* fix(skills): remove a leftover of another skill version only when it matches that version (#707)

Install removed any installed file at a path another team version of the skill
has, by path alone. A file a member added under that name, e.g. README.md
beside a namespace they never had, was deleted on every pull. Remove it only
when it is byte for byte that version's file; keep any other and name it.

* fix(env): edit no env file that does not parse, and no --project target after a failed refresh (#707)

env add and env remove read the target through parseEnvYaml, which answers an
empty list for a file that does not parse, then wrote that back: every
variable the file had was replaced. They now refuse and name the file.

--project resolves through manifest/projects.yaml. After a failed pull that
copy may be stale and name a namespace the project no longer uses, whose file
push would publish, so --project now changes nothing then. The root file and
--role do not depend on the manifest and still only warn.

* fix(pull): let no unusable namespace item replace the root one (#707)

A skill directory without SKILL.md replaced the root skill of its name:
install overlaid it and removed the installed SKILL.md as the other version's
leftover, while pull still counted the skill as synced. Such a directory is
no longer a skill; pull names it and keeps delivering the root one.

An agent file that does not parse delivers nothing, yet it still replaced the
root agent, and cleanup removed the unchanged root copy because no active
destination held that stem. The root agent now stays while its replacement
cannot be read or parsed.

* fix(push): take no namespace directory without SKILL.md for a member's skill (#707)

Pull stopped delivering such a directory in 6c7df8e3, but the push scan still
mapped the skill name to it. An unedited root skill then showed as modified,
and push wrote it into skills/<ns>/<name>/, deleting what was there and making
it a namespace skill that replaces the root one.

Also keep only a replaced root agent while its replacement does not parse, so
an unchanged copy from an inactive namespace is still removed, and put
renderedForTool's doc comment back on it.
2026-09-25 17:59:19 +08:00
Tremy.Wu 7d16175a77 feat(agents): support WorkBuddy subagent distribution (#826)
* feat(agents): support WorkBuddy subagent distribution

Register workbuddy as a supported tool for team agent sync:

- agent-format: add 'workbuddy' to ToolName/ALL_SUPPORTED_TOOLS,
  AgentSpec.tool_extras, and renderForTool; render/reverse in the
  Claude-style Markdown+frontmatter format with extras namespaced to
  tool_extras.workbuddy (mirroring the codebuddy pattern)
- types: default toolPaths.workbuddy gains agents: '.workbuddy/agents'
- agents: reverseByTool dispatches workbuddy to reverseFromWorkbuddy
- README (en/zh-CN/ja/ko/th): flip WorkBuddy agents cell to supported
- tests: workbuddy-agents.test.ts covers registration, render/reverse
  round-trip, and pull/push/remove via AgentsHandler

* docs(readme): keep WorkBuddy agents marked unsupported until product-side consumption lands

WorkBuddy does not yet consume .workbuddy/agents, so advertising the
capability as supported would make 'teamai pull' produce files the
product cannot discover. Flip the agents cell back to unsupported in
all five README locales; the CLI-side registration stays in place as a
latent capability and the cell flips back once consumption is verified.

Addresses review feedback on #826.

* chore: retrigger review (no changes)
2026-09-25 17:57:06 +08:00
Saul Moro f558b94614 fix(push): keep a teammate's update when pushing from a stale worktree (#812) (#819)
Before scanning, push syncs each rule and skill the member never edited to
the team repo's version, and "never edited" meant equal to the version at the
project's shared lastPullRev. state.json is shared by every worktree, so a
pull in another checkout moved that revision past the copy a stale worktree
still held: the unedited copy read as an edit, and push offered it as
modified, ready to send the teammate's change back as the old version.

Push now compares with the revision this checkout last synced, from its
lastPullByWorkspace entry (checkoutKey is exported from pull.ts), and falls
back to the shared lastPullRev for a checkout with no entry. For that entry to
survive, a pull at a new team revision no longer drops the other checkouts'
records: a checkout recorded at an older revision already misses the fast
path. When the pull finds lastPullRev cleared, it resets the other records
to an empty rev (FORCED_FULL_SYNC_REV), which matches no revision, so a forced
full sync reaches every checkout, single-repo mode included, while each record
keeps its push bases. Each full sync keeps only the records
of checkouts `git worktree list` still reports, and keeps them all when the
list comes back empty, so a removed or re-created worktree's entry does not
pile up.

The sync itself moves the unedited copies to the team repo's revision, so
push then adds that revision to the entry's pushBaseRevs (newest first, the
20 newest kept) and the next push compares with them; otherwise a copy synced
to R2 read as an edit against R1 once a teammate published R3. Push leaves the
entry's rev alone, since the pull fast path reads it and the checkout still
lacks that revision's docs and agents; the next pull rewrites the entry
without pushBaseRevs. The sync accepts a copy at any of pushBaseRevs or rev (a
skill only when all its files are at one of them), so a copy it left alone as
edited is synced again once the member undoes the edit, back to whichever
version a sync gave it. The base is recorded even when the sync stops
partway, which now warns, since the copies it did not reach still match an
older base; a revision push cannot save stops the push before the scan. A
checkout with no entry (a new worktree, or one last pulled by an older CLI)
can only sync against the shared lastPullRev, which may be another
checkout's or cleared: when the scan lists a team rule or skill as modified,
push stops before creating a branch and asks for a pull there, warning that
the pull replaces those files. A rule this machine placed does not count, and
config-only pushes and new resources go through.
2026-09-25 12:04:16 +08:00
dvd233 e7333778ff fix(push): keep config changes on explicit branch (#820)
* fix(push): address follow-up review findings

* fix(push): preserve config across reuse cleanup

* fix(push): handle config-only explicit branch with reused PR
2026-09-25 12:02:55 +08:00
yi111andYi-111-a 306e63f355 fix(push): preserve empty file additions (#821)
Co-authored-by: Yi-111-a <yi111@users.noreply.github.com>
2026-09-25 11:53:13 +08:00
Saul Moro c7723d652b fix(hooks): keep a removed worktree's hook events in its project (#810) (#824)
A hook resolves its scope from the payload's cwd, and resolveConfigForDir
answers the user scope for a directory that no longer exists. So once a
session's worktree was removed, its remaining events (tool_use, SessionEnd,
Stop) and skill uses were recorded under the user scope, which then counted
the session and reported the skills to its team, or were dropped when there
was no user scope. The project lost the session's last snapshot.

resolveHookConfig (dashboard-collector.ts) is the one resolver for the
hook dispatcher and the legacy dashboard-report, track and track-slash
entry points. For an existing (or absent) cwd it is resolveConfigForDir, as
before, and reads nothing else. For a cwd that is gone it reads this
session's last event that recorded a dataHomeKey, once per process, preferring
the events recorded at that same cwd (a detached Stop can run after the
session moved on to another repo), and
resolves the config at that event's projectAnchor, the main checkout,
which still exists (for a bare repo, whose anchor is the git directory,
at one of its worktrees that still exists). It uses that config only when it is still the scope the
recorded dataHomeKey names, so a worktree's own legacy .teamai never becomes
the main checkout's scope. If that config exists but cannot be read, the
event is dropped rather than given to the user scope (#748). With nothing
to match (no events, events from before #809 without an anchor), it is
today's answer.

The dispatcher's track and track-slash handlers now use the dispatcher's
config instead of resolving their own, and eventProjectAnchor gives an
event whose cwd is gone the session's last anchor, as process_exit does.
The legacy track-slash looks skills up under the resolved scope's tool
roots before the cwd's. The share reminder's gates (contribute-check on
Stop, pending-hint on the next prompt) ask about hookScopeDir, the
directory resolveHookConfig resolves from, so a removed worktree's
session gets the project's reminder settings, not the user scope's. The
legacy `teamai contribute-check` command gates the same way. The hook
session id has one implementation, deriveDispatchSessionId in
utils/session-id.ts, shared by the dispatcher and the event writers.
2026-09-25 11:51:40 +08:00
Saul Moro 4a65e3f676 fix(import): publish the learning import --from-mr extracts (#823) (#825)
`import --from-mr` wrote its learning into the teamai-learnings worktree,
then pushed with autoPushViaMR, which commits `.` in repo.localPath: the
knowledge clone, another checkout. That found nothing to commit, so the
learning stayed untracked on this machine and never reached the team,
while the command still reported the push step as done.

The draft now goes into the contribution queue, and a "Publish learning"
step calls publishQueuedLearnings, the path `teamai contribute` uses: it
commits and pushes the queue on teamai-learnings and drops an entry once
it is on origin. When publishing fails the learning stays queued, the step
says so, and the next `teamai pull` publishes it. As in contribute, the
queued file takes contribute's name (a random suffix keeps two learnings
with the same title and day apart), the recall index is rebuilt after the
publish attempt, the supersede check also reads the queue, and a read-only
(HTTP) source is refused up front instead of queueing a learning nothing
can publish; --dry-run and --output still work there. "Push changes via
MR" is left for the teamwiki update it was also for.

The learning also lands where contribute puts it: resolveLearningsSubdir
(now exported) picks learnings/<namespace>/ when exactly one active project
declares a learnings namespace, else the shared root. It used to go to the
root, where every project's members recall it.

Also, from the same follow-up issue:
- wiki slug: the main checkout's root takes its repo's name too, so one
  opened through a differently named symlink writes the same evidence as
  its worktrees. Subdirectories keep their own name.
- repo labels: a path is not qualified into a label a remote-form key
  already has (github.com/acme/api vs /x/acme/api), so the two no longer
  merge in `stats --by-repo`. The fallback is the repo's directory, so a
  bare repo's keys still share one row.
- local-agent tests use a session id unique per run: the hint markers are
  machine-wide files in os.tmpdir() keyed by session id, and overlapping
  runs deleted each other's.
2026-09-25 11:50:20 +08:00
Saul Moro c73d22147d fix(report): each scope reports its own dashboard sessions once, against its own snapshots (#785, #786) (#791)
* fix(report): each scope reports only the dashboard sessions recorded in it (#785)

Every scope read one machine-wide events.jsonl and picked its sessions out
by cwd prefix. The user scope excluded nothing, so a user-scope pull
reported every project's sessions (and, through the shared reported
snapshots, took them from the project's own report); Copilot sends no cwd,
so a project never reported its Copilot sessions; and a raw cwd under a
symlink or /tmp never matched the realpath'd projectRoot.

The hook now stamps each event's dataHome with the data home of the scope
the dispatcher resolved (the key the per-scope usage file already uses), and a
report keeps only its own scope's events, comparing realpath'd keys. A
project also owns its in-repo .teamai key, where hooks record until
migration moves it to a partition. Events written before this carry no
dataHome: a project keeps those whose realpath'd cwd is under its root, the
user scope never reports them. The log stays machine-wide for the
dashboard UI, stats --by-repo, session save and the contribute check.

Removes the excludeProjectRoots option, which pull only ever passed as []
(the user target exists only when no project config resolved), and the
projectRoot option now carried by selfConfig. The usage guide documents how
to remove by hand a skill an earlier release pushed into stats/<user>.yaml
from another project.

* fix(report): address pre-review findings (#785)

- Events record `dataHomeKey`, a hash of the realpath'd data home, instead
  of the path. A Copilot event persisted a workspace path through its data
  home (the raw root for a non-git project, the path-derived partition name
  otherwise), breaking the path-free Copilot contract from #666.
- A data home that no longer exists (an in-repo .teamai removed after
  migration) keys through its parent's realpath, so it still matches the key
  recorded while it existed.
- A non-git project's root is realpath'd before older events' cwd is matched
  against it, as the cwd already was.
- A key that is not a string (a hand-edited log) counts as absent instead of
  throwing and skipping the whole report.
- The legacy `dashboard-report` command's stamping is asserted.
- CHANGELOG and the comment say teamai does not record Copilot's cwd, not that
  Copilot sends none.

* docs(report): place the stats cleanup under usage reporting (#785)

The manual `stats/<user>.yaml` cleanup sat under single-repo mode, but the
pre-#748 leak hit every team with a git-kind repo, so it moves to "Usage
reporting" and notes where an `http` team repo keeps the file. The guide
also says the scope key is per event: hooks that run outside the project
(a worktree removed before the session ends) report to the scope they ran in.

* docs(report): name where unattributed sessions go (#785)

The CHANGELOG now says a session in a directory that resolves to no project
(a non-git project's subdirectory, a submodule or nested clone) is the user
scope's, as for skill usage. The usage guide drops the line on http team
repos: pull does not report usage to them, so no stats file there needs
cleaning.

* fix(report): each scope keeps its own reported dashboard snapshots (#786)

The report sends per-session deltas against reported-*.json snapshots that
every scope shared. A session whose events belong to two scopes (a cd into
another project mid-session) was then reported by the first scope, and the
second compared its own part with the first scope's totals and sent nothing.

Each scope now keeps its snapshots in <dataHome>/dashboard/, and the user
scope, whose data home holds the shared files, in user-reported-*.json. The
first time a scope needs one it copies the shared file, so the first report
after the upgrade sends nothing already reported; after that it reads only
its own. The user scope moves too, unlike the ticket proposed: had it kept
writing the shared file, a project seeding later would copy the user scope's
part of a split session and report nothing for its own. The shared file is
no longer written, except by an earlier release after a rollback, which only
a scope not yet seeded reads.

* fix(report): report each dashboard session once, from the scope it started in (#785, #786)

A Stop carries the whole transcript's totals (prompts, tokens,
interventions, request cost). Filtered per event, a session that moved into
another scope mid-session was reported whole again by the scope holding the
later Stop: 3 user-scope prompts then 2 in P reported 3 to the user team and
5 to P. Each session is now decided once, by its first keyed event, and
reported whole by that scope. This replaces #786's "a split session reaches
both teams with its part"; per-scope snapshots stay, so a session ID another
scope already reported (Copilot's PID fallback) still counts as new.

Unkeyed sessions from before the upgrade are decided by their first cwd. The
user scope now takes those whose directory still exists and resolves to it
(resolveConfigForDir, the dispatcher's rule) instead of dropping its whole
backlog; no cwd, or one removed since, is still no scope's.

The Copilot test also runs a payload without cwd from a hook in the project.

* fix(stats): read the scope's own dashboard filter and snapshots (#785, #786)

`teamai stats` (#771) still called filterEventsByScope with the old
{ projectRoot, excludeProjectRoots } options, synchronously, after #795
made it async and keyed by the scope config, so main no longer type-checks
and stats-scope fails. It also subtracted the shared reported-*.json, which
no scope writes since #786.

stats now filters with the config it resolved and subtracts that scope's
own snapshots (readReportedInterventions / readReportedPromptTokens, the
report's readers), so what it shows matches what pull reports. The user
scope leaves a project's older sessions out, as the report does (#785); the
stats-scope case that pinned "no exclusion in the user scope" now expects that.

* fix(stats): address CI review (#785)

A session ID now names one run up to its session_end or process_exit.
A PID-fallback ID (Copilot) comes back for a later run, maybe in another
scope, and the log keeps the ended run below the compaction threshold, so
grouping by ID alone gave the later run to the first run's scope. Each run
is still decided whole by its first keyed event.

Events written by main since #795 record the data home as a path
(`dataHome`); the report now keys them the way the writer derives
`dataHomeKey`, so pending Copilot sessions (no cwd) are not dropped.

* fix(stats): address CI review (#785)

A later run of a reused session ID (Copilot's PID fallback) was decided
on its own but returned under the same ID, so aggregation and the
per-scope snapshots merged two runs in one scope back into one session.
The filter now returns a later run as `<id>@<first event timestamp>`;
the first run keeps the bare ID, so existing snapshots still match.

An unkeyed event's cwd under a project root counted even when the
directory was gone (realpath fell back to the raw path). It now counts
only while it exists, as the docs and the user-scope rule already say.

* fix(stats): address CI review (#785)

Run identity no longer depends on which earlier runs compaction kept:
every run is `<id>@<first event timestamp>`, so a reused PID-fallback ID
is a new session even when the scope's snapshot still names the run
compaction dropped. Snapshot entries keyed by the bare ID (written by
earlier builds) are adopted by the first run of that ID in the log, so
the upgrade re-sends nothing; the next snapshot holds only run IDs.

An unkeyed event's cwd is now owned by the scope resolveConfigForDir
resolves it to, for projects as for the user scope, so a nested clone
under a project is no longer reported by both. The lexical root matcher
and its string-level tests go; the cases move to real repositories.

* fix(stats): address CI review (#785)

adoptBareKeys() read a legacy bare `pid-N` snapshot entry as the first
run's, but only in memory: the success writes merge into the file, and
with nothing new to report nothing was written, so the bare entry stayed.
Once compaction dropped that run, the next run reusing `pid-N` read it and
was suppressed. The report now writes each snapshot as soon as a bare entry
is retired, under the run ID only, even when there is no delta.

* fix(stats): address CI review (#785)

A bare snapshot entry is given only to a run an earlier release recorded
(its first event has no dataHomeKey). Only earlier releases wrote bare
entries, and a seeded one may be another scope's run under a reused
PID-fallback ID, so a run this release recorded takes none. A marker of
the seed time would miss the common case: a scope seeds at its first
report, usually the pull its first session's SessionStart triggers.

A second end of a run with nothing recorded since the first (the
dashboard monitor's process_exit after SessionEnd) joins the run it
closed instead of opening a terminal-only run counted as a session.

* fix(stats): address CI review (#785)

A scope's first snapshot is seeded only with the shared entries of its
own runs in the log, under their run IDs, and none for a run recorded
with a dataHome path: that release already kept per-scope snapshots, so
a shared entry under the same ID is another scope's. An unmatched entry
is dropped instead of copied, so a later reuse of the ID cannot inherit
it.

The dashboard monitor records processExitAfter, the last event it
observed, and the scope filter closes only that run. A delayed exit
appended after the next run of the same ID began no longer ends it and
splits it in two; an exit whose run compaction dropped is ignored.

* fix(stats): address CI review (#785)

An earlier release summed every run of a reused ID under its bare
snapshot entry, but only the first retained run adopted it, so the next
one was reported again. Each of those runs in the log but the last is
now taken as reported at its own totals and the last takes the entry,
in the report, in teamai stats and in the seed from the shared file.
The last run is undercounted by at most the other runs' share, once.

A session_start on a fallback ID from another monitorPid than its open
run's begins a new run, so a run that crashed with no dashboard running
no longer takes the next invocation, maybe another scope's. A tool's own
ID is not split: Claude fires SessionStart again on resume, in a new
process, and its Stop carries the whole transcript.

* fix(stats): address CI review (#785)

An end splits runs only on a fallback ID (pid-…). A tool's own session
ID is one session whatever ends it records: claude --resume continues it
in a new process, and its Stop carries the whole transcript, so a second
run counted it again, maybe in another scope.

* fix(stats): address CI review (#785)

A tool's own session ID is keyed by the ID itself again, as on main,
not by its first event's timestamp, so a session resumed after
compaction dropped its events still reads what its scope reported.
Only PID-fallback runs carry the timestamp.

A bare fallback entry is the sum of the runs of its ID in the log at
the earlier release's last report, and compaction keeps or drops an
ID's runs together. Those runs now consume it in log order, each up to
its own totals, so a later run that release never reported is sent
instead of taking the whole entry. The prompt-token snapshot decides
which runs it covered; interventions and daily follow it, and the first
run always takes a share.

* fix(stats): address CI review (#785)

Seeding a scope from the shared snapshot splits the whole log into runs,
lets every scope's runs of a bare ID consume its entry in log order, and
keeps the shares of the scope's own runs. The shared file summed every
scope's runs, so one scope consuming it alone could spend another
scope's baseline and suppress its own pending run.

The scope that first reports a tool's own session ID records itself in
~/.teamai/dashboard/session-owners.jsonl (the ID and its data home key,
no path), and a recorded session stays that scope's wherever it is
resumed, after compaction dropped its events too.

A dashboard started before processExitAfter existed reads the log and
appends its exit in one pass, so an unannotated exit less than one PID
check after the open fallback run began belongs to the run closed before
it instead of closing the next invocation.

* fix(stats): address CI review (#785)

A run taking its share of an earlier release's summed daily snapshot
keeps its own success and correction flags: the sum's are no single
run's (a successful run and an interrupted one sum to unsuccessful), so
an adopted run changed sessionsSucceeded without sessionsEnded.

An unannotated process_exit from a dashboard started before
processExitAfter no longer ends the open fallback run when more events
of that ID follow before the next start: a dead process records nothing
more, so it was observed before that run and belongs to the run closed
before it. This replaces the 15 s window, which a delayed callback or a
skewed clock could miss.

* fix(stats): address CI review (#785)

session-owners.jsonl is first written from the per-scope snapshots an
earlier release left: a tool's own ID in the user scope's or a
partition's prompt-token snapshot is that scope's, so a session main
reported in P, compacted and resumed in Q, stays P's instead of being
reported again to Q. An ID the shared snapshot also holds is left out:
main copied the shared file into every scope, so it names no owner, and
every scope already has its baseline. The file is created exclusively,
so a concurrent report in another scope reads the one written first.

* fix(stats): address CI review (#785)

Owner migration reconciles every per-scope baseline of an ID: a tool's
own ID in any of a scope's three snapshots is the scope's that holds its
greatest total (prompts, then tokens). A session main split per event
holds only part of it elsewhere, and a scope may have reported past the
shared total it was seeded with, so neither the first holder nor
leaving shared-held IDs out was right; a session reported with no
prompts, only its intervention count, is found too.

Besides the user scope and the partitions, it reads a project whose data
home is in its workspace that a session still in the log leads to, and
each report records the IDs of its own snapshots that no owner claims
yet, for such a project the log no longer leads to.

* fix(stats): address CI review (#785)

Owner migration assigns no owner when the greatest total ties across
scopes: main copied the shared snapshot into every scope, so equal
totals show only that copy, and each scope already holds the baseline.
A report records an ID of its own snapshots only when they show it
reported it (absent from the shared snapshot, or past its total there),
so a copy no longer claims it either.

A crashed fallback run a start from another process supersedes counts
as the run closed before it, so a late unannotated exit of it no longer
closes the new run.

* test(stats): pin a pre-upgrade exit reported before the next run's first prompt (#785)

The run split is recomputed from the whole log on every report, so once
the next run's first prompt follows the unannotated exit, the exit is
the earlier run's and the next run keeps its ID: the second pull
reports only its delta, not another session.

* fix(stats): give a compacted session resumed elsewhere to the project its transcript started in (#785)

Once compaction dropped every event of a project whose data home is in
its workspace, nothing outside it pointed to it, so a resume of its
Claude session in another project reported the transcript there again.
The transcript itself records where the session started: a Claude
transcript keeps its first cwd when resumed from another project (the
resume appends to the same file), as a Codex rollout keeps its
session_meta. Hooks now record transcriptPath on UserPromptSubmit and
SessionEnd as well as Stop (not SessionStart, whose path on such a
resume names a file that never exists; never Copilot's). A tool's own
session with no owner is the scope's that its origin resolves to, when
that scope's snapshots already hold it; otherwise it is decided as
before.

* fix(stats): address CI review (#785)

A session main split across scopes per event is credited once with
every part it reported: for each scope its `dataHome` names, the
shortest prefix of its events whose metrics reach its snapshot, and the
owner takes the metrics of their union as reported when they exceed its
own entry. Parts counted before any Stop carried the transcript's total
are no longer sent again by the owner, and cumulative Stops are not
credited twice.

A Copilot session with an explicit ID is traced to where it started by
Copilot's own session log, found by the session ID (TeamAI stores no
path of it, #666): its session.start context names the directory.

Compaction keeps a session whose tool process is still running, so a
run that an exit from a dashboard before processExitAfter marked stopped
keeps its start and its ID.

* fix(stats): address CI review (#785)

Owner migration takes a scope's entry as evidence only when its
snapshots show it reported the ID: the shared snapshots (interventions
included) hold none of it, or the scope is past their total. Main copied
the shared file into every scope it ran in, so a copy, even the only
one, names no owner, and the per-report recording follows the same rule.

A session main split across scopes whose events are gone is credited
from the parts' snapshots: a part whose daily entry shows a Stop holds
the transcript's cumulative total, so the greatest counts once; a part
with no Stop counted its own prompts, which add; intervention counts
add, tokens take the greatest. The credit rides on the owner's line in
session-owners.jsonl (numbers only) and is applied once as its baseline.

* fix(stats): keep a tool's own sessions in the first snapshot, parse legacy entries (#785)

Seeding a scope's snapshot from the shared one kept only the runs still
in the log, so a session reported before #795 and compacted before the
scope's first pull was sent again in full when resumed. Only fallback
entries need that filter, against a reused PID; a tool's own session ID
is one session, so its entry is copied whole, as main did.

Splitting a bare entry across runs read the snapshot entry as typed,
and one without `tokens` (hand-edited or truncated) threw and skipped
the whole report; the prompt-token and intervention shares now parse it
as the owner-migration path already does.

* fix(stats): report a resumed Codex rollout after compaction dropped the earlier one (#785)

A Codex build that writes a new rollout per resume restarts its
transcript counters, and the session summed only the rollouts still in
the log. Once compaction dropped rollout A, a resumed rollout B with
smaller counters was compared against A's reported total and reported
nothing until it passed it; routing the session back to the scope it
started in made that loss reach the resume in another scope too.

The prompt-token snapshot now keeps each rollout's reported prompts and
tokens under a hash of its path (no path stored), and a rollout that is
gone keeps its reported totals in the session's sum, so B is reported in
full. A session's prompts also sum its rollouts' Stop counts, which
restart per rollout like the tokens. An entry from before is compared
as a whole once, then kept per rollout.

* refactor(stats): move dashboard scope attribution and session owners out of team-push (#785)

No behavior change. src/dashboard-scope.ts holds which scope reports a
dashboard session (the log split into runs, each given whole to the
scope it started in, and the transcript origin); src/session-owners.ts
holds the machine-level owners index, its seeding from earlier
snapshots, and the snapshot files it reads. team-push.ts keeps the
snapshot adoption, deltas and push, and one reportedBaselines() now
serves both the report and `teamai stats`, which repeated the adoption
sequence.

* test(stats): real CLI resume of a compacted session from a workspace-data project (#785)

A non-git project W keeps its data home in its workspace. W reports a
Claude session; with no owners index and every W event compacted, the
session is resumed in git project Q through the real hook dispatcher,
appending to W's transcript. Q reports only its own session and W the
resumed turn; on a build without the transcript origin, Q reports both.
The fixture gains a second project and hooks sent as the installed
ones send them.

* fix(stats): credit a split session's Stop-derived interventions once (#785)

Interruptions and tool rejections come from Stops, which carry the
transcript's cumulative counts, so a compacted split session's credit
takes the greatest part, as it does for tokens; summing them made the
next cumulative Stop report nothing. Corrections are counted per prompt
in each part's own events, so they still add.

* fix(stats): place a compacted split session's parts by its transcript (#785)

A split session's credit added a part with no Stop to the greatest
cumulative Stop, which already counts that part when it came before the
Stop: after 3 prompts in P and a cumulative Stop of 5 in Q it credited
8, and the next Stop of 6 reported nothing. The credit now keeps each
part (scope key, prompts, whether it ended in a Stop; numbers only), and
the owner places them by the session's transcript, which keeps every
prompt in order with the directory it was typed in: the Stop covers the
first prompts, and only the part's prompts after those add. With no
transcript to place them, they all add, as before.

The Stop scan's human-turn test is now isHumanPromptEntry, shared by
both, so the two count prompts alike.

* fix(stats): keep a dropped Codex rollout's totals for daily and interventions, and migrate whole entries (#785)

An entry from before rollouts were kept is one total. An earlier
release rewrote every session in the log on each report, so it covers
the rollouts begun by the time its file was last written, read before
this report writes it: those still in the log consume it in order, what
is left is the dropped rollouts', kept as one prior rollout, and a
rollout begun later is new. Rollout B after a compacted A is no longer
compared against A's total and lost.

A rollout also keeps its Stop's interruptions and rejections, and its
dropped totals now reach the intervention and daily sums too, not just
prompts and tokens: daily prompt turns and intervention counts of a
resumed rollout were compared against the dropped one's.

* fix(stats): keep every metric of a dropped Codex rollout, with or without tokens (#785)

A Codex session is now kept per rollout whenever its Stops name a
rollout, not only once a Stop carries a token record, so a tokenless
resumed rollout is not compared against the dropped one's totals. Each
rollout also keeps its corrections (a correction goes to the rollout of
its prompt), its active time (each gap to the rollout of the event it
ends at) and its request costs, and a dropped rollout adds them to the
intervention and daily sums, with its cache tokens from its tokens.

The prompt-token snapshot, which holds the rollouts, is written with any
delta, so a rollout whose rejections alone moved keeps its new totals.

* fix(stats): sum a Codex session's rollout costs and keep a dropped rollout's failure (#785)

The daily snapshot took the request costs of the latest rollout only,
so with rollout A still in the log a rollout B was compared against A's
costs and clamped; a Codex session's daily costs now sum its rollouts.

Each rollout also records whether it failed (an error, an interruption
or a correction). A dropped rollout that failed keeps the session
unsuccessful, and one with a correction keeps it corrected, so a clean
later rollout does not turn it into a success.

* fix(stats): keep modern Codex rollouts, and their submit-counted prompts, per rollout (#785)

A Codex session whose tokens come from the thread-level counter
(tokenScope session) was not split into rollouts, so its prompts,
interventions, active time, costs and failure were compared against a
dropped rollout's. It is now kept per rollout like the others; the
counter already spans the rollouts, so no rollout holds tokens of its
own and the session total stays that counter's.

A Codex Stop may count no prompts, so a rollout's prompts are its
Stop's count or else its own submits: a dropped rollout's submit-counted
prompts are no longer lost.

* fix(stats): no legacy tokens on a spanning Codex counter; teamai stats writes no seed (#785)

A whole entry an earlier release left became a prior rollout carrying
its tokens, which were then added to a thread-level counter that already
holds them: rollout B's counter at 530 after A's 500 re-sent 500. A
session whose counter spans its rollouts now takes no tokens from a
dropped or prior rollout.

`teamai stats` only reads, but seeding a scope's first snapshot wrote it
with the current time, which a later report reads as the time an entry
from before covers, taking a rollout begun earlier as reported. A read
that does not persist now writes no seed, and a written seed keeps the
shared file's time.

* fix(stats): read a legacy daily entry's session cost fields as its day's costs (#785)

parseDailySnapshot() dropped the top-level pricedRequests, costMicros,
cache tokens and priceVersion a daily entry from before per-day costs
held, so an entry from before rollouts were kept lost its cost in the
prior rollout, and a later rollout's cost was compared against it and
omitted. They are now read as the session day's request costs, as
computeDailyStatsDelta already reads them.

* fix(stats): keep every Codex variant per rollout, and an older Stop's request cost (#785)

Rollout tracking recognized only `codex`, not `codex-internal` or
`tcodex`, which write the same rollouts; it now uses isCodexTool(). A
rollout's cost was read from requestDaily only, so an older Stop's
requestMetrics left the rollout without cost, and the daily snapshot,
which sums rollouts, omitted it; it is now that Stop's day's cost, as
outside rollouts.

* fix(stats): keep a Codex rollout's latest Stop by timestamp (#785)

A rollout's prompts, interventions and request costs took the last Stop
appended, though background Stop handlers may append an older scan
after a newer one, which then replaced the newer totals. They now keep
the latest Stop by its timestamp, as the rollout's tokens already do.

* fix(stats): an entry from before covers a running Codex rollout only as far as it had got (#785)

Migrating a whole entry from before rollouts were kept consumed it with
each covered rollout's current totals, so a rollout begun before the
entry was written but grown since had its later prompts taken as
reported: an entry of 6 (A's 5, B's 1) with B now at 3 reported nothing.
It now consumes it with each rollout's totals as of the entry's write,
the metrics of the events up to then; what a rollout has done since is
new.

* fix(stats): credit a split session counter by counter; read an old entry's cutoff before its push (#785)

Both credit paths applied only when the parts' prompts exceeded the
owner's, so a part that reported more active time, tokens or costs with
no more prompts was sent again by the owner. The owner's entry is now
raised counter by counter to at least the credit.

An earlier release wrote its snapshot after the push, so events that
arrived during the push predate the snapshot's time without being in
it. The team stats file in the scope's reports checkout was written
after that report read the log and before the push; the earlier of the
two times is now the cutoff an entry from before covers.
2026-09-25 10:36:37 +08:00
Saul Moro 5e5b86d9e9 fix(stats): count every worktree of a repo as that repo (#809) (#813) 2026-09-25 01:39:23 +08:00
Ruifeng Xue 49cf3fdd10 fix(docs): remove stale local documents during pull (#817) 2026-09-25 01:25:11 +08:00
Saul Moro a725574b34 fix(usage): cap usage.jsonl in scopes that never report (#788) (#790) 2026-09-24 23:43:42 +08:00
flowjzhandflowjzh b46563261c fix(hooks): give the detached hook child the bundled gits on PATH (#801)
The session-start pull is spawned through the WMI service to escape the
host's job object, and a WMI-created process inherits the provider's
environment, not the caller's. On a machine with no system git - every
WorkBuddy client - the PATH that ran `teamai` therefore never reaches the
pull: bare-name `git` lookups (simple-git, providers, mr-hint) fail with
`spawn git ENOENT`, the failure is reported only on the console spinner
(a detached pull has no console), and the clone silently stops advancing
while the postPull script - spawned by absolute path - keeps deploying a
stale tree.

Resolve the bundled git at startup instead, one resolver per host in the
git counterpart of BUNDLED_SHELLS: a qualifying root's `<root>/cmd` goes
on PATH first, with the msys dirs appended for the tools git itself shells
out to. A machine that already resolves git is left untouched, and both
outcomes are logged. The pull's outcome and the divergence notice are
persisted too - the spinner is console-only and log.warn never reaches
debug.log, which is why a stale clone had no explanation anywhere.

The suite builds its "no git on PATH" fixtures as empty dirs under the
per-test home, never a host directory like /usr/bin, which really holds a
git on macOS and Linux - the win32 simulation drops the execute-bit
requirement, so such a path silently satisfies the gate and the case
proves nothing.

Co-authored-by: flowjzh <flowjzh@users.noreply.github.com>
2026-09-24 22:23:35 +08:00
Saul Moro 9f81ae6751 fix(pull): deliver team resources to a worktree added after the last pull (#807) (#811)
state.json lives in the shared project partition, so a new worktree matched
the revision another checkout recorded and took the "Already synced" fast
path, leaving it without the team's skills, rules, agents and docs. The
shared tool targets also made two checkouts with different tool directories
force a full sync on each other on every pull.

Record the revision and targets per checkout in lastPullByWorkspace, keyed by
the checkout path plus the identity of its .git entry so a worktree re-created
at the same path does not inherit the old record. The fast path still requires
the shared lastPullRev, which exclude, tags, roles, projects, init and
bootstrap clear to force a full sync, and a pull that records a new revision
drops the other checkouts' records so every checkout does its own full sync.
2026-09-24 22:22:38 +08:00
Saul Moro ec56a67c1b fix(recall): search nothing in a project whose config cannot be read (#796) (#798)
* fix(recall): search nothing in a project whose config cannot be read (#796)

Detection skips a project config it cannot read and returns what loads
next: a legacy .teamai/ behind a broken partition, which may name another
team, or the user scope. recall searched that knowledge, recorded recalled
counts for it, and `recall --check` answered for it; with nothing behind
the broken file it printed NOT_RELEVANT, so the recall subagent told the
member the team had no knowledge and nobody learned the config was broken.

recall() now listens for the unreadable config before anything else,
searches and records nothing, prints the problem with
BROKEN_CONFIG_ADVICE and exits 1, `--check` included. A silent caller
records it in debug.log only, the rule pull follows since #784. The
teamai-recall agent relays that line instead of skipping the precheck.

* fix(recall): address pre-review findings (#796)

- The relayed line ends with "move it aside and run `teamai init`", and
  the main conversation may not have loaded the teamai skill that asks
  for consent first. The recall agent now tells it to show the line to
  the user and not act on it without their consent.
- Tools that run `teamai recall` directly (the Bash method of the recall
  rule, deployed to every tool) get the same instruction.
- CHANGELOG: only the subagent a pull from this release deploys relays
  the line; a project that broke before the upgrade keeps the old one
  until a pull succeeds there.
- The legacy-team test also asserts no votes land in that team's repo.

* fix(recall): address pre-review findings (#796)

- CHANGELOG: the entry covers `teamai recall <query>` and `--check`. The
  recall subcommands (enable, disable, status, feedback, maintenance,
  promote) still resolve their scope as before; that is a follow-up.

* fix(recall): address CI review (#796)

Reject a missing query before resolving the project, so a bare
`teamai recall` runs no detection (and no self-mode bootstrap). An empty
`--check` still resolves first: it must refuse rather than print
NOT_RELEVANT in a project whose config cannot be read.

Drop the silent branch: `recall` has no --silent flag and no caller
passes `silent`, so it was a contract nothing could invoke.

* test(recall): follow #787's per-scope votes (#796)

#787 moved recalled counts from the shared ~/.teamai/votes/ into each
scope's votes directory, so the #796 tests look for any votes directory in
the sandbox. #787's broken-project recall test expected a search to run;
#796 searches nothing there, which it now asserts, while its checks that no
scope received a vote stay.
2026-09-24 20:54:28 +08:00
Saul Moro 1fd400ecfa fix(pull): sync nothing in a project whose config cannot be read (#784) (#792)
* fix(pull): sync nothing in a project whose config cannot be read (#784)

Detection skips a project config it cannot read and returns what loads
next: a legacy .teamai/ behind a broken partition, which may name another
team, or the user scope. pull() deployed and reported for that team, and
the session-start hook did so on every session (reports-wt/ and
learnings-wt/ appeared in the legacy .teamai/).

pull() now listens for the unreadable config, syncs no scope, prints the
problem with BROKEN_CONFIG_ADVICE and exits 1. A silent pull (the
session-start hook, or a pre-dispatch hook running `teamai pull --silent`)
records it in debug.log only. Agent-root seeding and the package hint
refuse the same way, so a session start there does nothing. Hooks and
usage already follow this rule since #748. The message trimming detectTeam
used moves to config.ts as describeUnreadableConfig so both share it.

* fix(pull): address pre-review findings (#784)

- The session-start handler returns when the dispatcher resolved no
  config for the hook's cwd, which is what an unreadable project config
  resolves to since #748. That one guard replaces the unreadable-config
  sinks added to seedProjectAgentRoot and the package-hint context, and
  follows the #769 contract that handlers read their scope from the
  dispatcher. The handler tests that exercise cwd routing now pass a
  resolved scope; a new one pins that nothing runs without one.
- CHANGELOG and usage guide (en, zh-CN): a session start there runs no
  pull; only `teamai pull --silent` from a pre-dispatch hook writes the
  reason to debug.log.
- skill-data troubleshooting: what `Nothing was synced` means, and that
  moving the config aside and re-running init needs the user's consent.

* fix(pull): address CI review findings (#784)

- The session-start pull is registered with `requiresConfig` instead of
  returning early inside the handler: the dispatcher drops it wherever no
  config resolves, which covers an unreadable project config (#748), and
  spawns no detached pass for it. Where no teamai config exists at all
  it did nothing on main either (no scope to pull, no project root to
  seed, no config for a package hint). The #748 registry test and the
  docs no longer list it as machine-level work.
- `teamai pull --silent` exits 1 on the refusal too. Pre-dispatch hooks
  run it as `… 2>/dev/null || true` (`; exit 0` on Windows), so hosts
  still see success.
- The dispatch-scope test asserts the pull is skipped and resets the
  pull mock it queues.
- skill-serving design doc: `teamai pull` now reports an unreadable
  project config too. skill-data troubleshooting: `teamai doctor` can
  pass there.

* fix(hooks): still pull at session start where teamai is not set up (#784)

A null config from the dispatcher means either "no teamai here" or "the
project config cannot be read". Gating the session-start pull on
`requiresConfig` stopped it in both; only the second must stop it. The
handler now asks `findUnreadableProjectConfig` for the hook's cwd when
no config resolved (a cwd that no longer exists holds none) and runs
nothing when it reports a file. Everywhere else it runs as on main, so
the docs list it as machine-level work again.
2026-09-24 20:07:24 +08:00
Saul Moro dc233e4489 fix(votes): keep votes with the scope they were cast in (#787) (#793)
* fix(votes): keep votes with the scope they were cast in (#787)

Every scope recorded into one ~/.teamai/votes/<user>.yaml, so a vote cast in
one project (recall feedback, a recall search, a Stop whose push failed) was
pushed to the team of whichever scope synced next: the leak usage.jsonl had
before #758.

Votes now live in the data home of the scope that resolves for the session:
<dataHome>/votes/ for a project, ~/.teamai/user-votes/ for the user scope.
The Stop hook uses the config the dispatcher resolved; the pull report, recall
search, `recall feedback` and the vote view read only that scope's votes, and
the CLI readers resolve it with resolveConfigForDir, so an unreadable project
config falls back to no other scope: `recall feedback` exits 1 and the vote
view names the broken file.

The shared ~/.teamai/votes/ is never read. Its V2 `votes` map is the last
merged remote snapshot of whichever team synced, not this scope's history, and
seeding a scope from it would let `recall feedback --negative` push a
decrement and merged timestamps derived from another team. The remote
votes/<user>.yaml format is unchanged.

* fix(votes): address pre-review findings (#787)

- recall search: in a project whose config cannot be read, detection falls
  back to another scope; record no recalled count there, so the vote cannot
  reach that scope's team. Which scope the search itself uses stays #796's.
- recall feedback: with no project config and an empty or invalid user
  config, name the file and the fix (requireInit's error) instead of
  "not set up here".
- CHANGELOG: note the recall search case; the shared directory is never
  read or pushed "by this release" (an earlier release still pushes it).
- Design doc: getUserVotesDir() is the exception to "getters unchanged".

* fix(votes): address second pre-review round (#787)

- recall feedback: name an unusable user config through
  throwMissingOrInvalid (now exported) instead of re-running requireInit,
  which loaded the config twice and printed its parse error twice.
- recall search: a project detection that throws is treated like an
  unreadable config, so no recalled count lands in the fallback scope.
- The recall-search test now asserts the search ran and that neither the
  shared directory nor the broken project's votes/ was written; it fails on
  origin/main too.
- Design doc: re-wrap the edited paragraph.

* fix(votes): record recall votes from a deleted cwd in the user scope (#787)

The previous commit treated any throw from recall's project detection as an
unreadable project, including a cwd that no longer exists. Such a cwd holds
no project and resolves to the user scope everywhere else
(resolveConfigForDir, detectTeam), so its recalled counts belong there.

Tests pin that case and the single parse-error line for an invalid user
config in recall feedback.

* fix(votes): address CI review findings (#787)

- getVotesDir: a historical project-scoped ~/.teamai/config.yaml with no
  projectRoot (schema-valid, not backfilled) made getDataHome throw, so
  recall, feedback and the Stop hook recorded no vote. It lives in
  ~/.teamai, as recall and viz already treat it, so its votes go to the
  user scope's user-votes/.
- git-native-memory design doc: the local votes path is user-votes/.

* fix(votes): address CI review (#787)

- recall feedback --negative counts the upvotes the scope's own team
  already holds (its reports checkout's votes/<user>.yaml plus the
  deltas not yet pushed). A scope's file starts empty on upgrade, so a
  doc upvoted before it was rejected as not found, or as having no
  upvotes once a later recall counted it. The shared ~/.teamai/votes and
  other scopes' teams are never read.
- The adoption judge (#723, merged meanwhile) recorded into and synced
  from the user scope's votes in every scope, which pushed the user
  scope's pending votes to the project's team. It uses the scope's votes
  like the Stop handler.
- votes-scope tests: Stop transcripts prove adoption with a Read of the
  recalled file (#723), and the update.js mock keeps the real lock.
2026-09-24 20:06:35 +08:00
Saul Moro fd0e913814 fix(stats): await the async dashboard scope filter (#795) (#806)
* fix(stats): await the async dashboard scope filter (#795)

#795 made filterEventsByScope async and changed its argument from a
projectRoot/excludeProjectRoots filter to the scope config, while #771
still called it synchronously with the old filter. On main, tsc fails in
stats.ts and every stats-scope test throws "events is not iterable";
`teamai stats` crashes once there are dashboard events.

stats now awaits the filter and passes the scope config, the call pull
makes, which is what #771 set out to do: show what the report sends.
The cwd-based project-root resolution is gone with the old argument.

The user-scope test followed pull's rule before #795 (keep events that
carry no dataHome); it now follows the current one: the user scope
never reports them, so it does not count them.

* fix(stats): subtract the scope's own reported snapshots (#786)

Since #795 each scope reports against its own reported-*.json under its
data home, and the shared ~/.teamai/dashboard files are no longer
written. stats still subtracted the shared files, so after the upgrade
every session reported since counted twice in the headline.

stats reads the snapshots through team-push's readers with the scope
config, including the one-time seed from the shared file.
2026-09-24 19:50:37 +08:00
Saul Moro 8cee7ab23e fix(migrate): keep the legacy .teamai/ while the partition config cannot be read (#797) (#799)
* fix(migrate): keep the legacy .teamai/ while the partition config cannot be read (#797)

planMigration took a partition config.yaml that merely existed as a built
partition and planned a retire-only cleanup, so the first init/pull/push after
the partition file broke renamed the legacy directory to .teamai.bak although
it held the only config that still loaded. Only a partition config that
detection's own reader (readConfigFrom) accepts now counts as built; one that
exists but cannot be read plans nothing, and the next write command after the
fix retires the legacy dir as before. The re-check under the sync lock in
runMigration uses the same rule, so a broken file that appears between planning
and locking skips instead of retiring.

* fix(migrate): address pre-review findings (#797)

- Warn with the file and the first line of the reason when an unreadable
  partition config holds the migration back. The skip was silent, so a member
  had no signal which file to fix, including under --dry-run. The text says what
  happens next and hedges for a file caught mid-write.
- Stand down when the partition dir exists without a config.yaml. Keeping the
  legacy dir made "move it aside and run teamai init" (BROKEN_CONFIG_ADVICE)
  reach the full copy, which removes the partition dir before renaming the
  staged copy in and so deleted its pending learnings, env and clone. The
  re-check under the lock stands down on an existing partition dir too.
- Decide built / unreadable / absent in one helper that uses detection's own
  onUnreadable report, so the plan and the re-check cannot drift.
- Tests: a real YAML syntax error, a partition config that is not scope:
  project, the moved-aside sequence, a fresh project still planning 'full', and
  the re-check retiring when a readable partition appears after planning.
- Design doc and CHANGELOG: list every cause detection reports and the guard.

* fix(migrate): address CI review (#797)

The upgrade note in both usage guides promised an unconditional migration;
it now says a partition whose config.yaml cannot be read, or is missing,
keeps .teamai/ and what the member does next. The full copy's re-check
under the sync lock logged only at debug level when it stood down; it now
gives the same actionable warning as the planner.
2026-09-24 19:20:06 +08:00
Saul Moro 352cfc4ccc fix(report): each scope keeps its own reported dashboard snapshots (#786) (#795)
* fix(report): each scope reports only the dashboard sessions recorded in it (#785)

Every scope read one machine-wide events.jsonl and picked its sessions out
by cwd prefix. The user scope excluded nothing, so a user-scope pull
reported every project's sessions (and, through the shared reported
snapshots, took them from the project's own report); Copilot sends no cwd,
so a project never reported its Copilot sessions; and a raw cwd under a
symlink or /tmp never matched the realpath'd projectRoot.

The hook now stamps each event's dataHome with the data home of the scope
the dispatcher resolved (the key the per-scope usage file already uses), and a
report keeps only its own scope's events, comparing realpath'd keys. A
project also owns its in-repo .teamai key, where hooks record until
migration moves it to a partition. Events written before this carry no
dataHome: a project keeps those whose realpath'd cwd is under its root, the
user scope never reports them. The log stays machine-wide for the
dashboard UI, stats --by-repo, session save and the contribute check.

Removes the excludeProjectRoots option, which pull only ever passed as []
(the user target exists only when no project config resolved), and the
projectRoot option now carried by selfConfig. The usage guide documents how
to remove by hand a skill an earlier release pushed into stats/<user>.yaml
from another project.

* fix(report): each scope keeps its own reported dashboard snapshots (#786)

The report sends per-session deltas against reported-*.json snapshots that
every scope shared. A session whose events belong to two scopes (a cd into
another project mid-session) was then reported by the first scope, and the
second compared its own part with the first scope's totals and sent nothing.

Each scope now keeps its snapshots in <dataHome>/dashboard/, and the user
scope, whose data home holds the shared files, in user-reported-*.json. The
first time a scope needs one it copies the shared file, so the first report
after the upgrade sends nothing already reported; after that it reads only
its own. The user scope moves too, unlike the ticket proposed: had it kept
writing the shared file, a project seeding later would copy the user scope's
part of a split session and report nothing for its own. The shared file is
no longer written, except by an earlier release after a rollback, which only
a scope not yet seeded reads.
2026-09-24 19:18:09 +08:00
Jeffandreview a84bf6fbba chore(review): scope e2e to behavior changes, drop provider×agent matrix, add P3 (#781)
* chore(review): scope e2e requirement, drop provider×agent matrix, add P3

The auto-review gate over-flagged: it required a full provider × agent
e2e matrix as [P1 blocking] on every PR, raised theoretical/low-confidence
risks at blocking severity, and repeated already-resolved findings.

- e2e record is [P1 blocking] only for behavior-changing PRs; docs-only /
  tests-only diffs need none. One representative real-CLI run is enough;
  missing extra provider/agent coverage the author flagged as untestable
  here or deferred to CI is at most [P3 nit], never blocking.
- Add a third severity [P3 nit] (minor/optional polish, theoretical edge
  case, coverage deferred to CI).
- Tell the reviewer not to over-review: report only high-confidence
  findings, prefer few high-signal over exhaustive, never restate a
  resolved finding.

Kept in sync across AGENTS.md (## Code Review Rules, the trusted base
criteria) and the workflow prompt in codex-review-on-assign.yml.

* chore(review): relax PR 前测试 matrix, gate P1 by evidence not count

Address SaulMoro's changes-requested on #781:

- The provider × agent matrix requirement still lived in `## PR 前测试`
  (AGENTS.md/CLAUDE.md), which the Codex reviewer loads whole — so the
  matrix could return as a P1 despite `## Code Review Rules` relaxing it.
  Relax `## PR 前测试` to match: one representative real-CLI run suffices,
  docs/tests-only needs no e2e, extra provider/agent coverage defers to CI.
- Replace "prefer a few high-signal findings" (which caps real bugs and
  pushes them to later passes) with evidence-gated severity: every
  [P1 blocking] must cite a concrete failure scenario or the exact rule it
  breaks, else it drops to P2/P3. Count is uncapped; real bugs surface in
  one pass. Synced into the workflow prompt.

---------

Co-authored-by: review <review@local>
2026-09-24 19:04:58 +08:00
JinandClaude Opus 5.5 97be092633 feat(projects): add admin commands to manage the projects manifest (#756) (#774)
`manifest/projects.yaml` could only be edited by hand. Add
`teamai projects add/update/remove`, mirroring `roles add/update/remove`:
each edits the manifest, validates it with the same checks a load applies,
and opens a PR, with --dry-run to preview. The first `add` creates the file.

- `add --namespaces` sets one namespace set on every resource type.
- `update --add-namespaces/--remove-namespaces` edits each type's own list,
  so hand-edited per-type layouts survive; emptying a project is refused.
- `remove` warns about directories that still have the project active.

The pull/branch/PR plumbing moves from roles-cmd.ts into manifest-edit.ts
so both commands share the single-repo worktree handling.

An e2e test drives add -> pull, update -> pull and remove -> pull through
the built CLI: after `remove`, a member that still has the project active
has its deployed skills, rules and agents reclaimed on the next pull.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 19:01:14 +08:00
JinandClaude Opus 5.5 aa6a8830b4 fix(webhook): report failed deliveries in webhook test and log the last attempt (#777)
sendToEndpoint swallowed every failure without telling its caller, so
`teamai webhook test` printed "Webhook test successful" for an endpoint that
answered 4xx/5xx or timed out. A 5xx or 429 on the final attempt was not
logged at all, and a timeout was retried immediately instead of backing off.

sendToEndpoint now resolves to whether the endpoint accepted the event,
logs the final failure with its cause, and backs off after a timeout like
after any other failure. `webhook test` reports success only for a delivery
that went through; sendWebhook keeps never failing the calling command.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 18:59:31 +08:00
ydflowandydflow c3bfef38f8 fix(stats): count each session once and keep other projects out (#771)
* fix(stats): count each session once and keep other projects out

showStats added the whole machine's local events.jsonl metrics to the
scope's already-reported totals, so every session a pull had reported —
and that stays in the event log until compaction — was counted once by
the team total and once again locally, and sessions whose cwd belonged
to another project were added to this scope's totals too.

Filter the event log the way `pull` reports it (a project scope keeps
only sessions under its own root, the user scope excludes them) and add
only what the scope has not reported yet, derived from the same
per-session reported-* snapshots the report path advances, so the local
figure agrees with the team's instead of exceeding it. The per-repo and
by-hour breakdowns use the same filtered log, keeping them consistent
with the headline numbers.

* fix(stats): align the breakdowns with the headline and resolve the project root independently

Three points from the review of the double-count fix:

1. The project root is resolved on its own with detectProjectConfig(),
   the same call `pull` makes, instead of being read off the resolved
   scope config. A user-scope config carries no projectRoot — that field
   is attached only when a PROJECT config is detected — so the old
   expression could never populate the user scope's exclusion list.

2. `--by-repo` and `--by-time` now describe the same local part the
   headline adds to the team totals. They consumed the whole retained
   event log, so with a reported session still on disk the headline said
   "1 session, 300 tokens" while the breakdown said "2 sessions, 1.8K
   tokens" — two numbers from one command that could not both be right.
   unreportedDashboardStats now also returns the set of sessions the
   scope still owes the team, and the breakdowns filter to it.

3. The regression tests assert token totals, not only sessions and
   conversation turns, since over-counted tokens were half the bug.

Verified through the built CLI in an isolated HOME: the headline and the
per-repo breakdown now report the same sessions, turns and tokens.

* fix(stats): keep the breakdowns reading the scope's own event log

The previous round filtered the breakdowns down to unreported sessions
so they would match the headline. That was the wrong trade: the headline
adds this machine's unreported sessions to totals that already include
other machines and sessions compaction has dropped, so it can never
equal a per-repo or per-hour view of the local log. Filtering made the
breakdowns show neither a total nor a delta — a fully reported project
vanished from `--by-repo` entirely.

Restore the full scoped log for the breakdowns and say in the comment
what question each answers. What both must share is the SCOPE, and the
breakdown tests now pin that: removing the scope filter turns the
cross-project case red, which the earlier assertion missed because it
read only the first matching row.

* fix(stats): subtract reported totals only when they were read

The delta path keyed off `config` alone, so a scope whose team stats
could not be read at all — no stats file yet, an unreadable one, a
reports worktree that is not there — still had its local snapshot
subtracted. The snapshot records what this machine pushed, not what the
team holds, and with the reported side null it hid sessions the member
could see happening, down to "No usage data yet."

Require `reported` as well, so the local aggregate is shown when there
is no team total to reconcile against. The changelog entry no longer
claims the breakdowns match the headline; they read the same scoped log
and answer a different question, and the heading says so.

* docs(stats): say which log the optional breakdowns read

The `--by-repo` heading was changed to name the local event log, but
the flag descriptions and the generated command reference still said
"Break usage down per repository", and `--by-time` kept the old
"(local time)" heading — so the two optional views disagreed with each
other and with the changelog entry about them.

Name the source in both flag descriptions, label the by-hour view the
same way, and regenerate `commands.md` per AGENTS.md.

* test(stats): cover the project scope against its reported totals

The end-to-end shape the review asked for was missing: a project scope
with reported team totals, one reported session still in the event log,
one new session in this project, and one session belonging to a
different project. It now asserts 3 reported + 1 new = 4 sessions, 301
turns, and a breakdown holding only this project's rows.

Dropping the scope filter turns four tests red at once (the new one
reporting 5 instead of 4), so the filter is pinned rather than assumed.

* fix(stats): do not subtract a snapshot the team file never received

The previous guard only checked that the reported totals were readable.
The local snapshots are machine-global while the team file is per-scope,
so a snapshot can name a session this team never got — an empty team
file alongside a populated snapshot. Subtracting then undercounts, down
to "No usage data yet." with sessions sitting in the event log.

Trust the snapshot only when the team total is non-empty: the report
path writes the team file and advances the snapshot under the same lock,
so a non-empty total is what licenses the subtraction.

* docs(stats): correct the changelog's trust and scope claims

The note inverted the guard's condition: the code trusts the reported
snapshots only when the team total is NON-empty, while the wording said an
empty total is what licenses trusting them. Say what the code does.

It also claimed the user scope excludes project roots. A user config
resolves only when detectProjectConfig() found no project, so there is no
root to exclude and the log passes through unfiltered. State that instead of
asserting an exclusion the branch cannot perform.

No behaviour change; wording only.

---------

Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-24 18:55:32 +08:00
Jiahe GengandClaude Opus 4.8 89ac24a3d4 fix(votes): collect upvote adoption from tool-use evidence + opt-in LLM-judge (#723) (#744)
* fix(votes): collect upvote adoption from tool-use evidence + opt-in LLM-judge (#723)

Fixes #723. upvoted_count was structurally near-zero because collecting an
upvote depended on the main agent voluntarily emitting the
<!-- teamai:referenced-doc-ids: [...] --> marker (~3.7% in the issue's data).
This collects adoption WITHOUT AI self-declaration and removes that mechanism
entirely.

Signals:
- Tool-use evidence (always on): a recalled doc is adopted when the MAIN agent
  opens its file (Read/Grep/Glob/Bash). Sidechain tool calls excluded; gated to
  the recalled set; full-path or >=2-segment suffix match (no bare-basename
  cross-attribution); relative `./x` normalized; Bash harvests FILE-OPERAND
  tokens only (grep patterns, option values, `#` comments and `>`/`>>`/`2>`
  redirection targets never credit; `-e/-f` frees the file operand); a FAILED
  tool_result (is_error) revokes its refs; Glob/Grep matched files are harvested
  from the reader RESULT (input path is often just a directory).
- Opt-in background LLM-judge (TEAMAI_UPVOTE_JUDGE=1, off by default): a detached
  Stop handler asks the local signed-in CLI whether the latest reply used each
  still-uncredited recalled doc, grounded in the doc's real content. Grounded-only
  (a candidate whose excerpt cannot be securely read is dropped, so a forged
  recall marker cannot earn an upvote from its id alone); fail-closed excerpt read
  (lstat + realpath + .md-in-trusted-root); trusted roots derive from
  learningsRoots + pendingLearningsDir; prompt fences excerpts/reply as untrusted
  data. Each recalled doc is judged at most once per session via a per-doc
  judged-record (crash-safe: recorded AFTER the CLI call, so a killed run retries;
  later turns still judge NEW docs) — no exclusive claim marker.

Scope attribution: recall labels each hit [project]/[user]; while a project is
active a doc recalled from the inherited USER scope is read-only and is NOT
upvoted into the project team (matches recall.ts's recalled_count scoping and the
documented rule). Recall regions from a reader tool_result (file content the agent
opened) are UNTRUSTED — parsed into throwaway sinks so a forged region can neither
manufacture a doc-id nor poison a real doc's scope/path; only assistant text,
non-reader results, toolUseResult.stdout and plain-string content are trusted.

Concurrency & idempotency:
- All vote mutators serialize on one cross-process file lock; the votes file is
  written atomically (temp+rename) so a killed detached judge cannot leave a torn
  file that loadUserVotes would read as empty. creditedDocIdsForSession reads the
  YAML directly (never loadUserVotes) so a v1 file is not auto-migrated by an
  unlocked read.
- Per-session dedup ledger lives inside the votes file under the same lock; its
  TTL is measured from the session's first-seen time (firstTs) so a long/resumed
  session cannot re-credit an already-adopted doc; TTL-pruned every Stop; local-
  only (never synced to the team repo via mergeDeltas).

Also: deterministic "[teamai] Adopted team knowledge this session: <ids>" summary
(once per session, tools that print the Stop payload); finalAssistantText joins
all text blocks of the final message and accumulates across records sharing a
message id; recall lock-exhaustion is logged honestly; removed the
referenced-doc-ids parser/nudge/stash and its rules text in builtin-rules.ts /
pull.ts. gitOnly and promote thresholds unchanged.

Docs: document TEAMAI_UPVOTE_JUDGE (usage-guide en+zh); note OpenCode does not
participate in adoption (session.idle carries no JSONL transcript_path);
git-native-memory design flow + both decision tables updated to adoption-driven.

* fix(votes): tighten adoption evidence in transcript parser (#723)

Address the #744 review's precision findings in adoption collection:

- Reader tool results no longer harvest every Markdown-looking string;
  only whole-line file paths and grep `path:line:` prefixes count, so an
  unrelated notes.md that merely mentions a recalled doc's name cannot
  credit it.
- Relative tool-call paths are resolved against the transcript entry's
  cwd before matching, so a relative `learnings/setup.md` opened in one
  checkout is not misattributed to a recalled doc under another checkout.
- Shell command splitting is quote-aware, so a `|` inside a quoted grep
  pattern no longer forges a synthetic `cat` segment that falsely credits
  a doc named only in the pattern.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(votes): dedup the LLM-judge via the vote ledger, drop session marker (#723)

The per-session "judged-docs" marker introduced crash-unsafe and
once-per-session-violating behavior (#744 review): a transient CLI
failure was recorded as judged and never retried, a positive verdict
was marked judged before the vote write so a busy lock lost it
permanently, an early negative verdict permanently excluded a doc a
later reply actually used, and the 24h marker TTL re-credited docs on
a resumed session.

Remove the marker entirely and rely on incrementUpvoted's atomic,
lock-protected per-session ledger, which already dedups credits across
the foreground and background passes. A verdict is now recorded only
after the atomic credit succeeds; an unadopted recalled doc may be
re-judged on a later Stop (the judge is opt-in), which is the accepted
cost of removing the crash-unsafe marker.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-24 18:53:02 +08:00
dvd233 9d3a91cf76 fix(push): honor explicit branch and protect dirty team clones (#690) 2026-09-24 18:43:14 +08:00
Yu Geigei 57afe76810 feat(models): merge show into list with an optional profile (#782)
Model catalogs are small, so `teamai models list` now prints every profile in full: API key source, gateway, models by protocol, compatible agents, and where it is active. `teamai models list <profile>` narrows the output to one profile, and the separate `show` command is removed.
2026-09-24 15:37:53 +08:00
Jeff 82ecf4153d docs(skill): harden credential and consent guidance in teamai skill (#779)
ClawHub's SkillSpector scan flagged three doc-level issues in the teamai
skill guidance:

- PE3 (Credential Access): the GitLab setup showed a literal
  `export GITLAB_TOKEN=glpat-...`, which lands in shell history and process
  listings. Switch to a no-echo prompt, recommend a short-lived api-scope
  token, and unset it after init.
- SQP-2 / P4 (silent install + behavior manipulation): the TGit login guidance
  told the agent to run install/login itself and 'never tell the user'. Require
  disclosing that it installs a binary and stores a credential, and getting the
  user's OK first (or showing the commands if they prefer).
- SDI-4 (trigger ambiguity): publishing a skill from a plain-language request
  had no confirmation gate. Require showing what will be shared and getting an
  explicit go-ahead before teamai push / contribute.

Doc-only; edits land in skill-data/ (served by `teamai skill get`).
2026-09-24 15:30:01 +08:00
Saul Moro 2ab697d062 fix(usage): discard pre-upgrade usage and stop falling back past an unreadable config (#748) (#758)
* fix(usage): discard pre-upgrade usage and stop falling back past an unreadable config (#748)

Follow-up to #753, from its review.

- resolveConfigForDir returns null when any project config was reported
  unreadable, even if a lower-priority one (a legacy .teamai/ behind a broken
  partition) loads: that one may name another team.
- The user scope's usage.jsonl is the old shared file. #753 only emptied it on
  a machine's first user-scope init, so a machine that already had a user
  scope reported every project's pre-upgrade usage to it. The first access
  after the upgrade now discards what an earlier release left there and
  writes ~/.teamai/usage-per-scope. One process discards, under acquireLock;
  concurrent hooks wait for the marker, so none deletes what another recorded.

* fix(review): parse the empty usage file in its test; scope the fallback wording to team hooks (#748)

- "handles empty file" wrote no marker, so the discard removed the file and
  the read passed on a missing file. A first read now settles the file as
  the scope's own, and the test asserts the file survives.
- The session-start pull still resolves its project on its own, so the
  "never falls back to a lower-priority config" rule is stated for team
  hooks and skill usage only (CHANGELOG, usage guide en/zh-CN).

* fix(usage): keep the user scope's usage in its own file, safe across a rollback (#748)

The usage-per-scope marker could not tell a pre-upgrade event from one an
earlier release appends after a rollback, so a reinstall reported those to
the user-scope team. The user scope now records in ~/.teamai/user-usage.jsonl,
which no earlier release writes; ~/.teamai/usage.jsonl is removed, never read.
Drops the marker, its lock and the bounded wait.

* fix(usage): a failed removal of the shared usage file does not stop the user scope (#748)

Also names the user scope's own file where comments and the design diagram
still described every scope's usage as <dataHome>/usage.jsonl.

* docs(changelog): drop the claim that teamai doctor reports an unreadable project config

resolveDoctorContext falls back past an unreadable project config the way
detection does, so doctor diagnoses the config it falls back to and says
nothing about the broken one (#752).

* fix(usage): leave the shared usage file in place instead of removing it on every access (#748)

getUsagePath deleted ~/.teamai/usage.jsonl on every user-scope call,
including each hook append and the read-only `teamai stats`. The user scope
never reads that file, which is what keeps its events off the team; the
delete added a side effect to a path getter and a failure path to guard.
2026-09-24 15:04:18 +08:00
RererrandClaude Fable 5.1 a2f93ae3d2 fix(config): release only Claude's MCP servers on a root move; read the recorded root in import and skill tracking (#775)
A re-init that moved the Claude Code root handed the full team config to
reconcileMcpForConfig({ removeAll }), which walks every MCP-capable tool,
so Codex, Cursor and the rest lost their teamai-managed servers until the
next pull. The release now narrows the team config to Claude.

import --from-claude scanned ~/.claude/rules and skill-use tracking only
knew the static ~/.claude/skills; both now resolve the recorded root. The
resolution (project config governing the directory, else user scope) moves
into resolveMemberToolRoots so the local agent, import and tracking agree;
tracking resolves it from the hook's reported directory. The helper checks
that the directory exists before probing, as resolveConfigForDir does, so a
hook from a deleted worktree still records — this also stops the local
agent from throwing on a missing workspace path.

Follow-up to #728 (third review pass, findings 1 and 5).

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 15:01:41 +08:00