`reconcileMcpForConfig` wrote team servers into every installed tool
returned by `resolveMcpTargets`, without checking `isAgentExcluded`. A
member who ran `init --agent claude` still got the team's servers in
`~/.cursor/mcp.json` when Cursor was installed, and `uninstall --agent
cursor` did not stop the next pull from injecting new ones there.
Skills, rules, agents, hooks and builtin deploy already apply this gate
(#504, #523, #592).
The reconcile loop now skips excluded targets. Their manifest entry is
left as is, so servers injected before the upgrade stay where they are,
and `removeAll` (uninstall, `mcp remove`) still reaches every tool.
In project scope tclaude has no MCP file of its own and reads the
<root>/.mcp.json written by the claude target, so that target stays
live while tclaude is enabled.
`teamai tags subscribe` and `teamai tags unsubscribe` save the new
subscription list and tell the user to run `teamai pull`, but they left
`lastPullRev` in place. When the team repo has not moved since the last
pull, `pull` takes its unchanged-revision fast path and prints "Already
synced", so rules filtered out by the new subscriptions stay installed and
rules brought back by an unsubscribe are never restored.
`teamai skill exclude` already clears `lastPullRev` for the same reason.
The tags commands now do the same when they save the config.
`removeItem` in `RulesHandler` and `SkillsHandler` deleted the resource
from every tool in `toolPaths`, without checking `isAgentExcluded`. A
member who whitelists `enabledAgents: ["claude"]` still lost the Codex
copy. The tombstone pass in `teamai pull` checks that whitelist, so the
two sides of the same removal disagreed.
PR #589 added the gate to the agents handler while closing #576, and
left these two out so the tombstone regression stayed reviewable. This
applies the same gate to both.
Gating the two removal loops is not enough on its own. `RulesHandler
.removeItem` ends by calling `pullAllRules`, whose stale-file sweep
iterates every tool and deletes any local rule missing from the team
set, including the one just removed. `pullItem` in the same file already
skips excluded tools, so the sweep now does too. Without that, the gate
in the removal loop is bypassed whenever the team still holds another
rule, which is the normal case. The sweep also runs during `teamai
pull`, so an excluded tool keeps its rule files there as well.
For skills the check sits above the OpenClaw branch, so an excluded
tool's workspace copy and Codex's shared `.agents/skills` destination
are both covered.
`MCPHandler.removeItem` needs no gate. It only rewrites the team repo's
`mcp.yaml` and never touches a tool directory.
Fixes#590
execVersion ran `<agent> --version` through the native execFile. On
Windows, npm installs codebuddy, claude and openclaw as `<name>.cmd`
shims with no `.exe`, so execFile fails with ENOENT, the error is
swallowed, and the local-agent report carries an empty agent_version.
Use probeBinary from utils/exec, which already launches through
cross-spawn and is what env-detect uses to probe `claude --version`.
`teamai remove agents <name>` deleted the `.md`, `.toml` and `.json`
renders on the machine that ran it, but the tombstone pass in `teamai
pull` tried only `.md`. The Codex `.toml` and Kiro `.json` copies of a
removed agent stayed on every other machine.
The cause was duplicated extension knowledge. `removeItem` and the
tombstone pass each carried their own list, and the two drifted.
`rule-format.ts` already documents the same split for rules, a per-tool
function for writers and one shared list for scanners and deleters.
Agents now get that split through `AGENT_FILE_EXTENSIONS`, and the three
other sites that hardcoded the same triple read it too.
The same code had two further gaps:
- `removeItem` ignored `isAgentExcluded`, so `teamai remove` deleted
agents from tools the member had excluded through `enabledAgents`.
- `pull` returns early when the team repo rev is unchanged, before the
tombstone pass. Every machine this bug affects is in that state,
because it already pulled the tombstone with the older CLI, so the
upgrade never reached it. The cleanup now also runs on that path,
next to the built-in deploys that are there for the same reason.
Fixes#576
* feat(roles): add agents resource namespaces to roles and projects manifests
Optional `agents:` key on roles.yaml and projects.yaml resources, merged
into the active namespace set. Absent means no agents namespaces, which
keeps today's behaviour for existing manifests. Part of #563.
* feat(agents): scan one level of agents/<namespace>/ in the team repo
Root files stay shared. Subdirectory files carry a namespace like
rules/<namespace>/ so pull can filter them by role. Part of #563.
* feat(pull): filter team agents by active namespaces and reject stem collisions
Root-level agents ship to everyone; agents/<ns>/ ships only when <ns> is an
active agents namespace, mirroring filterRulesByKnowledgeNamespaces. Two
kept agents with one stem would overwrite each other on disk, so pull
fails loudly like scanRoleAwareSkills does for skills. Part of #563.
* feat(agents): remove deployed agents of namespaces that stop being active
A role or project change now revokes the previous namespaces' agents on
every installed tool. A deployed file is deleted only when it is byte-equal
to what pull renders from the team source; a local edit is kept and
reported, the same data-safety rule inactive skills follow. Part of #563.
* feat(agents): locate team agents by stem across namespaces in push, remove and uninstall
A modified agent whose source lives in agents/<ns>/ is written back there
instead of creating a root duplicate; remove and uninstall find namespaced
sources the same way. New agents still land at the root. Part of #563.
* docs: describe role-scoped agents in both READMEs, usage guides and the changelog
* refactor(agents): share the agents directory walk and drop dead nullish guards
Review follow-ups: one listTeamAgentDirs used by scan, stem lookup,
uninstall and the self-mode pickup (which now sees agents/<ns>/ too);
removeItem iterates every match once; the roles/projects resolvers rely on
the zod default for agents instead of runtime guards; design doc example
gains the optional agents key.
* fix(agents): resolve active sources and clean up per tool
* fix(reports): merge report writes onto origin's latest reports data
Writers (session save --push, Stop votes, member registration, auto-report)
now sync the reports worktree under the reports lock before they
read-merge-write, so the same member on two machines cannot lose an
entry to a stale checkout.
Closes#561
* refactor(reports): keep the #561 writer path closer to existing style
Drop the init roster state machine, restore the original merge/restore
comments in auto-report, and keep commitAndPushReports' docs and retry
loop as they were. The write still happens after sync under the reports
lock.
toZcodeEntry took its timeoutMs only from the ZCODE_TIMEOUT_MS table added in
#555 and never looked at def.timeout, so on ZCode a timeout the team wrote down
was silently discarded: a `timeout:` on a hook in hooks/hooks.yaml, and the
`builtin.overrides.<key>.timeout` knob documented in the usage guide. Before
#555 the writer honored it (`timeoutMs: def.timeout * 1000`); the table is the
right default for the builtin dispatch hooks, but it took over the explicit
values too.
Observed end-to-end with the built CLI: a team hook declaring `timeout: 300`
and `builtin.overrides."Hook dispatch stop".timeout: 240` both landed in
.zcode/cli/config.json as `timeoutMs: 60000`. Every other tool honors the same
fields, so a hook long enough to matter on ZCode was capped at the default and
killed mid-run.
The table now serves as the default only, and a stated timeout wins. Builtin
defs for zcode carry no timeout (builtinHookDefs sets withTimeout false for the
tool), so the rendered output for a config without overrides is unchanged --
hooks-golden stays green.
* feat(init): activate every project with --project all
`--project all` expands to every id manifest/projects.yaml declares and
persists that snapshot, so a monorepo keeps one line in its onboarding docs
instead of a project list that has to be edited whenever the manifest
changes. `all` is reserved: a mixed list (`all,backend`) is rejected as
redundant or a typo, and a project whose id is literally `all` is still
covered by the expansion but selected on its own through
`teamai projects set all`, which takes plain ids.
It stays an explicit operator choice — activate everything, project-private
learnings included — rather than the auto-activation the multi-project design
rules out. Snapshot semantics also keep the existing rule that re-running
`init` is what re-resolves the active set.
Scope is `init` only; `teamai projects set --all` is deliberately not added.
Docs: the usage guide (en + zh-CN) flag table and multi-project section, the
CLI help text, and the multi-project design note.
* test(init): real-CLI e2e for --project all
The unit cases in `init-projects.test.ts` drive `resolveActiveProjects`
directly. This adds the end-to-end leg for the user-facing command: build the
CLI, run `init <repo> --project all` against a real team repo, and read the
ids back out of `.teamai/config.yaml`.
Offline, without weakening the test: the remote is a synthetic HTTPS URL that
git's `url.<base>.insteadOf` rewrites to a local bare repo. init rejects plain
HTTP outright, and a plain-HTTP static server (the first attempt) never gets
past the URL validator. The rewrite target is a filesystem path rather than
`file:///C:/…`, which on Git for Windows parses as `/C:/…` and fails the
clone.
Two cases: `all` expands to every manifest id in manifest order (the fixture
is deliberately not alphabetical, so an accidental sort is caught), and an
explicit id still replaces the previous selection instead of merging with it.
On `main` the first case fails with `Unknown project "all". Available
projects: gamma, alpha, billing` while the second passes — the case that
fails is the new selector, not the harness.
* fix(init): refresh reused clone before resolving --project all
Re-running init --project all against an existing clone previously expanded the stale local manifest. pullRepo before resolveActiveProjects so newly added remote projects are picked up.
* test(init): cover remote manifest update on clone reuse for --project all
Regression for PR #518 review: re-init with an existing clone must pick up projects added to the remote between runs.
* fix(init): non-destructive clone refresh; repair --project all e2e
Address jeff-r2026 review on #518:
- init clone-reuse path uses pullRepoFastForward (ff-only only; never reset --hard)
- e2e reads repo.localPath from project config; git wrapper keeps remotesMatch on synthetic origin
- regression: tracked local edits survive a failed refresh
* fix(init): drop misleading --force hint; canonicalize e2e clone-path assert
A matching-origin clone is always reused regardless of --force, so
suggesting it as a replacement after a failed refresh was wrong; point at
manual recovery instead. The e2e assertion now compares realpathSync
canonicalized paths so it holds on macOS where tmpdirs live under
/private/var.
Two independent defects caused team-wide upvote counts to stay at 0
while recall counts accumulated normally:
- mergeDeltas merged upvoted_count via remote+delta but last_upvoted_at
via max(local,remote), producing orphan timestamps (count 0 with a
last_upvoted_at set). Now clears the timestamp when the merged count
is 0, and recallFeedback clears last_upvoted_at when a downvote brings
the count to 0. Prevents the inconsistent state from propagating.
- parseStdin returned null on malformed STDIN, short-circuiting the
entire hook dispatch so votes-sync (and other handlers) never ran when
a concurrent hook delivered corrupt STDIN. Now degrades to {} and
continues; handlers that need STDIN fields self-skip. Also normalizes
valid-but-non-object JSON (null/number/array) and logs a diagnostic
preview of the unparseable payload.
Adds unit tests for both paths.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(mcp,hooks): parse an optional roles list on mcp.yaml servers and hooks.yaml hooks
Carried through to McpServerDef and HookDef; not yet applied. Part of #563.
* feat(roles): add activeRoleIds and matchesRoles for per-entry role filters
Shared by the MCP and hooks reconcilers. Omitted roles match everyone, an
empty list matches nobody (like tools), no configured role means no
filter, the same fallback skills and rules already use. Part of #563.
* feat(mcp): apply the roles filter when reconciling team MCP servers
A server with roles: ships only to members whose primaryRole or
additionalRoles it lists; a role change removes it on the next reconcile
through the existing desired-set rebuild. Unknown role ids get one warning
per id so a typo does not hide a server in silence. Part of #563.
* feat(hooks): apply the roles filter when resolving team hooks
resolveTeamHooks drops hooks whose roles: does not list one of the member's
roles, before the security gates so the transparency print shows only what
will run. A role change removes the previous role's hooks through the
existing managed-slice rebuild. Unknown role ids warn once. Part of #563.
* feat(mcp,hooks): show the roles restriction in mcp list and hooks list
* docs: describe the roles field for MCP servers and hooks in both guides, READMEs and the changelog
* refactor(roles): review follow-ups for the roles filter
Warn once per process for an unknown role id so a member with user and
project scopes reads it once per pull; matchesRoles accepts an undefined
active set; list commands print nobody for an empty roles list, matching
the docs; docs example uses the devops role the guide already defines.
* fix(zcode): strip unknown keys from the hooks block
ZCode validates the hooks block against a strict schema and rejects the
entire block on any unrecognized key (config_file_invalid → hookCount 0 →
no hook fires, with the runner still reporting installed). A single
hand-added annotation key therefore silently kills every hook.
reconcileZcodeFormat now heals the config by keeping only the keys the
schema knows (enabled/events), so a poisoned config recovers on the next
inject/pull instead of staying dead.
* fix(zcode): use network-scale hook timeouts
Session-start dispatches carry a network pull to the team host; on slower
links that exceeds the 10-15s builtin shell-hook defaults, so ZCode killed
the hook mid-pull (observed: hook.run.failed at durationMs 10023 with a
10000ms timeout, last pull left stale). toZcodeEntry now applies a
per-event timeout table (SessionStart 180s, Stop/UserPromptSubmit 60s,
PostToolUse 30s) instead of inheriting the shell-hook defaults.
* fix(zcode): route Windows hook entries through cmd.exe
On Windows, CreateProcess resolves bare `bash` to System32's WSL launcher
before any PATH directory. The WSL side has a different $HOME (no teamai
state) and often no Node >= 20, so the spawned hook no-ops or dies —
observed on a real machine: hook.run.failed for every ZCode hook while
the same payload succeeded through Git Bash.
On win32, render entries as cmd /c <payload> (cmd.exe always exists in
System32; the npm .cmd shim resolves via PATHEXT). POSIX entries keep
bash -lc. Payload stays verbatim in args[1] on both variants, keeping the
manifest-invariant that fixed the team-hook matching bug.
* fix(reports): keep teamai-reports reads fresh and read-only
#489 moved independent git clones onto the teamai-reports worktree but kept
the single-repo reader/writer patterns, which regressed git-kind clones:
read-only commands published a missing teamai-reports branch (#558), pull
indexed votes and read stats from a stale reports checkout (#557), and
writers merged per-member files into that stale copy, so a second machine of
the same member stayed diverged and its reports never reached origin.
- Read-only callers (members, digest, projects members, stats, viz, the vote
read in contribute, pull's votes/stats reads) pass pushIfCreated: false.
- A cold start reuses an unpublished local teamai-reports branch instead of
failing with "a branch named 'teamai-reports' already exists".
- refreshReportsWorktree syncs under the reports lock: offline keeps the local
copy, unpushed report commits are rebased onto origin and dropped only when
they conflict, and a busy lock reads the local copy.
- pull refreshes the reports worktree once per scope before the search index
and skill recommendations.
- updateReports() replaces ensure + write + commitAndPushReports: it syncs,
runs the merge/write, commits and pushes under one lock. Pull auto-report,
Stop-hook votes, session save --push, init and bootstrap member
registration use it; auto-report and the Stop hook skip the round-trip when
nothing is pending.
Closes#557Closes#558
Claude-Session: https://claude.ai/code/session_01TFtw1GRR7pJNTGzNCikoKe
* refactor(reports): narrow the fix to #557/#558, move writer changes to #561
Keep this PR to what the two issues ask for: read-only callers never publish
teamai-reports, cold start reuses an unpublished local branch, and pull
refreshes the reports worktree (under the reports lock) before reading votes
and stats. Pull's auto-report runs after that refresh, so its stats merge
also starts from fresh data.
The writer-side refactor (updateReports replacing commitAndPushReports in
init, bootstrap, session save, the Stop hook and auto-report) is reverted
here and proposed separately in #561.
Because writers still write before taking the lock, the refresh no longer
discards uncommitted report files: it fast-forwards, or rebases unpushed
commits with --autostash, and only resets a clean worktree whose unpushed
commits conflict with origin.
Claude-Session: https://claude.ai/code/session_01TFtw1GRR7pJNTGzNCikoKe
* fix(reports): restore uncommitted files after rebase autostash conflicts
git rebase --autostash can exit 0 when the rebase itself succeeds but
reapplying the autostash conflicts. That left UU paths, conflict markers
in report YAML, and a leftover autostash stash, so a read-only command
could serve invalid stats.
After rebase/merge, inspect the worktree. If the rebase has finished and
stash-apply conflicts remain, restore the original uncommitted content
(the stash --theirs side) and drop the autostash instead of treating
exit 0 as proof that the checkout is clean.
* fix(reports): drop only the rebase autostash, not shared stashes
Git worktrees share refs/stash. dropRebaseAutostash scanned the whole
stash list and dropped every entry whose message contained autostash, so
a self-mode members refresh could delete a business-tree stash created
with `git stash push -m autostash`.
Snapshot stash SHAs before rebase --autostash and drop only objects
created by that refresh. Do not select stashes by message text.
* fix(reports): carry dirty reports without touching refs/stash
stashShasCreatedSince treated every stash SHA that appeared after the
pre-rebase snapshot as this refresh's autostash. Worktrees share
refs/stash, and the reports lock does not serialize ordinary git stash
in the business tree, so a stash created during rebase was deleted.
Stop using rebase --autostash. Snapshot dirty tracked files with
git stash create (a dangling commit, not stored in refs/stash), rebase
on a clean tree, then stash apply that object. Never drop from the
shared stash list.
* fix(dashboard): match correction keywords as whole words, add team keywords
`isCorrectionPrompt` used raw substring matching, so the built-in `undo` and
`redo` fired on ordinary Spanish and Portuguese words ("segundo", "mundo",
"redondo"). One false correction scores 20, which is the share-learnings nudge
threshold on its own.
Keywords in a space-separated script now match as whole words (Unicode-aware,
so accented letters count as letters). Keywords containing Han, Hiragana,
Katakana or Hangul keep substring matching.
Teams can add their own words via `sharing.intervention.correctionKeywords` in
teamai.yaml. The prompt_submit hook resolves them and stores a `correction`
flag on the event, because the machine-level events file mixes sessions from
every team. Events without the flag fall back to the built-in list.
For #564
* fix(dashboard): resolve team keywords from the hook cwd, treat _ as a word char
Cursor runs hooks from ~/.cursor and sends the project in workspace_roots, so
resolving the team via autoDetectInit() (process.cwd()) silently dropped the
team's correctionKeywords there. Resolve the project from resolveHookCwd(stdin)
and fall back to the user-scope config.
Underscore joins the word boundary so identifiers such as "test_undo" do not
count as `undo`. Drop the regex cache and fold both keyword lists into one
`.some`. CHANGELOG separates the whole-word fix from the team-keywords feature
and notes the new `correction` event field; both usage guides say the built-in
list still covers only zh/en/ja.
* fix(dashboard): keep autoDetectInit for team keywords in the prompt hook
hook-dispatch-cli already chdir's to the hook payload's cwd before running
handlers, so autoDetectInit() resolves the right project on Cursor too. The
explicit detectProjectConfig(resolveHookCwd(stdin)) path added in the previous
commit was redundant; drop it and its tests.
* fix(pull): refresh the team clone before reading teamai.yaml
A clone missing teamai.yaml exited early before git pull, so it could
never fetch the file from the remote, even with --force.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(pull): drop the redundant teamai.yaml reload after refresh
The config is now read once, after the clone refresh, so the later
'reload after pull' block was a duplicate read with a stale comment.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(e2e): cover pull self-heal for a clone missing teamai.yaml
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(pull): reuse the loaded team config in the skip-sync fast path
The fast path re-read teamai.yaml that was already loaded right after
the refresh, and re-checked dryRun that its enclosing guard excludes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: hanhnt2 <hanhnt2@hblab.vn>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(git): route independent-clone reports onto teamai-reports
Independent git clones currently push members/sessions/votes/stats onto the
default branch, so every member needs write access on main. Reuse the existing
orphan-branch worktree writer: git-kind reports live beside the clone on
teamai-reports; self stays nested under .teamai/; HTTP is unchanged. Empty-repo
init may still seed teamai.yaml on the default branch; leftover report files
on main are ignored.
Closes#484
* fix(git): recreate stale sibling reports worktree after re-clone
isGitRepo only checks that reports-wt has a .git file. After init replaces
the clone, that file still exists but the backing gitdir is gone, so
ensureReportsWorktree returned the husk and commitAndPushReports failed
silently. Probe revparse --is-inside-work-tree and fall through to the
existing remove+recreate path.
The recall rule body instructs the AI to "invoke the teamai-recall
subagent". In tools like Cursor that leak always-apply rules into
subagent sessions, the teamai-recall subagent itself reads this and
spawns another teamai-recall subagent, looping until it hangs.
Add a bilingual self-exemption block at the top of both the rule body
(builtin-rules.ts) and the CLAUDE.md-injected block (pull.ts): if you
ARE the teamai-recall subagent, the rule does not apply — proceed
directly to the search instead of invoking recall again.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(partition): adopt pre-#546 legacy-named partitions instead of stranding them
#546 widened the partition slug prefix from the anchor's basename to its
whole path but kept the sha256 suffix — without migrating existing
installs. On an upgraded CLI, a partition still named <basename>-<hash>
(every install created before #546) stops matching projectSlug(anchor):
detectProjectConfig finds no config (the project looks uninitialized),
planMigration would re-copy a retired workspace into a second, empty
partition, and status --all flags the perfectly good data as "corrupt —
dir name does not match anchor".
Both formats share the same hash, so an anchor's legacy name is
computable exactly — no directory scanning. resolvePartitionDir becomes
the seam for "the partition that actually holds this project's data"
(detection, init, migration): when the canonical current-format
directory is absent and a legacy-named one exists, it ATOMICALLY RENAMES
the legacy partition into place — a same-parent metadata move, no data
copied, and an interruption leaves either name intact. An authoritative
current-format partition is never clobbered by a leftover legacy one; a
rename that is genuinely impossible (read-only home) keeps serving the
legacy directory so no data is stranded. status --all stays read-only
and reports a legacy-named partition as "active (legacy name; renamed
automatically on next command)" instead of corrupt.
Verified end-to-end with the real CLI in an isolated HOME: a legacy
.teamai migrates into the readable whole-path slug; status --all
reverse-resolves it as [active]; a partition renamed to the pre-#546
format is reported active-legacy (no rename, no corrupt), then adopted
by the next command (detection) with data intact, and pull runs through
the adopted partition.
* fix(partition): rebase repo.localPath when adopting a legacy partition
A pre-#546 partition stores repo.localPath as an ABSOLUTE path to its
team-repo clone (<legacyPartition>/team-repo). resolvePartitionDir renamed
the directory but left config.yaml pointing at the now-gone old path, so
every later `pull` read the team config from a dead directory and silently
skipped the sync ("Team config (teamai.yaml) not found. Skipping.", exit 0)
— the project could never sync again after an upgrade.
Adoption now rebases repo.localPath onto the new partition (mirroring
migrate.ts's rebaseConfigPaths). The rewrite is idempotent and self-healing:
a modern install whose localPath already sits in the canonical dir is left
untouched, an external clone outside the partition is left untouched, and an
adoption interrupted between the rename and the config rewrite is finished by
the next command.
Reproduced with the real CLI (isolated HOME, real local git team repo): a
legacy-named partition + two `pull --force` runs printed "Team config not
found. Skipping." (exit 0) before this fix; after it both runs print
"Synced 1 skills" and localPath points at the live clone.
Regression tests: localPath rebased off the legacy dir / external localPath
left alone / modern-install no-op. Verified they fail when the rebase is
disabled.
* fix(partition): write the adopted config.yaml atomically to prevent truncation
The localPath rebase overwrote config.yaml with a plain (non-atomic) write.
By that point the legacy partition has already been renamed away, so
config.yaml is the partition's ONLY copy — a write that fails partway
(ENOSPC, EFBIG, crash mid-write) truncates it with no source to recover from,
and the CLI still prints "Your original data is unchanged, re-run to retry"
while the data is in fact corrupt and cannot self-recover.
Add writeFileAtomic (same-dir temp + rename, preserving mode) alongside the
existing writeJsonAtomic, and use it for the adopted config.yaml. rename(2)
is atomic, so a failed write removes the temp file and leaves the original
config.yaml byte-for-byte intact; the next command retries the idempotent
rebase and converges.
Reproduced with the real CLI (isolated HOME, real local git team repo, NO fs
mock) by injecting a write failure with RLIMIT_FSIZE=128:
before: pull --force reports EFBIG, config.yaml truncated 453 -> 128 bytes,
legacy partition already moved, retry after lifting the limit exits
0 but never syncs and cannot recover
after: config.yaml preserved at 453 bytes, retry after lifting the limit
prints "Synced 1 skills" and localPath points at the live clone
Regression test: the localPath rewrite failing mid-write leaves config.yaml
intact with no leftover temp file. Verified it fails when the write is made
non-atomic. Full suite: 3109 passed, 0 regressions; tsc clean.
Follow-up to #502: self-update of a copy installed WITHOUT -g (flat
layout, <root>/node_modules/<pkg>) still ran `npm install -g` into npm's
default global prefix - resolveInstallPrefix only recognized the
lib/-nested global layout, returned null for flat roots, and doUpdate
logged success while the running copy stayed stale.
Flat installs come in two shapes, and neither can be reinstalled by the
CLI itself:
- project dependencies: the version is owned by the project manifest -
a self-bump diverges node_modules from package.json/lockfile and is
reverted by the next npm install/ci;
- manifest-less vendored dirs: a non-global install reconciles the root
as a project and prunes every undeclared sibling (measured on npm
11.13: "removed 144 packages" including a co-located npm; --no-save
does not prevent the prune), while -g always lands in <prefix>/lib.
resolveInstallPrefix now reports non-global (vendored) targets, and
doUpdate refuses them with explicit guidance instead of silently
updating the wrong copy. After a successful -g install, the running
entry's package.json version is verified against the registry latest -
a mismatch warns instead of claiming success (covers linked checkouts
and the null-target fallback).
Co-authored-by: flowjzh <flowjzh@users.noreply.github.com>
Project-scope machine data lives in a partition under
~/.teamai/projects/<slug>/, but the slug prefix was only the anchor's
basename (e.g. teamai-cli-<hash>), so the directory name did not reveal
which project it belonged to — unlike Claude Code's ~/.claude/projects/,
whose names encode the full path.
Build the prefix from the WHOLE anchor path (leading separator dropped,
path separators -> '-'), e.g. /Users/x/Project/app ->
Users-x-Project-app-<hash>. The 16-hex sha256 suffix is unchanged, so
uniqueness (escape-ambiguous paths, shared basenames) and the
slug===dir corrupt check are untouched; the prefix is length-bounded so
a very long path can never overflow NAME_MAX. The authoritative reverse
lookup remains the per-partition 'anchor' file.
Verified end-to-end with the real CLI in an isolated HOME: a partition
is laid down under the readable slug, 'status --all' reverse-resolves it
to the project path and reports [active], a mismatched dir name is
flagged [corrupt], and detection from the project cwd loads the
partition config via the new slug.
`enabledAgents` (from `teamai init --agent`) is documented as gating the CLI
built-in skills/rules/agents and CLAUDE.md-class injects. `deployRecallArtifacts`
deploys all four, but only the first three went through the whitelist, so
`teamai recall enable` still injected the recall block into a tool the user had
excluded — provided its root directory already existed.
Add the same `isAgentExcluded` guard the sibling deploys use. The injection loop
in `injectRecallBlockIntoTools` (src/pull.ts) is already guarded this way, so
this brings `recall enable` in line with what `pull` and `init` do.
* fix(mcp): resolve requires from PATH including Windows PATHEXT
teamai mcp inject skipped servers with requires: [uvx] on Windows because
requirementsMet spawned /bin/sh -c 'command -v', which ENOENTs when /bin/sh
is absent even if uvx.exe is on PATH. Scan PATH (and PATHEXT on win32)
instead of a POSIX shell, and keep rejecting unsafe names.
Fixes#539
* docs: link MCP requires Windows fix to PR #540
Self-update shells out to a bare 'npm' and relies on the default global
prefix; both guesses break in PATH-less subprocesses (bundled-runtime GUI
hosts) and non-default install layouts.
- resolveNpmCommand: prefer the npm-cli.js co-located with process.execPath,
probing all standard layouts (flat bundled layout, <nodeDir>/lib, and the
canonical POSIX <prefix>/bin + ../lib) before falling back to npm from PATH.
- resolveInstallPrefix: derive the prefix from this module's own path via the
<prefix>/[lib/]node_modules/<pkg> marker, sanity-checked; POSIX hands npm
the slice above lib/ (npm re-adds the component itself). Returns null for
linked checkouts so dev setups keep default behavior.
- Post-update hook refresh runs the freshly installed entry script directly
(resolveTeamaiEntryScript) instead of a bare teamai from PATH, falling back
for dev setups.
Unit tests cover the three npm layouts and the prefix extraction (POSIX lib
nesting, Windows, non-npm paths).
Co-authored-by: flowjzh <flowjzh@users.noreply.github.com>
PR #520 changed cnbExec to resolve the CLI path (resolveCliPath) and
launch it via cross-spawn instead of a bare spawnSync('cnb', ...). PR
#536 landed cnb-create.test.ts mocking node:child_process. Both were
green on their own branches, but once merged the test mocks the wrong
layer: resolveCliPath returns null in CI, cnbExec short-circuits with
'cnb CLI not found', and the canned responses never reach the code —
so cnbOrganizationExists returns false and 7 cases fail on main.
Mock cross-spawn's default.sync + utils/cli-path.resolveCliPath (the
pattern #520 already uses in cnb-login-host.test.ts). Assertions are
unchanged; only the mock injection point moves.
A pushed skill whose source has an empty subdirectory (e.g. an unused
assets/) left that subdirectory behind in the team repo working tree:
git tracks files, not directories, so when pushGroup checked out the
default branch again it removed the tracked files but could not prune
directories git never saw. The leftover shell is named after the skill
and has no SKILL.md, so scanTeamRepoNamespaces() reported it as a
namespace — and with a single candidate push silently selected it,
nesting every subsequent skill under skills/<skill-name>/<name>/,
compounding on each push.
- pruneEmptyDirs(): remove directories holding no files at any depth.
- push: prune each pushed path after returning to the default branch.
- scanTeamRepoNamespaces(): only report a directory as a namespace when
it actually holds at least one skill, so an existing leftover shell in
a repo already in this state can no longer capture new skills.
Claude-Session: https://claude.ai/code/session_01F4cW3p176qKkNwh7zWAohv
Co-authored-by: liuhaitao <liuhitao@yonyou.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The cnb login OAuth token carries neither group-manage:rw (create org) nor
group-resource:rw (create repo), so create-repo failed with a raw 403/404 that
gave users no next step.
- Detect a missing organization up front (cnb get-group, a read-only scope the
login token has) before prompting to create the repo.
- Surface OrganizationNotFoundError / RepoCreatePermissionError from the CNB
provider, each carrying the platform's web create URL.
- init now prints the create page (cnb.cool/new/groups or /new/repos) and exits,
instead of a confusing scope error.
Adds GitProvider.organizationExists?/getOrganizationCreateUrl? (CNB only) and
tests; docs updated (providers.md, usage-guide.*).
* fix(providers): resolve and launch the gh/cnb CLIs portably
`gh-cli.ts` located gh with `execSync('which gh')` and handed the result to
`spawnSync`. On Windows that lands on Git's `which`, which prints an MSYS path
(`/c/Program Files/GitHub CLI/gh`); Node resolves it as `C:\c\Program Files\...`,
so the spawn fails with ENOENT. `isGhInstalled()` therefore answered "installed"
while every `ghExec()` returned status 1 with an empty stderr — the GitHub
provider was unusable on Windows even with gh correctly installed on PATH.
`cnb-cli.ts` has the mirror-image problem: `spawnSync('cnb', ...)`. The package is
`bin: { cnb: 'bin/cnb.js' }`, so npm only writes `cnb.cmd` / `cnb.ps1` shims on
Windows — there is no `cnb.exe` for CreateProcess to find (ENOENT), and handing
the resolved `cnb.cmd` to `child_process.spawnSync` fails with EINVAL. A failed
`cnb status` came back as `cnbIsAuthenticated() === false`, i.e. "not logged in".
Both wrappers now use the resolver that #496 introduced for the AI clients, which
was previously private to `utils/ai-client.ts`:
- `src/utils/cli-path.ts` (new): `resolveCliPath` / `pickWindowsCommand`, using
the native `where` on Windows (accepting only `.exe` / `.cmd` / `.bat`) and the
`bash -lc` -> `zsh -lc` -> `which` chain on POSIX. `ai-client.ts` re-exports
both names, so existing imports and its tests are untouched.
- `gh-cli.ts`: `getGhPath()` returns `resolveCliPath('gh')`, and both invocations
launch through `cross-spawn`. Resolving alone is not enough: the picker accepts
`.cmd`, so the launcher must be able to start what the resolver can return.
- `cnb-cli.ts`: `cnbExec()` resolves first and launches through `cross-spawn` (the
same remedy the AI client uses — only a shell can start a `.cmd`). A missing CLI
now returns status 127 with `cnb CLI not found on PATH` instead of a bare status
1 and an empty stderr. `isCnbInstalled()` uses the same resolver, so "installed"
and "runnable" can no longer disagree.
`gf-cli.ts` is deliberately untouched: gf only supports macOS/Linux
(`gf-cli.ts` throws `Unsupported platform` otherwise) and its MSYS path is
consumed by `bash -c`, so the Windows resolver would break it.
Tests:
- `src/__tests__/cli-path.test.ts` (13 cases) mocks the resolver and the launcher.
- `src/__tests__/e2e/providers-cli-path.test.ts` (5 cases, no mocks) plants real
executables — an extensionless POSIX shim plus the `.cmd` shim npm writes — in a
temp dir, prepends it to PATH, and drives the real resolver and launcher. Run it
with `npm run test:e2e`; see the PR notes for why it is not CI-gated.
Reverting the launchers to native `spawnSync` fails 4 unit cases and the cnb e2e
case, the latter with the original symptom (`status=1`, empty stderr).
* test(cnb): follow the resolver in cnb-login-host
cnbLogin now launches through `resolveCliPath('cnb')` + cross-spawn, so
stubbing `node:child_process.spawnSync` no longer intercepts the call. The
suite fell through to the real binary, got a non-zero status, and both cases
failed with "cnb login failed. Please try again." on every CI runner.
The stub moves to the two layers the provider actually uses, and the
assertion moves from the bare command name to the resolved path:
`expect(cmd).toBe('cnb')` would also have passed for the direct spawn this
branch removes, so it could not tell the fix from the bug.
Reverting only the launchers in cnb-cli.ts turns both cases red again, which
is the check that the suite still holds the regression.
* test(cnb): follow the resolver in cnb-login-host
`cnbLogin` now launches through `resolveCliPath('cnb')` + cross-spawn, so
stubbing `node:child_process.spawnSync` no longer intercepts the call. The
suite fell through to the real binary, got a non-zero status, and both cases
failed with "cnb login failed. Please try again." on all four CI runners.
The stub moves to the two layers the provider actually uses, and the
assertion moves from the bare command name to the resolved path:
`expect(cmd).toBe('cnb')` would also have passed for the direct spawn this
branch removes, so it could not tell the fix from the bug. A case for the
missing CLI (status 127) comes along.
The mocks carry explicit function types: `vi.fn(() => RESOLVED_CNB)` infers a
zero-argument signature, which leaves the spread site failing type check.
Reverting only the launchers in cnb-cli.ts turns two of the three cases red
again — that is what keeps the suite honest.
Cover the two gaps in real-command e2e:
- data-layout-migration.test.ts: seed a legacy <repo>/.teamai (real
team-repo clone + plaintext env + localPath into legacy), run the
compiled CLI 'pull --force', and assert the preAction hook migrated
data into ~/.teamai/projects/<slug>/, kept the git clone intact,
rebased repo.localPath onto the partition, retired the source to
.teamai.bak, and pulled through the migrated clone.
- multi-project.test.ts: 'projects set' + 'pull --force' deploys only
the active project's skills and prunes an inactivated one; 'push
--project' routes a skill into the project's skills/<ns>/ on the
remote, matching what pull reads.
Hermetic: sandbox HOME + local bare remote (provider: git), no network,
tokens, or PR creation. Land in src/__tests__/e2e/ so vitest.e2e.config
picks them up.
The cnb CLI infers its platform URL from the first git remote of the
current directory when --host is not given. Running `cnb login` inside a
repo whose remote points at a non-CNB host (e.g. an internal git server)
therefore sends the OAuth2 device-auth request there and fails with 401,
while the same command succeeds in a directory with no such remote.
Pass --host ${CNB_HOST} explicitly. CNB_HOST (TEAMAI_CNB_HOST, default
cnb.cool) is already the single source of truth for every other CNB
operation (clone / create-repo / PR), so anchoring login to it keeps the
provider's auth consistent and self-hosted-friendly.
`teamai init` cloned CNB repos via `git -c credential.helper='!cnb git-credential'
clone`, but the `-c` flag only applies to that single invocation — it never
reached the cloned repo's `.git/config`. So `remote.origin.url` stayed
credential-free, and the subsequent push (member registration) plus `teamai pull`
fell back to an interactive Username/Password prompt despite the user being
logged into the `cnb` CLI.
GitHub/TGit solve this by embedding the token in the clone URL (persisted into
remote.origin.url); CNB's CI path (CNB_TOKEN) does the same. Only the
interactive-login path was missing persistence. Persist the helper into the
repo's local config after a successful interactive clone so every later git
operation authenticates transparently.
Verified end-to-end against a real cnb.cool account (logged-in, no CNB_TOKEN):
`teamai init https://cnb.cool/test1122444/test` now pushes member registration
without prompting, and `teamai pull` reports up to date instead of
"couldn't find remote ref".
Follow-up to #504. Resource handlers already use isAgentExcluded, but
builtin skills/rules/agents, CLAUDE.md-class injects, lastPullTargets,
and pull cleanup still gated only on disabledAgents. An already-installed
tool outside init --agent kept receiving writes, and adding a tool to the
whitelist with unchanged HEAD skipped the first team-resource sync.
Cleanup of leftover copies on out-of-whitelist trees is skip-not-delete.
Closes#510
Extract now writes teamwiki/evidence/code/<project>/_manifest.json from
module or component facts when AI enrichment is skipped, fails, or finds
no qualifying modules. Public empty-manifest errors no longer tell the
user to re-run extract. Hidden teamai deep-enrich exits non-zero when
there are no components.
Fixes#508.
`refreshTeamRepo` reports a failed submodule update so the caller holds
`lastPullRev` back and retries next pull (#501). It does not cover the
successful update: `getInstalledResourceTargets` only records which AI
tools are installed, not which skills appeared, so filling an empty
submodule leaves the parent SHA — the fast path's only cache key —
untouched and the unchanged-rev fast path skips the deploy.
That is the upgrade path in #525: a member pulls with a CLI that predates
`submodules: true` (the unknown yaml key is stripped), so the parent SHA
is cached with the submodule directories still empty. Upgrading the CLI
and pulling fills the clone, then reports "Already synced at <rev>,
skipping" and the tool directories stay empty until `pull --force` or an
unrelated parent commit.
Compare `git submodule status` before and after the update and carry the
result into the fast path as `submodulesChanged`. The leading status char
is `-` while uninitialized and ` `/`+` once checked out, so a changed
string means the on-disk tree the deploy step reads is not what was
cached. The fast path only skips when the update left the tree alone, so
a submodule-using team still takes it in steady state instead of paying a
full re-deploy on every pull.
An unreadable status degrades to "changed": a redundant full sync is
cheap and self-correcting, whereas wrongly skipping re-pins the empty
tool directories this fixes. It is only the status read that degrades —
an update failure still reaches the outer catch and holds the rev back.
Fixes#525
Teams that distribute skills as git submodules get an empty directory after
clone/fetch - the submodules are never populated, so every resource deploy
silently misses their content.
Add a teamai.yaml knob (off by default):
submodules: true
After each repo refresh (and before the resource deploy step, so freshly
checked-out content is what gets deployed), pull runs
`git submodule update --init`. Deliberately not shallow: submodules are
pinned to exact SHAs, and a shallow fetch only brings the remote tip, so
checking out any older pin fails with "reference is not a tree" - the full
history guarantees the pinned commit is always present.
A failed submodule update does not abort the pull (best-effort warn), but it
also must not persist the new rev: the parent-repo SHA is the
incremental-sync cache key, and caching it while the tree is incomplete
would let the unchanged-rev fast path suppress the retry on every later
pull - tool directories would stay empty forever. refreshTeamRepo therefore
reports submodulesFailed and pullForScope skips the rev persistence for that
run, so the next pull re-runs a full sync and retries the update.
Documented in docs/usage-guide (en + zh-CN), including the auth caveat for
private submodules on token-injecting hosts. Unit tests cover the
submodulesFailed -> rev-not-persisted invariant.
Co-authored-by: flowjzh <flowjzh@users.noreply.github.com>
GitHub's automatic base selection takes the newest tag reachable from the
tagged commit. Release commits are created detached and never land on a
branch, so no release tag is an ancestor of the next one and every recent
release compared back to v0.22.0 — the last release that happened to be
merged into main. The 0.23.x stables are worse: they were cut from an
internal mirror sync, so their merge-base with main sits at v0.17.4 and
their notes span three months of unrelated PRs.
Resolve the base explicitly and feed it to the generate-notes API. A
prerelease compares against the tag before it; a stable release compares
against the previous stable tag. Both skip tags whose parent is absent
from this release's history, so a mirror-sync tag can no longer poison
the range.
Co-authored-by: Cursor <cursoragent@cursor.com>
detached: true without windowsHide: true gives the child its own console
window on Windows — every hook dispatch flashed a transient "npm" window
when WorkBuddy (a GUI host) fired session hooks. Set windowsHide: true on
both detached spawn sites (hook-dispatch background child and the
local-agent plugin reconcile worker).
Co-authored-by: flowjzh <flowjzh@users.noreply.github.com>
Two paths ignored the team's --agent opt-in and deployed into every
configured tool directory:
- hooks inject reconciled every configured tool - re-seeding injections
for tools the team never opted into (e.g. unconditional hermes
writes). Route it through reconcileTeamHooksForConfig so it applies
the same enabledAgents whitelist minus disabledAgents scoping as
pull and init.
- isAgentDisabled only checked disabledAgents, so the per-tool resource
handlers (skills/rules/docs/agents) copied skills into e.g.
~/.hermes/skills. Add isAgentExcluded (disabled or outside the
whitelist) and use it in the handlers, mirroring the hooks gate.
Co-authored-by: flowjzh <flowjzh@users.noreply.github.com>