72 Commits
Author SHA1 Message Date
Saul Moro 85ff72c775 feat(recall): link an OMP subagent's session to its parent for adoption (#935)
* feat(recall): link an OMP subagent's session to its parent for adoption

OMP gives a subagent a session of its own, so a recall the teamai-recall
subagent ran never credited the main agent's later reads. OMP writes a
subagent's session to <parent>/<agent id>.jsonl beside the parent's
<parent>.jsonl, whose header names the parent session. The generated
extension now reads that header on a session's first tool_result and
sends the session_link the OpenCode plugin already sends, so the reducer
credits the root session. A session kept only in memory has no file and
still gets no link.

* fix(recall): keep looking up an OMP subagent's parent until it is linked (PR review)

A session was marked as looked up before its parent was found, so a parent session file not yet on disk lost the link for the whole session. Also: skill-data troubleshooting lists OMP's subagent path, the docs name the no-link case (--no-session), a row shows a linked subagent's own read still scores 0, and the header read limit is named.

* fix(recall): read an OMP parent's session header after its title slot (PR review)

OMP writes a padded title slot line before the session header, so the
parent lookup parsed the title line and never linked a subagent. The test
harness now writes OMP's real layout; verified in a live OMP 18.4.8
session.
2026-10-01 20:27:50 +08:00
Saul Moro 9bcd53e915 fix(rules): give Codex the team rules: session hooks in a project, its own AGENTS.md in user scope (#940)
* fix(rules): inline rules without frontmatter, naming path-scoped globs

hermesRulesText becomes inlinedRulesText, the one renderer for tools that read rules from a single instructions file. It strips frontmatter and leads a path-scoped rule with 'Applies to files matching: <globs>'. Hermes SOUL.md and doctor's Hermes check use it. Part of #938.

* fix(rules): inline team rules into Codex AGENTS.md in project scope

Codex reads instructions from AGENTS.md, not from .codex/rules, so the
copies pull wrote there never reached it (#938). The codex defaults drop
rules and gain claudemd (AGENTS.md, ~/.codex/AGENTS.md at user scope).
pullAllRules syncs a [teamai:team-rules] block into that file once per
path while an enabled, installed Codex-family tool maps it, and removes
it otherwise. syncManagedInstructions probes rules ?? settings, so a
machine without Codex gets no ~/.codex.

* fix(rules): render the Codex team-rules block with inlinedRulesText

Follows the T3 rename, so a path-scoped rule reaches AGENTS.md without frontmatter, after its Applies to line.

* test(rules): cover Codex AGENTS.md in user scope, culture, shared instructions and recall

Codex now reaches AGENTS.md through its default claudemd path; these tests
pin user scope (~/.codex/AGENTS.md), a full project-scope pull writing all
four teamai blocks into one AGENTS.md, and recall on/off. Also corrects the
injectRecallBlockIntoTools comment, which listed Codex as skipped.

* fix(rules): remove Codex team rules and their old copies on uninstall

Uninstall strips the [teamai:team-rules] block and scans the legacy .codex/rules directory for teamai's copies. A shared instruction file now keeps only the blocks a remaining enabled, installed tool would write, so --agent codex drops the team-rules block while Pi keeps its own.

* fix(doctor): check the team-rules block in Codex AGENTS.md

Codex reads no rules directory, so doctor now compares the team-rules
block in each AGENTS.md an enabled, installed Codex-family tool maps
with what pull renders. It fails on a missing or stale block, on an
AGENTS.override.md that Codex reads instead, and on an entry with no
claudemd path. The file list comes from RulesHandler, the same
resolver the pull writes through.

* fix(rules): remove the .codex/rules copies earlier pulls delivered

Codex never read the <rule>.md copies teamai wrote to .codex/rules (and
.codex-internal/rules, .tcodex/rules). Every rules sync now removes the
ones that still hold what teamai delivered, including teamai-recall.md,
keeps edited copies and names them, and leaves *.rules alone.

* docs(rules): document where Codex reads team rules (AGENTS.md)

* fix(rules): give Codex its AGENTS.md rules on an already-synced pull

A CLI upgrade with the team repo unchanged takes the "Already synced" fast path, which never ran the rules sync. So a member upgrading from 0.22.0 got no team-rules block in AGENTS.md and kept the old .codex/rules copies until the repo moved or they ran pull --force. The fast path now runs the Codex part of the rules sync (block + reclaim) and nothing else.

* fix(doctor): tell a Codex member a plain pull restores the team-rules block

The already-synced pull now rewrites the block, so the fix text no longer sends the member to pull --force.

* fix(rules): give codex-internal and tcodex the codex instructions layout

codex-internal and tcodex drop their rules path and read team rules from
AGENTS.md like codex. Codex-family instruction writers probe
rules ?? settings ?? skills, so a team entry with neither rules nor settings
no longer counts as installed. writesInstructionBlock is the one answer to
which blocks a tool gets, shared by pull's writers and uninstall.

* fix(rules): keep an empty AGENTS.md that held no team-rules block

removeClaudeMdSection now reports whether it removed a section and can delete a file left holding only whitespace. The team-rules sync and recall off use it, so a pull that removed nothing no longer deletes a member's own empty AGENTS.md.

* fix(doctor): pass the Codex team-rules check when no rule has a body

When every team rule is frontmatter only, pull writes no team-rules block, so a missing block is correct and an AGENTS.override.md shadows nothing. Also note that an empty override shadows AGENTS.md too.

* fix(doctor): a plain pull applies the Codex claudemd toolPaths fix

Editing teamai.yaml moves the team repo, so a plain pull does the full sync; --force is not needed. Matches the other Codex block checks.

* fix(uninstall): name the instruction file uninstall cleans

The summary and result lines said CLAUDE.md for every instruction file, including Codex's AGENTS.md. They now print the file's path under an 'Instruction-file blocks' heading.

* refactor(recall): gate the recall block on writesInstructionBlock

recall on now asks the same question as pull's recall writer and uninstall, so the three cannot drift. No behavior change.

* fix(rules): reclaim .codex/rules copies of rules the team removed

The tombstone cleanup walks toolPath.rules, which Codex no longer has, so a copy of a rule tombstoned after the member's last pre-#938 pull stayed forever. The legacy reclaim now removes tombstoned names too, keeping and naming a copy the member changed since delivery (#822).

* test(rules): cover codex-internal and tcodex in reclaim, doctor and uninstall

Parametrizes the project-scope legacy reclaim, the four Codex doctor failure cases and the full uninstall over the Codex family.

* test(pull): cover the project-scope Codex fast path after an upgrade

* test(local-agent): pin the prompt sync into Codex's AGENTS.md

With Codex's default claudemd (#938), an HTTP prompt command reported by Codex writes the shared-instructions block into ~/.codex/AGENTS.md when ~/.codex exists, and creates nothing when it does not.

* refactor(rules): move rulePaths into rule-format

It reads a tool-neutral rule's paths: frontmatter, which both the Copilot renderer and the inlined-rules renderer use; it no longer lives in the Copilot module.

* fix(doctor): ask nothing of a Codex entry that delivers no rules to it

A team toolPaths entry such as codex: { agents: .codex/agents } has neither
rules nor claudemd: the team sends Codex no rules on purpose, like any tool
without a rules path. Doctor failed it on every member's machine (E2E
doctor-delivery-cli). It now fails only an entry that still lists the
pre-#938 rules path without claudemd. Part of #938.

* fix(rules): keep a teamai-recall.md the member edited since the last pull

The legacy Codex reclaim removed any .codex/rules/teamai-recall.md that
opened with the recall heading, so text a member appended after the last
pull that wrote it was deleted. The built-in rule is not in the delivery
ledger, so match the file against the exact bodies teamai deployed
instead; an edited copy is kept and named like other legacy copies.

* fix(rules): keep a member's empty AGENTS.md when the team-rules block goes

Removing the last block deleted any instructions file it left empty, so
an empty AGENTS.md a repository tracked was deleted once Codex was
disabled or the team had no rules left. injectClaudeMdSection now creates
a file as just the block, and deleteIfEmpty deletes only a file that
opens with it; a member's empty file is written back empty.

Also describe the per-block cleanup of shared instruction files in the
deployed uninstall skill.

* fix(rules): address review on uninstall, remove, doctor and markers

- uninstall and the legacy [teamai:rules] strip remove blocks through
  removeClaudeMdSection, so a member's empty AGENTS.md survives them as it
  does pull. Removing a block at the top of a file keeps the next one
  there, so a file teamai created still goes with its last block.
- `remove rules` refreshes with the rules this member's pull delivers
  (role, project and tag selection) instead of every rule in the repo.
- doctor reports a Codex team-rules block left behind when the team has
  no rules, and nothing else in that case.
- teamRulesBlock strips the block markers wherever they appear in a rule,
  not only as whole lines, since the block is located by substring.

* refactor(rules): split legacy Codex rule-copy ownership out of the reclaim

legacyRuleCopies classifies each .codex/rules copy (and the codex-internal
and tcodex dirs) as teamai's or edited, on the ledger, team-render,
tombstone and shipped-recall proofs reclaimLegacyRuleCopies used inline.
Read-only and public so uninstall can apply the same check. No behavior
change.

* fix(uninstall): keep edited legacy Codex rule copies, remove removed rules' copies

Uninstall picked every .md in a legacy rules dir whose name matched a
current team rule or a built-in, so it deleted copies a pull had kept as
the member's edits, and left copies of rules the team had tombstoned.
It now takes the legacy dirs from RulesHandler.legacyRuleCopies, the
check pull reclaims by: owned copies go, edited ones stay and are named
in one English warning. The tool's current rules dir is unchanged.

Addresses the codex-review findings on #940.

* fix(rules): give Codex the team rules through its session-start hook

The project AGENTS.md the earlier revision wrote the rules into is read by
Cursor, Copilot, OpenCode and Claude too, which already get the rules in
their own format. The Codex family now gets them from the teamai
session-start hook instead, and pull writes no rule file for it.

- A team-rules handler adds the user-scope rules, then the cwd project's,
  as pull resolves them, skipping a scope that does not enable the tool.
  It adds nothing on resume, whose history already holds them.
- A SubagentStart entry gives a fresh subagent the same rules.
- Both entries set additionalContextLimit: 0; past 2,500 tokens Codex
  keeps only the start and end of a hook's context.
- doctor checks the limit on the session-start entry; the AGENTS.md
  block check, its markers and its uninstall handling are gone. The block
  never shipped.
- Culture, shared instructions and recall still go to AGENTS.md; the old
  .codex/rules cleanup and the remove-rules role filter are unchanged.

For #938.

* test(tool-roots): expect Codex's instruction file under CODEX_HOME

#942 pinned Codex's default paths before #938 moved its rules to the
session-start hook: no rules path, and an AGENTS.md that follows the
root in user scope.

* fix(rules): keep Codex's project content out of AGENTS.md (#945)

The project AGENTS.md is the owners' file, and Cursor, Copilot, OpenCode
and Claude read it too. The Codex family now takes its project content
from the session hooks and its user content from its own AGENTS.md:

- Project scope: no `claudemd` in the Codex-family defaults, so pull
  writes nothing to the project AGENTS.md. The session-start and
  subagent-start hooks add the project's rules plus the culture,
  claudemd/ and recall blocks, leaving out a block another tool already
  wrote into the project AGENTS.md.
- User scope: the team rules go into a team-rules block of
  ~/.codex/AGENTS.md (and the variants' homes), beside the other blocks;
  the hook adds nothing outside a project. The "Already synced" pull
  writes the block too.
- doctor checks that block in user scope, and both hook entries' limit
  in a project. uninstall removes the block.

For #938.

* fix(rules): keep tombstoned Codex copies without delivery records
2026-10-01 20:26:31 +08:00
Saul Moro b2d3598b38 fix(config): honor CODEX_HOME through toolRoots (#942)
* fix(config): honor CODEX_HOME through toolRoots

Codex keeps its home in $CODEX_HOME, but teamai resolved every Codex
path under ~/.codex, so a member who exports the variable got no
skills, hooks, agents or MCP in the directory Codex reads, and doctor
stayed green.

Extend the toolRoots relocation built for CLAUDE_CONFIG_DIR (#725) to
Codex. A RELOCATABLE_TOOLS table maps each tool to its variable; init
records CODEX_HOME into toolRoots.codex, a re-init that moves the root
releases the old hooks and managed MCP servers, and doctor adds a
"Codex root matches CODEX_HOME" check. The two Codex writes outside
toolPaths now follow the root: the co-author setting in config.toml and
the skill-existence probe used by usage tracking.

The doctor check compares the paths each root resolves to rather than
requiring a record, so an unrecorded CODEX_HOME=~/.codex passes while
an unrecorded CLAUDE_CONFIG_DIR=~/.claude still fails (it moves
.claude.json).

Closes #939

* fix(config): address review on CODEX_HOME relocation

Derive the Codex co-author file from the scoped paths, list relocatable
skill roots from RELOCATABLE_TOOLS, and skip the root-release warning
for a tool the member does not sync or whose old root does not exist.

* fix(coauthor): resolve the Codex config root, not the skills mapping

A team mapping toolPaths.codex.skills to .agents/skills sent the Codex
co-author setting to ~/.agents/config.toml. Resolve each Codex's own root
(toolRoots.codex or ~/.codex, ~/.tcodex, ~/.codex-internal) instead.
2026-10-01 16:30:07 +08:00
Saul Moro 9d8ef7d81f fix(tests): clear agent root env vars once in the vitest setup file (#936) 2026-10-01 16:28:34 +08:00
Saul Moro 3030ffd8a8 fix(test): give each unit test file its own HOME (#927)
Commands resolve ~/.teamai from HOME, so a unit test that reached the real
home shared the sync lock with every teamai session on the machine.
push-env.test.ts failed while another session held ~/.teamai/.sync-lock
(#924). A setup file, listed first so module-level home paths see it too,
now points HOME/USERPROFILE at an empty temp dir per test file. The e2e
config keeps the real home, which CI prepares.
2026-10-01 14:48:05 +08:00
Saul Moro fb9f5a2c28 feat(mcp): keep project MCP configs with resolved tokens out of git (#882) (#886)
* feat(mcp): keep project MCP configs with resolved values out of git (#882)

A project-scope MCP config that carries a resolved ${VAR} sat untracked
and unignored in the business repo, one `git add -A` from committing the
token. After the reconcile writes such a file and git would track it,
teamai lists its path in the clone's .git/info/exclude inside a marked
block (resolved via `git rev-parse --git-path`, so linked worktrees and
submodules work). The committed .gitignore is never touched; an ignored
path or a config with no resolved value adds nothing; dry runs write
nothing. Project-scope uninstall removes only teamai's block, and doctor
reports such a file git would still commit.

The hook sits after the appliers in reconcileMcpForConfig, outside
desiredMcpForTarget/applyJson/applyCodex, so it merges cleanly with #880.

* fix(uninstall): count the .git/info/exclude block in the removal plan (#882)

The plan now records whether the project's .git/info/exclude holds
teamai's MCP config block (gitExcludeBlock). It counts toward
isPlanEmpty, is listed in the summary and dry run, and gates the
removal, so a plan whose only teamai leftover is the block removes it
instead of reporting "Nothing to uninstall".

* fix(mcp): keep the exclude block until every MCP config is clean (#882)

- uninstall keeps a repository's .git/info/exclude block while a config in
  it could not be parsed and still holds teamai servers, and warns
- uninstall finds and removes the block in nested repositories holding an
  MCP config, across every worktree
- the block opens at the last start marker, so an orphaned start never
  pairs with a later block's end and takes the member's lines
- doctor counts only servers the ownership manifest records, not a member's
  own server under a team name

* fix(uninstall): remove the exclude block only once its files are proven clean (#882)

A missing or unreadable managed-mcp.json made the MCP cleanup return early
without reporting anything, so uninstall removed the block while .mcp.json
still held the resolved token.

Uninstall now inspects every path the block protects after the cleanup. The
block goes only when each one is missing, or parses and holds none of the
team's servers that need a resolved ${VAR}. Anything it cannot check keeps
the block, with a warning naming the file. This replaces the leftInPlace
report from the reconcile, which the check subsumes.

* fix(mcp): protect every project MCP config holding a resolved value (#882)

Pull, doctor and uninstall each skipped a case they had not inspected and
treated it as safe. Now:

- pull lists a config in .git/info/exclude whether or not it delivered to
  it this run: a disabled or undetected tool's file, a team with automatic
  delivery off, an unreadable mcp.yaml (any teamai entry counts), a failed
  write to another tool's config, and a lost ownership manifest (the
  resolved value found in the file)
- doctor checks the same files, including one that does not parse, and
  counts a git error as a failure
- git check-ignore failing inside a repository is no longer read as "not
  tracked": the path is excluded anyway, or teamai warns with git's error
- uninstall also keeps the block while a file contains the value (8+
  characters, not a path or the login name) of a variable still set in the
  environment, which finds a server since dropped from mcp.yaml

* fix(mcp): judge MCP configs by disk and manifest, not current config (#882)

Pull, doctor and uninstall still decided "clean" from the current team
config in places. Now one function, resolvedValueEvidence, decides for all
three:

- a teamai-owned entry still in the file counts when its server has left
  mcp.yaml, as well as when it needs a resolved ${VAR} or mcp.yaml cannot
  be read (doctor no longer skips that case)
- targets include the built-in location of a tool the team dropped from
  toolPaths or moved
- doctor names a file two tools share once
- exclude updates take the existing acquireLock helper, re-read the file
  and write it atomically, so concurrent commands keep each other's paths
- uninstall inspects every worktree of each repository owning a block,
  including a nested repository's linked worktrees, and applies the
  manifest rule per worktree
- the kept-block warning names each file and why, such as the variable
  whose value matched

* fix(mcp): skip the .git/info/exclude write while another command holds its lock (#882)

After the 2.5 s wait for the exclude file's lock, updateExclude wrote without
it, so two writers could drop each other's pattern and leave a plaintext MCP
config committable. It now writes nothing and reports 'locked': pull warns that
the file is not excluded yet and to run `teamai pull` again (doctor's exclude
check keeps reporting it meanwhile), and uninstall keeps the block and warns.

* fix(uninstall): keep an exclude entry unless its MCP config is proven free of teamai's servers (#882)

Uninstall judged a protected file clean from the current mcp.yaml, manifest
and resolvable values, so with the manifest lost, the server gone from
mcp.yaml and its value unset, a plaintext token looked like the member's own
server and the exclusion went. It now fails closed and works per entry: a
pattern goes only when its file is gone, holds no server, or holds none of
teamai's servers with managed-mcp.json still there to say what teamai wrote.
A kept entry is named with its file, why, and how to clean it by hand, since
a rerun of uninstall finds no config after a full uninstall.

* fix(mcp): exclude a project MCP config from git before writing a resolved value into it (#882)

Pull listed the file in .git/info/exclude only after writing the plaintext,
and a failed exclusion only warned, so the secret-bearing file stayed
eligible for git add -A. The exclusion now comes first; when it cannot be
established (exclude file or .git/info not writable, lock held past the
wait, file already tracked, git error) the file is left as it was and the
warning names the reason and the fix.

* fix(mcp): report a server withheld from a file git would commit in mcp list and doctor (#882)

* fix(mcp): report a tracked file on a dry run and a withheld server already installed (#882)

Backports #880's merge 0d9f7fa7: a dry run (doctor, mcp list) names a
tracked file before any pull has listed it, mcp list reports withheld for a
server an earlier pull installed, and doctor's withheld note carries the
exclusion's own fix instead of the pull --force advice.

* fix(mcp): name a tracked MCP config before an unwritable .git/info/exclude (#882)

A tracked file needs `git rm --cached` whatever else is wrong, so
ensureExcludedFromGit checks gitTracks before the writability check, on a
pull and a dry run alike, and lists nothing for it.

* fix(mcp): take a project MCP config's exclude line back out once it holds no resolved value (#882)

A pull that lists a config in .git/info/exclude and then writes no value
into it (it does not parse, a member's server holds the team's name, the
write fails) removes the line it added. After a pull or `teamai mcp
remove`, a line whose configs are proven clean in every worktree, by the
proof uninstall uses (moved to mcp-reconcile.ts), is removed under the
lock; one not proven clean stays. A config listed before its write is
listed again after it, so a concurrent uninstall that dropped the line
between the check and the write does not leave the value unprotected.

* fix(mcp): judge a project MCP config by the manifest as it stood before the pull rewrote it (#882)

A pull whose manifest was lost before it ran recreates managed-mcp.json
while reconciling, so the clean-file proof read the new record and took
the exclude line out of a file still holding a teamai server that left
mcp.yaml with its variable unset. The proof now uses this worktree's
manifest as read before the reconcile.

* fix(mcp): log a rolled-back exclude line at debug level (#882)

A line this pull added and took back out, because it wrote no resolved
value into the file, was reported as removed although the member never
saw it added. Only removing a line an earlier run added stays at info.

* fix(mcp): keep a shared exclude line while another worktree's config holds a server (#882)

A pull or `teamai mcp remove` judged every linked worktree's MCP config
with today's definitions and values. Once a server's ${VAR} became a
literal, and the value was no longer set, a pull in worktree A took
worktree B's stale token-bearing entry for clean and removed the shared
/.mcp.json line, so `git add -A` in B staged the token.

These commands now release a line only when the current worktree's file
passes the full proof and every other worktree's file is missing or holds
no MCP server. `teamai uninstall` keeps its full proof in each worktree.

* test(mcp): build the other worktree from the real temp path so the test checks what it names (#882)

* fix(uninstall): list no worktrees for a project root that no longer exists (#882)

buildRemovalPlan now lists every worktree to find teamai's exclude blocks,
outside the MCP cleanup's try. simple-git throws synchronously for a missing
directory, so a project uninstall whose root is gone crashed; main's #878 test
caught it after the merge.

* fix(mcp): treat an empty, unparsable or tool-less managed-mcp.json as no record (#882)

The clean-file proof took any managed-mcp.json on disk as teamai's record,
so an empty or truncated one let a pull, `mcp remove` or uninstall judge a
file still holding a stale secret-bearing entry clean and drop its exclude
line. A file now counts as recorded only when the manifest parses and holds
an entry for that tool's file. A project record teamai empties stays as []
so a file left with only the member's own servers can still be released.

* fix(mcp): keep the exclude line of a config this pull wrote when a later step fails (#882)

* fix(mcp): read a check-ignore error as unsafe unless ls-files proves the config untracked (#882)

* fix(mcp): report withheld only for targets delivery would write the server to (#882)

* docs(mcp): describe the .git/info/exclude block in the setup skill and the stricter git check (#882)

* fix(mcp): keep the exclude line of an entry a pull wrote with a resolved value after its definition turns literal (#882)

* fix(mcp): judge a nested repository's linked worktree config by its sibling's tool (#882)

* feat(mcp): record the project MCP configs a pull wrote a resolved value to in managed-mcp-files.json (#882)

* fix(mcp): keep protecting a config a pull wrote under a toolPaths mapping the team has since changed (#882)

* fix(mcp): keep a config's exclude line past the pull that rebuilt its lost record (#882)

* test(mcp): pin today's exclude rules for a missing, corrupt or locked managed-mcp-files.json (#882)

* docs(mcp): describe managed-mcp-files.json and the configs it keeps protected (#882)

* refactor(mcp): keep the #882 record edits off the lines #880 changes (#882)

* fix(mcp): protect a config an older teamai wrote under a mapping an earlier teamai.yaml made (#882)

A teamai from before managed-mcp-files.json kept no record of the path it
wrote a resolved value to. Once the team changed that toolPaths mapping, no
pull visited the file. The first pull on this version now reads every
mcpProject path the team repo's history of teamai.yaml mapped, once per
worktree: a file under the project root that no current mapping or record
reaches, and that holds a resolved value, is listed in .git/info/exclude and
recorded. A git error leaves the read for the next pull; a shallow clone
reads the history it has.

* fix(mcp): keep a rebuilt record from persisting without its note of the file's other servers (#882)

When a pull rebuilt a lost managed-mcp.json and could not note the other
servers in the file (managed-mcp-files.json locked, an I/O error), it still
wrote the rebuilt record, so no later pull knew the record was rebuilt and a
stale server's line could go. The same manifest write now marks those records
unnoted: the file counts as having no record, so it keeps its line while it
holds a server, and the next pull notes them and clears the mark.

* fix(mcp): take back a managed-mcp-files.json record for a config the pull then did not write (#882)

A pull records a config before writing a resolved value to it. When the
write failed or did not happen (the file does not parse), the record stayed,
and once the mapping changed a config of the member's own at that path was
kept excluded while it held any server. The pull now takes back a record it
added for a file it did not write, as it does the file's exclude line; the
settle after records it again if the file holds a resolved value anyway.

* docs(mcp): describe the teamai.yaml history read, the unnoted rebuilt record and the record a failed write takes back (#882)

* refactor(mcp): keep the r3 edits off the lines #880 changes (#882)

* fix(mcp): also protect a config an older teamai wrote under a built-in default it has since changed (#882)

* fix(mcp): judge a project MCP config under a symlinked directory where the write lands (#882)

The appliers replace the file itself (tmp + rename) but follow its
directories. Every git check now judges realFilePath(file), the one
resolver the release keying already used: a directory linked out of any
repository no longer withholds the servers on git's "not a git
repository", and a tracked file there is named with both paths, with a
git rm --cached that works (git refuses the path through the link).

* docs(mcp): describe how a config under a symlinked directory is kept out of git (#882)

* refactor(mcp): keep realFilePath next to existingAncestor, without an import cycle (#882)

* fix(mcp): judge a config under an earlier teamai.yaml mapping as a recorded file, not by today's records (#882)

* fix(mcp): have doctor check the configs earlier teamai.yaml mappings reach until a pull reads them (#882)

* docs(mcp): describe how a config under an earlier teamai.yaml mapping is judged, and doctor's check of it (#882)

* fix(mcp): record a config under an earlier teamai.yaml mapping that git tracks, and judge it once git no longer does (#882)

* fix(mcp): keep judging a recorded config for a tool the team moved while another tool still maps it (#882)

* docs(mcp): describe the tracked config an earlier mapping reached, and a moved tool's config another tool still maps (#882)

* fix(mcp): find a config an older teamai wrote under an earlier mapping another tool maps today, and judge it by that tool's records (#882)

* docs(mcp): describe the history read's configs another tool maps today (#882)

* fix(mcp): prove a shared config clean only while every tool that wrote a resolved value there has its record (#882)

* fix(mcp): judge a built-in location no mapping reaches today as an earlier-mapped file, and a shared config no pull recorded by every tool mapping it (#882)

* docs(mcp): describe the built-in location of a moved or dropped tool, and a shared config no pull recorded (#882)

* fix(mcp): name the ignore rule that re-includes a config teamai just listed, instead of saying git tracks it (#882)

* fix(mcp): judge a moved tool's built-in location another tool maps for that tool too, and hold a config's line while no managed-mcp.json claims its servers (#882)

* docs(mcp): describe the no-manifest rule, a moved tool's built-in location another tool maps, and a re-including ignore rule (#882)

* fix(mcp): keep a tool's record as it was when its config does not parse, and take back each tool a write that did not happen recorded (#882)

* fix(mcp): mark the records a pull with no managed-mcp.json writes as unnoted until the servers no record claims are noted (#882)

* fix(mcp): have doctor judge a record marked unnoted like a missing managed-mcp.json (#882)

* fix(mcp): judge a config tools of different formats share in each of their formats before releasing its line (#882)

* fix(mcp): note the servers no record claims in the file of a tool whose record a pull writes first, as with no managed-mcp.json at all (#882)

* fix(mcp): take back a tool managed-mcp-files.json recorded before a write unless its own records hold a resolved value there (#882)

* fix(mcp): treat an installed tool's missing record as lost at every pull and in doctor, not only an empty managed-mcp.json (#882)

* fix(mcp): note the unclaimed servers under every format of a shared config, and pin an uninstalled tool's leftover config (#882)

* fix(mcp): count an uninstalled tool's missing record when no installed tool maps its file, and name a re-including .gitignore rule on a dry run (#882)

* fix(mcp): scope a shared config's claims to the tools reading the same key, and settle its notes on every format's view (#882)

* fix(mcp): keep suspect an uninstalled tool managed-mcp-files.json lists as a writer, though an installed tool maps the file (#882)

* fix(mcp): judge a moved tool's file by the records of the tools reading the same key today (#882)

* fix(mcp): read a Copilot project config's bare servers beside the mcpServers another tool added, and remove teamai's there (#882)

* fix(mcp): keep a project config the local agent writes a header or env value to out of git, and have doctor check it for HTTP teams (#882)

* fix(mcp): document the local agent's project-scope exclusion and Copilot's bare servers beside mcpServers (#882)

* fix(mcp): tell the local agent's withheld install to install the MCP server again, not to pull (#882)

* fix(mcp): replace a Copilot server's bare copy when pull or the local agent writes it again under mcpServers (#882)

* fix(mcp): count a local-agent install's arguments, URL user or query and command line as credentials, and scope the rebuild's claims by key (#882)

* fix(mcp): count any URL in a local-agent install as a credential, and have doctor judge a recorded HTTP-team file with no record (#882)

* fix(mcp): read OpenCode's command array, remove only teamai's own bare Copilot copy, hold a shadowed one, and judge a partly lost HTTP record by its entries (#882)

* fix(mcp): list a project config an older local agent wrote a credential into on the next sync or pull of an HTTP-backed team (#882)

* fix(mcp): document that the local agent's sync and pull list an older install's credential file (#882)

* fix(mcp): have the local agent's sync say a new session tries again and call the file's content a credential (#882)

* fix(mcp): run the local agent's sync protection after an uninstall_teamai too: a failed or partial one leaves what to keep out of git (#882)

* fix(mcp): judge a local-agent entry a failed write left by what it holds, not by the new install's resolved: false (#882)

install_mcp records the new entry before writing the config. When that
write fails, the older entry, credential included, stays in the file
while the record says resolved: false. localAgentCredentialFiles now
trusts resolved: false only while the entry on disk has the recorded
hash; otherwise it checks the entry itself. The doctor fixture that
used a placeholder hash now records the hash an install writes.

* fix(mcp): keep a moved Copilot's stale bare server apart from a name another tool owns under mcpServers (#882)

recordedMcpFileEvidence merged a Copilot project file's bare servers with
those under mcpServers by name, so Claude owning a nested jira read as
owning a bare jira Copilot wrote with a token before it moved. Bare
servers now count as owned only by a Copilot record: no other tool
writes there.

* fix(mcp): have the local agent's protection read CodeBuddy's former default and a bare Copilot entry apart (#882)

An HTTP team has no teamai.yaml history, so localAgentCredentialFiles never
visited a built-in default teamai has since changed: a credential a local
agent from before 57636a27 wrote to .codebuddy/mcp.json stayed committable.
It now also reads EARLIER_BUILTIN_MCP_PROJECT (earlierMappedMcpTargets with
history: false), in sync and doctor alike.

It also judged a Copilot project file through installedMcpEntries, which
merges bare servers with mcpServers by name, the keyed one winning. A
tokenized bare entry beside a credential-free mcpServers entry of its
name, both under one record, went unseen. Each entry is now judged on its
own (mcpEntriesByPlacement, which recordedMcpFileEvidence uses too).

* fix(mcp): prove bare ownership and persist installs before exclusions (#882)

* fix(mcp): keep bare ownership from claiming keyed Copilot entries (#882)

* fix(mcp): require evidence for unmarked Copilot ownership (#882)

* fix(mcp): preserve ownership across failed config updates (#882)
2026-09-30 22:10:35 +08:00
Saul Moro 260f1945e8 feat(agents): model aliases so one agent works in every tool (#830) (#924)
* refactor(agents): each tool reads its own tool_extras key (#830)

Qoder, Qoder CN, ZCode and OMP rendered tool_extras.claude while push
wrote their edits to their own key, so an edit never reached them.
tclaude and tcodex now read their own key over claude and codex; push
writes only the values that differ from the base and reports an edit
that removes an inherited field. Reverse parsing keys extras by tool,
so a new native agent's extras land where its renderer reads them.

* feat(agents): resolve team model aliases on pull for Claude and Codex (#830)

An agent can set model: strong, model: fast, or an alias the team defines
in models/aliases.yaml, which maps each alias per tool to that tool's own
model value and an optional effort. Pull writes the mapped model into each
tool's agent file, with effort as `effort` for the Claude family and
`model_reasoning_effort` for the Codex family. A tool the alias does not
map gets no model field; tool_extras.<tool>.model skips the alias.

Push compares deployed copies with the same resolved rendering, so a
pulled alias agent is not reported as edited. A non-string model fails
to parse, a legacy .md agent with an alias model is copied with a
warning, and an unreadable aliases file holds the agents that depend on it.

* feat(agents): resolve model aliases for every agent tool (#830)

* feat(agents): record alias resolution and redeploy on the pull fast path (#830)

* feat(agents): apply the member's local model alias override (#830)

* feat(agents): let tools switched to a model profile inherit their model (#830)

* feat(agents): keep model aliases intact on push and report drift (#830)

* feat(agents): hold alias agents on invalid alias files and drop bad entries (#830)

* feat(agents): namespaced model alias files (#830)

* feat(doctor): show how each agent's model alias resolved (#830)

* fix(agents): deliver agents held by a broken aliases file on an ordinary pull (#830)

* test(agents): stateful acceptance for model aliases and adoption docs (#830)

* fix(agents): keep push from promoting alias-owned or delivered fields (#830)

* fix(agents): deliver held agents and quiet the pull fast path (#830)

A pull that holds an agent for model resolution no longer records the team revision, so a local fix is delivered by the next pull. The fast path skips copies the member edited, re-renders untouched copies an older CLI rendered differently, and names what it wrote. Records carry the alias name, so removing an alias that gave a tool no model warns. A changed model no longer fails doctor's delivery check, and a switched tool gets no extras effort.

* refactor(agents): dedupe alias helpers (#830)

One extras base-tool table (EXTRAS_BASE_TOOL) for render and push, one model-record comparison for pull and doctor, and the per-tool renderers renderForTool no longer used are removed, their tests moved to renderForTool.

* test(agents): cover delivering held agents after the cause is fixed (#830)

* fix(agents): scope local override failures to alias agents and warn on inactive-only aliases (#830)

* fix(agents): reject aliases files without an aliases key, report fast-path holds, keep extras pins out of adoption (#830)

* docs(agents): note that an extras model pin never adopts an alias on push (#830)

* fix(pull): preview model holds in dry-run and do not report a held pull as complete (#830)

* fix(agents): hold only tools that fail resolution and keep resolved values out of alias adoption (#830)

* fix(agents): name the held tools when a broken aliases file holds only some of an agent's tools (#830)

* fix(agents): write a model edit over a tool's own extras pin back to that pin on push (#830)

* fix(agents): deliver a deleted recorded copy on the fast path, adopt an alias only for a changed model, keep unchanged own extras keys (#830)

* test(agents): cover push of a tclaude hand edit under each kind of model pin (#830)

* fix(pull): name an edited agent copy the fast path keeps for a changed model (#830)
2026-09-30 22:09:37 +08:00
Saul Moro 5fbaa377c7 feat(recall): attribute recall and adoption to the agent session in every agent (#925)
* refactor(usage): extract the locked JSONL store from the usage file (#884, ticket 01)

Move the usage file's lock, side-record, fold and rewrite protocol into
src/utils/jsonl-store.ts so the recall log can reuse it. The usage file
keeps its behavior: same waits, mode, side-file names and gitignore healing.

The store adds what the recall log needs: reads that include unfolded side
records, owner-only files by default, and a newline before an append when a
killed writer left a torn last line.

* feat(recall): attribute a Claude direct recall to the doc it opens (#884, ticket 02)

recall prints run=<id> on the region start line and appends a run record
(session from the environment, each doc's vote key, scope, printed path and
eligibility; never the query) to the active scope's recall log. A new
PostToolUse handler appends a claim for a Bash call that ran teamai recall,
and evidence for a Read under the knowledge roots. votes-sync now credits
through the recall-log reducer instead of the transcript parser and no longer
needs transcript_path.

* feat(recall): exclude the recall subagent's own reads, mark its runs with --caller (#884, ticket 03)

Claims and evidence carry the actor (session, agent_id, agent_type). A run
is marked as the teamai-recall subagent's by its hidden --caller or its
claim's agent_type; reads by that actor never count for it, while reads by
the main agent or any other subagent do. The recall agent passes
--caller teamai-recall on every recall and drops the unread
referenced-doc-ids instruction.

* feat(recall): settle each run by a valid claim or an unambiguous env session (#884, ticket 04)

A run now records its agent family, via (env or none) and whether the
environment held a single session id. The hook records a claim for every
run id a shell call printed, noting whether its command ran teamai recall
itself (parsed, so a quoted mention does not count); it also reads Codex's
string tool_response. The reducer applies only the earliest direct claim for
a run in the log, else the env session when unambiguous; unsettled runs
never vote. agentSessionIdFromEnv callers are unchanged.

* feat(recall): count a shell reader's read of a recalled doc, the Codex path (#884, ticket 05)

Add the tool-call classifier (src/utils/tool-call.ts): category, paths read
and status per PostToolUse, replacing recordToolCall's Bash and Read branches.
The shell tokenizer moves to src/utils/shell-command.ts and now returns each
simple command's operator and redirects; invokesRecall and the reader check
share it. A read is one reader command (cat, bat, batcat, less, more, head,
tail, nl, sed -n printing lines) alone or at the head of a pipeline; any ;,
&&, || or & makes the call no read. A failed read is not recorded; a read of
unknown status (Codex's string tool_response) counts only when simple.

* feat(recall): count a search that shows a recalled doc's lines, never a listing or a count (#884, ticket 06)

A search tool in content mode, or a shell grep/rg/ag/ack/git grep, records
evidence for each file under the knowledge roots whose lines its output
shows: a line that starts with the file's path and ':', under a searched
root, or the lone operand when the output is non-empty. The hook only scans
the output it already holds; it logs resolved paths, never the command or
the output. Listings (Glob, ls, find, fd, tree, rg --files, git ls-files,
grep -l/-L, a Grep file list) and counts (-c, --count, count mode) record
nothing. OMP's and Cursor's search tools produce no evidence yet.

* feat(recall): count Windows reads: PowerShell readers, drive paths, Git Bash /c/ (#884, ticket 07)

PowerShell (Claude, CodeBuddy, Copilot) is a shell tool. Get-Content, gc,
type and cat read their -Path, -LiteralPath and positional files. Paths are
resolved and compared by the new agent-path module, independent of the host:
drive-letter case, either separator and Git Bash /c/ name one file, and a
search line's drive colon is part of its path.

* fix(recall): compare Windows paths without case, look up shell verbs by own key (#884, ticket 07 follow-up)

Windows paths ignore case, so a drive-lettered path's comparison key is now
lowercased whole, as isUnderRoots did through path.win32.relative before
ticket 07. A verb such as `constructor` no longer resolves to an
Object.prototype member in the reader and search tables.

* feat(recall): count Cursor, Copilot, CodeBuddy, WorkBuddy, Qoder and ZCode reads (#884, ticket 08)

Session derivation reads Cursor's conversation_id after session_id and
sessionId. The classifier reads Cursor's tool_output JSON string and
Copilot's tool_result (text_result_for_llm, result_type), a read's path from
file_path, filePath or path, and the names view, ReadFile, bash, Shell and
run_in_terminal. Unverified payload shapes are marked in the row names.

* feat(recall): OpenCode bridge sends its session, output and status, and links task children to their parent (#884, ticket 09)

* feat(recall): Pi and OMP bridges send their session, output and status; OMP subagents their agent id and type (#884, ticket 10)

* feat(recall): credit late reads at SubagentStop and pull, and prune the recall log at pull (#884, ticket 11)

- Register SubagentStop for Claude Code, Codex, CodeBuddy and Qoder (and
  their internal builds); votes-sync runs the reducer there, without the
  end-of-turn summary.
- `teamai pull` credits sessions whose evidence is still pending, then
  prunes the log under its lock: 30 days, then 5000 lines, never dropping
  pending evidence younger than 24 h or the run, claim, credited reads and
  links it needs to vote. Evidence and its consumed lines go together.
- A run's doc is credited once, so a late read drained after the ledger's
  24 h window cannot vote twice.
- The store gains appendJsonlBatch (one lock, one write, one side record);
  a hook call and a reducer pass each make at most one write. Folds now
  work per line.

* refactor(recall): the upvote judge skips docs in the session's ledger, and the transcript parser loses its adoption logic (#884, ticket 12)

The opt-in judge (TEAMAI_UPVOTE_JUDGE) excluded the parser's adoptedDocIds on top of the session's upvote ledger. Since votes-sync credits from the recall log, a doc the parser counted but the hook path did not (a Glob listing, say) was neither credited nor judged. The judge now excludes only the ledger, and the parser keeps its recalled-doc detection, region parsing and recalled-doc-ids comment for the judge and older-format sessions.

* fix(recall): the upvote judge credits the turn before reading the ledger (#884, ticket 12 follow-up)

The dispatcher starts the detached judge before the foreground votes-sync
credits the turn, so a doc opened in that turn was sent to the judge CLI
once. The judge now runs the same reducer pass first; the ledger and the
consumed marks make a second pass a no-op.

* feat(stats): a recall section lists the last 10 sessions' runs, recalled and adopted docs (#884, ticket 13)

* fix(stats): name the agent whose hook claimed a recall run (#884, ticket 13 follow-up)

A claim now records the dispatch tool id of the hook that sent it, and the
recall section prefers it over the run's environment family. A Codex run
started from a Claude shell, or an OMP run, showed '-' before.

* docs(recall): one adoption section with the per-agent table and known limits (#884, ticket 14)

The usage guides gather adoption, the recall log, run ids, run ownership,
subagents, vote timing and the stats section under "Recall adoption and
upvotes", with a direct-recall / subagent-path table per agent and the
known limits. The core skill's troubleshooting reference says where a
recalled doc gets its upvote, and the recall subagent no longer says the
Stop hook parses its doc-id comment for adoption.

* docs(skill): CodeBuddy and WorkBuddy hooks are installed, not skipped (#884)

The core skill's hooks table said CodeBuddy and WorkBuddy hooks are skipped
by design, which contradicts the settings.json injection in src/hooks.ts and
the per-agent adoption list this PR adds to the same file.

* test(hooks): Claude settings carry the SubagentStop hook (#884, ticket 11 follow-up)

Ticket 11 registers votes-sync on SubagentStop for Claude; the hooks e2e
still expected four events. Cursor keeps four.

* fix(recall): a search line counts only from a file-name prefix, and the log keeps only .md paths (#884, review finding 1)

A one-file search's lines carry no path, so text before a colon (Cause:,
a -n false Grep line, Pi's basename) no longer becomes <doc>/<text> in the
recall log. Other lines need a prefix ending in a file name, then :<line>:,
or : in grep/rg path:text and OpenCode's header. The recorder drops any
evidence path that is not .md.

* fix(recall): a one-file search counts only when it shows a line, not a tool's no-match text (#884, review finding 2)

OpenCode's No files found and Pi's and CodeBuddy's No matches found made a
grep of the doc that found nothing credit it. The target now needs a line
that is neither blank nor a search tool's status line. Shell searches
print none, so their lines are never taken for one.

* fix(recall): the upvote judge keys each doc by the key its run recorded, not the parser's basename (#884, review finding 3)

The transcript parser names learnings/setup and a skill's SKILL.md by
their basename, so a doc the hook path had credited was judged again and
upvoted under another doc's key. recalledKeyOf maps the printed path to
the run's key for the ledger check and the upvote; a path no run printed
keeps the parser's id.

* fix(recall): SubagentStop credits locally and leaves the vote push to the next Stop (#884, review finding 4)

The main agent waits on SubagentStop, so the git round trip of the
reports push (and the ledger prune) blocked it mid-turn, up to the
handler budget, after every subagent while deltas were pending. Stop
and pull push what it credited. Stop is unchanged.

* docs(hooks): state only what is known about Codex and an unknown hook event key (#884, review finding 6)

Codex 0.159 knows SubagentStop, and Codex main's HookEventsToml has no
deny_unknown_fields. How older hook-capable builds parse an unknown key is
unverified.

* docs(skill): restore the Claude hooks row to its main wording (#884, review finding 7)

The CodeBuddy/WorkBuddy row keeps its fix, which the adoption section
below needs; the Claude row's rewording was unrelated to #884.

* refactor(recall): the adoption window and retention constants are module-private (#884, review finding 8)

Nothing imports them, tests included.

* fix(recall): an unquoted # that starts a word begins a shell comment (#884, PR review finding 1)

`cat other.md # /kb/doc.md` reads only other.md, yet the doc counted as
an operand. The comment runs to the end of the line; `a#b` and a quoted
`#` stay literal. A claim with a trailing comment still claims.

* fix(recall): an unquoted backslash escapes a space or shell character, and keeps a Windows path whole (#884, PR review finding 2)

`cat /kb/my\ doc.md` was split into two operands. Outside quotes a
backslash now escapes whitespace and ' " \ $ ` ( ) & ; | < > * ? # !;
before any other character it stays literal, so an unquoted
`C:\kb\learnings\x.md` is unchanged. Double quotes already followed
POSIX for \" and \\, and single quotes stay literal.

* fix(recall): a claim skips the root program's options before recall (#884, PR review finding 3)

`teamai -v recall "q"` is a valid recall, but its claim was marked
indirect, so a run with no session variable (OMP, via: none) never
settled. The root options now live in one table that index.ts registers
on the program and the claim parser reads through Commander's Option, so
it also skips the value of an option that takes one. `npx teamai-cli[@v]`
is handled as before.

* fix(recall): with no status, a shell call that printed only its command's errors failed (#884, PR review finding 4)

Codex sends no status, so `grep needle /kb/doc.md` answering
"grep: /kb/doc.md: Permission denied" counted as the doc's content and
upvoted it; `cat doc` answering "cat: doc: No such file or directory"
did the same. When the status is unknown, output lines that start with
the command word (as written or by name) and ": " are errors, not
content. A read or search that printed nothing else is a failure and
records no evidence. Both usage guides say so.

* fix(recall): Copilot's SessionEnd credits and pushes the session's votes (#884, PR review round 2 item 1)

Copilot's last turn can end with SessionEnd and no Stop, so votes-sync now
runs on session-end too: foreground, git only, with Stop's timeout. It
pushes as Stop does and prints no adopted summary. The Copilot rows end with
the real SessionEnd payload instead of a synthetic stop.

* fix(recall): a search with no hits prints its run id, so its shell call can claim it (#884, PR review round 2 item 2)

The no-hit line now ends in run=<id>, and the claim pattern reads it after
the logger's glyph. An OMP recall with no hits and no session variable
settles through its bash call's claim and shows in teamai stats. --check
still prints no run id and records no run.

* fix(recall): type and gc read only in PowerShell or for a Windows path, and a shell's own diagnostic is an error line (#884, PR review round 2 item 3)

In bash, type is a builtin that prints what a name is, so Codex's
type /posix/doc.md under Linux gave a false upvote with no status. type
and gc now count under a PowerShell tool, or when every file is a Windows
path. Get-Content counts everywhere. With no status, bash:, sh: and zsh:
diagnostics are error lines too.

* fix(recall): with no status, a read drops a file its command's error names (#884, PR review round 2 item 4)

In cat a.md doc.md where only doc.md errors, doc.md still counted. An error
line <verb>: <file>: … now removes that file from the read's paths, matched
as written and as resolved against the cwd; the other files still count.
2026-09-30 22:08:44 +08:00
Saul Moro 7c929a996b feat(env): team-declared secrets with member-local values (#875) (#880)
* refactor(entries): let a namespaced entry reader declare its own layout (#879)

An entry reader took its directory, file name, activation key and failure
wording from its EntryType. It can now declare them as an EntryLayout,
defaulting to entryLayout(type), which gives today's values. This lets a
later reader read env/secrets.yaml and env/<ns>/secrets.yaml activated by
resources.env. No behaviour change: env, hooks, MCP and models resolve and
report as before.

Part of #875.

* refactor(models): share the team values path and stdin reader with a second store

getTeamValuesPath takes the store directory (defaulting to models/teams)
and keeps its <team>-<hash>.json naming. The piped-stdin reader moves to
utils/prompt.ts as readStdin; the --api-key-stdin checks and messages stay
in the models command. No behaviour change.

Refs #879 (S2), #875

* feat(env): declare team secrets in env/secrets.yaml and show their state (#879)

A team repo can declare the secrets its members need, with no value, in
env/secrets.yaml and env/<ns>/secrets.yaml (key, optional description and
url). They resolve like env.yaml: active through resources.env, a namespace
entry replaces the root entry with the same key. The declarations are
absent, valid or failed; a broken file fails the secrets only, is reported
in secret wording by pull, env list and doctor, and env variables are still
delivered.

- env list and list env show each declared secret as environment or
  missing, never its value, --reveal included.
- doctor fails "Team secrets can be resolved" on a broken file, and its
  notes name env/secrets.yaml, not env/env.yaml, for an override or a key
  repeated in legacy mode (describeEntryNotes takes the reader's layout).
- push lists a changed secrets.yaml, in single-repo mode too.
- docs/designs/team-secrets.md and .zh-CN.md start here, with the #818
  boundary; usage guide, product overview, multi-project, management
  backend and the admin reference updated.

Part of #875.

* feat(env): declare a team secret with env add --secret (#879)

env add <key> [value] --secret [-d] [--url] [--role|--project] writes
env/secrets.yaml or env/<ns>/secrets.yaml with no value; a value is
rejected and never printed. env remove removes a declared secret when
env.yaml does not set the key, and --secret removes only the declaration
for a key both files carry. entryNamespaceFromFlags takes a layout so
the --role warning names secrets.yaml.

* feat(mcp): keep the entry an earlier pull wrote when a declared secret is missing (#879)

The session-start pull inherits the agent's environment, which often lacks
the member's shell export, so it removed the MCP entry the interactive pull
had written. A server whose only missing variables are declared secrets now
keeps its entry and ownership record; it is removed when it leaves mcp.yaml
or by removeAll. A failed secrets declaration keeps managed MCP state.

* feat(env): keep a member's value for a team secret and resolve it in MCP servers (#879)

teamai env set KEY (hidden prompt, --stdin, --from-env VAR) and env unset KEY
store a member's value per team repo in ~/.teamai/secrets/teams/, 0600,
accepting only keys the scope declares as secrets. ${VAR} in MCP servers
resolves a declared secret from that value, then from the member's own
environment, which leaves out values a teamai env.sh exported (Conflict 10).
A key declared as a secret and set in env.yaml resolves as the secret: its
repo value leaves env.sh, the env backup, both list renderers and doctor's
expected set (Conflict 13). A failed declaration leaves env.sh and the backup
as they are (Conflict 14). env list shows team.

* docs(env): document team secret values, storage and MCP resolution (#879)

* feat(env): set one value for a team secret for every team on the machine (#879)

env set/unset --global keep the value in ~/.teamai/secrets/machine.json.
Resolution becomes team value > machine value > the member's environment,
for MCP servers and env list (state `global`). In a scope --global still
accepts only a declared secret; outside any scope it accepts any valid key
and notes that no team declares it yet.

* docs(env): document the machine value for team secrets (#879)

* feat(mcp): keep project MCP configs with resolved values out of git (#882)

A project-scope MCP config that carries a resolved ${VAR} sat untracked
and unignored in the business repo, one `git add -A` from committing the
token. After the reconcile writes such a file and git would track it,
teamai lists its path in the clone's .git/info/exclude inside a marked
block (resolved via `git rev-parse --git-path`, so linked worktrees and
submodules work). The committed .gitignore is never touched; an ignored
path or a config with no resolved value adds nothing; dry runs write
nothing. Project-scope uninstall removes only teamai's block, and doctor
reports such a file git would still commit.

The hook sits after the appliers in reconcileMcpForConfig, outside
desiredMcpForTarget/applyJson/applyCodex, so it merges cleanly with #880.

* feat(env): name a missing team secret and the command that sets it (#879)

Interactive pull, mcp list, env list and doctor print one line per declared
secret with no value, naming the MCP servers that use it, `teamai env set KEY`
and the declared url. doctor prints it as a note and no longer fails the MCP
delivery check for a server skipped only for a missing declared secret. Pull
and doctor also note a kept entry that may hold an old value and a key
declared as a secret and set in env.yaml. The silent pull prints nothing.

The lines come from one envAdvisories() result that later pull notices extend.

* docs(env): document the missing-secret advisory in pull, doctor and the lists (#879)

* fix(uninstall): count the .git/info/exclude block in the removal plan (#882)

The plan now records whether the project's .git/info/exclude holds
teamai's MCP config block (gitExcludeBlock). It counts toward
isPlanEmpty, is listed in the summary and dry run, and gates the
removal, so a plan whose only teamai leftover is the block removes it
instead of reporting "Nothing to uninstall".

* feat(env): run a command with the directory's team env and secrets via env exec (#879)

* docs(env): document env exec, what the command can reach and the ANTHROPIC_* caveat (#879)

* fix(tests): isolate Claude config dir from model tests

(cherry picked from commit e567ac31e5)

* fix(env): warn when env add --secret updates a declaration an unknown key keeps undeclared (#879)

* feat(env): tell agents at session start which team secrets exist and to run their CLIs through env exec (#879)

* docs(env): teach agents to use team secrets through env exec and leave values to the member (#879)

* feat(env): resolve plain env variables in one order, with a member override (#879)

The environment no longer overrides an env.yaml variable in MCP servers,
env exec or env.sh. A member sets their value for a team with
`teamai env set KEY`; an interactive pull and doctor say when an export
differs and is ignored. env.sh exports a literal override and leaves out a
--from-env one; doctor's env delivery check expects the same.

* docs(env): document the variable order and the member override (#879)

* fix(mcp): keep the exclude block until every MCP config is clean (#882)

- uninstall keeps a repository's .git/info/exclude block while a config in
  it could not be parsed and still holds teamai servers, and warns
- uninstall finds and removes the block in nested repositories holding an
  MCP config, across every worktree
- the block opens at the last start marker, so an orphaned start never
  pairs with a later block's end and takes the member's lines
- doctor counts only servers the ownership manifest records, not a member's
  own server under a team name

* fix(env): discount an old env.sh export in every later command (#879)

A shell opened before a pull keeps the values env.sh exported then. Only
the process that rewrote env.sh discounted them, so the next command read
the team's old token as the member's own (Conflict 10). Each env.sh now
keeps env.sh.exports.json beside it: SHA-256 of KEY=VALUE for the last 20
values per key, mode 0600, never a value. memberEnvironment discounts any
recorded export, which replaces the in-memory snapshot.

* fix(uninstall): remove the exclude block only once its files are proven clean (#882)

A missing or unreadable managed-mcp.json made the MCP cleanup return early
without reporting anything, so uninstall removed the block while .mcp.json
still held the resolved token.

Uninstall now inspects every path the block protects after the cleanup. The
block goes only when each one is missing, or parses and holds none of the
team's servers that need a resolved ${VAR}. Anything it cannot check keeps
the block, with a warning naming the file. This replaces the leftInPlace
report from the reconcile, which the check subsumes.

* fix(mcp): protect every project MCP config holding a resolved value (#882)

Pull, doctor and uninstall each skipped a case they had not inspected and
treated it as safe. Now:

- pull lists a config in .git/info/exclude whether or not it delivered to
  it this run: a disabled or undetected tool's file, a team with automatic
  delivery off, an unreadable mcp.yaml (any teamai entry counts), a failed
  write to another tool's config, and a lost ownership manifest (the
  resolved value found in the file)
- doctor checks the same files, including one that does not parse, and
  counts a git error as a failure
- git check-ignore failing inside a repository is no longer read as "not
  tracked": the path is excluded anyway, or teamai warns with git's error
- uninstall also keeps the block while a file contains the value (8+
  characters, not a path or the login name) of a variable still set in the
  environment, which finds a server since dropped from mcp.yaml

* refactor(env): resolve a scope's env once per command (#879)

resolveTeamEnv (src/env-resolution.ts) reads env.yaml, secrets.yaml,
both value stores and the env.sh exports once. buildVarTable, the MCP
reconcile, envAdvisories, env exec, doctor and pull take that result:
an interactive pull resolves once per scope in its env stage and reuses
it for MCP and the advisories (was up to four secrets.yaml reads).

- Variable and secret values move out of resources/secrets.ts, which now
  holds declarations only (R3); one StoreResolution<V> and storeEntry
  (R2); resolveSecretDeclarations takes { active } instead of a
  tri-state array (R4); reportMissingSecrets lives in env-advisories (X1).
- ENV_KEY_RE moves to resources/env-key.ts, so env.ts imports
  SECRETS_LAYOUT statically (EV1).
- env list and `teamai list env` print one listing (env-listing.ts);
  a variable shows the value it resolves to and its source, team or
  env.yaml, like a secret (S1).
- A failed declaration is never "no secrets": the listings show no
  variable value then, --reveal included, and a server skipped for a
  missing variable is kept (S2).
- SecretState gains `unreadable` for a store that can't be read,
  handled with a never check (R1).

* fix(mcp): mcp list reports a broken secrets.yaml and exits 1 (#879)

A failed declaration read as no secrets, so every server variable came
from the raw environment, another team's token included, and mcp list
said 'all set' while pull kept MCP frozen. It now names the file, exits
1, and shows a server's variables as 'not resolved' (Conflict 14, C3).

* fix(env): env commands outside a scope say so and exit 1 (#879)

env list, add, remove, set and unset threw NotInitializedError with a
stack trace outside any teamai scope. They now print its message and
exit 1; env set/unset --global still work there.

* fix(env): env exec refuses a command without -- before it (#879)

Commander drops --, so `teamai env exec gh pr list --dry-run` read the
command's flag as teamai's and ran nothing. env exec now takes what was
typed after `exec`, and without -- before the command it says
"Put -- before the command: teamai env exec -- <command>" and exits 2.
teamai's own options may still come before --.

* feat(doctor): check that the member's secret values can be read (#879)

While teams/<team>.json or machine.json can't be read, every secret has
no value and MCP keeps what the last pull wrote, yet the MCP check
passed and only a pull warning said why. doctor now fails
'Your team secret values can be read' with the reason, for a scope whose
secrets or variables read those files.

* fix(env): name the unset variable a --from-env secret reads (#879)

The missing-secret line told the member to run `teamai env set KEY`,
which replaces the reference they chose. For a secret whose deciding
entry reads an unset variable it now says: KEY reads VAR, which is not
set. Set VAR, or run `teamai env set KEY [--global]` to replace the
reference.

* test(mcp): env unset keeps the entry, still owned, until a new value (#879)

env set, pull, env unset with nothing exported, pull: the entry keeps
the old value and the manifest still claims it, so the next value
replaces it instead of colliding with a user-owned server.

* refactor(env): expected env set/remove failures as values, one wording (#879)

- secretInput returns a tagged result; an unexpected stdin or prompt
  failure propagates instead of printing as a user error (E1).
- Every error path in env-commands uses fail() and invalidKeyMessage(),
  so each exits 1, env add's invalid key included (E2).
- removeSecret returns removed | absent | reported: env remove of a
  variable env.yaml lacks still says it is not found next to a broken
  secrets.yaml (E3).
- env add branches on --secret, not on a missing value (E4).
- env unset with declarations that can't be read no longer calls the key
  a variable (E5).
- "Secret is not declared" names the file and the next step (E6).
- Messages call the machine store "global value (every team on this
  machine)", matching --global and the env list label (R5).

* refactor(entries): EntryFailure always carries its layout (#879)

An optional layout fell back to the type's, so a failure site that forgot
it would word a broken secrets.yaml as env's. activeEntryNamespaces now
takes the layout and every failure sets it (N1).

* refactor(doctor): name entry resolution checks with their layout (#879)

The check names were looked up by the layout's message label, so a
wording change to a label would silently drop its doctor check. Each
entry set now carries its check name (DD1).

* refactor(secrets): report an unparsable value store by path only (#879)

- Drop src/utils/json-position.ts, a second JSON grammar kept only to
  print a line and column: a parse error now reports the path alone,
  which is what the spec asks (never the input), and a disagreement with
  JSON.parse can no longer leak the parser's quote (SS3).
- The unreadable-store error names what to check and the way out, and
  reads the error code without a cast (SS1).
- SecretStore is inferred from its schema (SS2).
- The design doc says an unreadable store leaves secrets 'unreadable'.

* test(secrets): typed fixtures instead of casts in the new tests (#879)

The LocalConfig and TeamaiConfig fixtures in env-exec, mcp-secrets and
secret-values are now checked against the types, so a new required
field breaks them instead of testing a shape production never has (T1).

* docs(env): env set takes variables, listings show sources, exec needs -- (#879)

- Usage guide: `env set` accepts an env.yaml variable without --global,
  not only a declared secret (D1/C2); `env exec` links the resolution
  order instead of "the order above", which the section never gave (C4).
- team-secrets.md intro states the variable order; a --from-env variable
  override leaves the key out of new shells entirely.
- Document the shared listing (variable source, `unreadable`, no values
  while the declarations fail), `mcp list`'s `not resolved`, the doctor
  check for an unreadable values file, and `env exec`'s exit 2 without --.

* fix(env): refuse env set/unset/list when the project config can't be read (#879)

Detection returned null for an unreadable project config, so env set fell
back to the user scope and stored the value for that team. Report the file
and why, exit 1, and write nothing, as pull and env exec do.

* fix(env-exec): exit 128 + signal for a command killed by SIGPIPE or SIGUSR1 (#879)

Re-raising SIGPIPE on teamai does nothing (Node ignores it) and SIGUSR1
starts the inspector, so teamai exited 0. Set the shell's exit code first
and re-raise only the signals Node ends on.

* fix(env-exec): don't send the command a second SIGINT on Ctrl-C (#879)

The terminal sends Ctrl-C and Ctrl-\ to the whole foreground process group,
so the command already has them; forwarding sent a second SIGINT, which
tools such as terraform take as force-quit. Ignore SIGINT and SIGQUIT while
the command runs and keep forwarding SIGTERM and SIGHUP.

* docs(env): exec signal handling and env set on an unreadable project config (#879)

* fix(mcp): judge MCP configs by disk and manifest, not current config (#882)

Pull, doctor and uninstall still decided "clean" from the current team
config in places. Now one function, resolvedValueEvidence, decides for all
three:

- a teamai-owned entry still in the file counts when its server has left
  mcp.yaml, as well as when it needs a resolved ${VAR} or mcp.yaml cannot
  be read (doctor no longer skips that case)
- targets include the built-in location of a tool the team dropped from
  toolPaths or moved
- doctor names a file two tools share once
- exclude updates take the existing acquireLock helper, re-read the file
  and write it atomically, so concurrent commands keep each other's paths
- uninstall inspects every worktree of each repository owning a block,
  including a nested repository's linked worktrees, and applies the
  manifest rule per worktree
- the kept-block warning names each file and why, such as the variable
  whose value matched

* fix(env): the env commands and env exec load config through --dry-run (#866)

#866 threaded { dryRun } through the loaders pull, push, status and list
use. The scope lookups this branch added did not take it, so
`teamai --dry-run env list|set|unset|add|remove` and `env exec --dry-run`
still persisted the legacy role migration, adopted a pre-#546 partition
or ran the self-mode bootstrap.

- resolveConfigForDir takes LoadOptions and passes them to both loaders.
- scopeHere/requireScope (env commands) and commandEnvironment (env exec)
  forward options.dryRun. Without the flag nothing changes: env exec still
  runs the migrations every command runs (spec #879 Conflict 12).

Five cases added to dry-run-load-path.test.ts; each failed before this
change (config.yaml rewritten, partition renamed).

* fix(secrets): name the team values file by the repo identity hash alone (#879)

Renaming team: in teamai.yaml changed <team>-<hash>.json and orphaned every
member's values. The secrets store now uses ~/.teamai/secrets/teams/<hash>.json;
the models key files keep their names. The path has not shipped, so there is
no migration.

* fix(env): keep a __proto__ env key through the store, MCP and env exec (#879)

ENV_KEY_RE accepts __proto__, but z.record dropped it from the value store,
assigning it on an ordinary object hit the inherited setter, and reading it
unset returned Object.prototype. The store, the MCP var table and the env exec
overlay are now built without a prototype (envTable), and the member's
environment and --from-env references are read as own keys (envValue).

* fix(env): env list loads config with --dry-run always (#879)

env list only reads, so like status and list since #866 it never migrates
the config it loads, with or without --dry-run. The dry-run load-path test
now runs env list without the flag, which is the case that used to migrate.

* revert: drop the #890 cherry-pick from this branch

2e15c3f6 was picked only to protect the local Claude config while this
branch's tests ran; #890 lands it on main on its own.

* fix(env): env exec applies no team env while the declarations fail (#879)

A legacy GITHUB_TOKEN in env.yaml plus a secrets.yaml that fails gave the
command the repo value, though the key may be a declared secret. Like
env.sh and the backup (Conflict 14), env exec now overlays nothing on a
failed declaration: the command gets the inherited environment, and
stderr names the failure.

* fix(fs): create the atomic-write temp file with the target mode (#879)

writeFileAtomic and writeJsonAtomic wrote the temp file with the umask's
default mode (0644 under umask 022) and narrowed it by chmod afterwards,
so a secret or model key was readable by other users until then. The temp
file is now opened exclusively with the target mode; the chmod stays for
the bits the umask removes.

* fix(env): env set --dry-run previews without asking for the value (#879)

`teamai --dry-run env set KEY` prompted for the value, or read stdin with
--stdin, before printing the preview, so it failed without a terminal. The
preview now comes first, and no value is read.

* fix(env): serialize env set and env unset on the values file (#879)

Two env set or env unset runs at once each read the store, changed it and
wrote it back, so the later write dropped the other's change. Both now go
through updateSecretStore, which takes <store>.lock (the acquireLock
helper), re-reads the file inside it and writes the result. A --dry-run
takes no lock.

* fix(env): keep a secret's stored value from becoming a variable override (#879)

Secrets and variable overrides share the team store. env set now records
kind: secret | variable from what the scope declares; variable resolution
uses only variable entries, secret resolution only secret ones, and an entry
without kind counts as a secret. env list flags an entry of the other kind
with the fix.

* fix(mcp): write an MCP config that holds a resolved value 0600 (#879)

writeJsonAtomic preserved an existing file's mode, so a 0644 .mcp.json or
~/.claude.json kept 0644 after teamai wrote a resolved secret into it. A
config holding a resolved ${VAR} value (a kept entry included) is now
written 0600; one without keeps its mode.

* fix(mcp): create the Codex config temp file 0600 (#879)

applyCodex wrote config.toml.<pid>.tmp with the umask's mode and chmodded it
after, so a resolved secret was briefly readable at 0644. Both Codex writes
now go through writeFileAtomic with mode 0600 (random temp name, created
with the mode, symlinked targets written through).

* docs(env): say env exec ignores a SIGINT or SIGQUIT sent to teamai alone (#879)

* fix(mcp): tighten an unchanged config that holds a resolved value to 0600 (#879)

* fix(env): mark what each env.sh exported so a value from one no scan finds is not the member's (#879)

* fix(env): pass on a SIGINT or SIGQUIT sent to env exec outside the terminal's foreground (#879)

* fix(env): keep marking what an env.sh exported before a rewrite dropped it (#879)

* docs(env): say a SIGINT sent to a foreground env exec alone is not passed on (#879)

* fix(secrets): name the team values file by the configured team repo URL, not teamai.yaml's repo: (#879)

A copied or hostile team repo could claim another team's repo: and receive
that team's stored values. The secrets file now hashes the URL from the
member's own config, normalized so ssh, https and credentialed forms match.
The models key store keeps its naming (#894).

* fix(env): remove what a teamai env.sh exported from env exec while the declarations fail (#879)

An invalid secrets.yaml left the inherited environment untouched, so a
shell that had sourced an env.sh still passed the repo's GITHUB_TOKEN to
the command. Every inherited value the member-environment rule discounts
is now removed, named by key on stderr; the member's own exports stay.

* fix(mcp): never write a declared secret into a project MCP config git tracks (#879)

.git/info/exclude (#882) stops git add, not a file git already tracks. For
such a file the server is skipped for that tool and an entry an earlier pull
wrote stays as it is; pull warns, mcp list shows it as withheld and doctor
fails the delivery check, each naming the file and git rm --cached.

* fix(env): leave no duplicate declaration of a key env add --secret or env remove edits (#879)

A key declared twice fails every read of secrets.yaml, and both commands
edited only the first declaration. env add --secret now updates the first
and removes the rest; env remove removes every one; both say how many.

* test: keep the real normalizeRepoUrlForCompare in utils/git mocks the secrets store reaches (#879)

* fix(env): keep the port in the URL that names a team's secrets file (#879)

normalizeRepoUrlForCompare drops explicit ports, so two team repos on one
host with different ports shared one values file and one team's secret
reached the other. The file is now named by the URL's scheme family,
lowercased host, non-default port and path; only credentials, the ssh user,
a trailing .git and slashes are dropped. The scp form and the ssh URL of a
repo still share a file; its ssh and https URLs no longer do.

The utils/git mocks that kept the real normalizeRepoUrlForCompare for the
store are no longer needed and are reverted.

* fix(mcp): skip the .git/info/exclude write while another command holds its lock (#882)

After the 2.5 s wait for the exclude file's lock, updateExclude wrote without
it, so two writers could drop each other's pattern and leave a plaintext MCP
config committable. It now writes nothing and reports 'locked': pull warns that
the file is not excluded yet and to run `teamai pull` again (doctor's exclude
check keeps reporting it meanwhile), and uninstall keeps the block and warns.

* fix(uninstall): keep an exclude entry unless its MCP config is proven free of teamai's servers (#882)

Uninstall judged a protected file clean from the current mcp.yaml, manifest
and resolvable values, so with the manifest lost, the server gone from
mcp.yaml and its value unset, a plaintext token looked like the member's own
server and the exclusion went. It now fails closed and works per entry: a
pattern goes only when its file is gone, holds no server, or holds none of
teamai's servers with managed-mcp.json still there to say what teamai wrote.
A kept entry is named with its file, why, and how to clean it by hand, since
a rerun of uninstall finds no config after a full uninstall.

* fix(env): keep http and https team repos in separate secrets files (#879)

The store identity mapped https and http to one `http` family and dropped
each default port, so `http://host/team.git` and `https://host/team.git`
read one file: if the two endpoints serve different repos, one team got the
other's stored secrets. The identity now keeps the scheme, still dropping
443 and 80; the ssh forms (`ssh://`, `git+ssh://`, `ssh+git://`, scp) stay
one family.

* fix(env): strip teamai env.sh exports in env exec when the project config can't be read (#879)

With an unreadable project config, env exec passed the inherited
environment unchanged, so a value another scope's env.sh exported (a legacy
token, a member override) reached the command while the warning said no team
values were applied. It now removes what the member-environment rule
discounts (env.sh file, record or marker, the env.sh beside the unreadable
config included), keeps the member's own exports, and names the removed keys,
never values, as the failed-declaration path does. With no config at all the
inherited environment is still passed as is.

* fix(mcp): exclude a project MCP config from git before writing a resolved value into it (#882)

Pull listed the file in .git/info/exclude only after writing the plaintext,
and a failed exclusion only warned, so the secret-bearing file stayed
eligible for git add -A. The exclusion now comes first; when it cannot be
established (exclude file or .git/info not writable, lock held past the
wait, file already tracked, git error) the file is left as it was and the
warning names the reason and the fix.

* fix(mcp): report a server withheld from a file git would commit in mcp list and doctor (#882)

* fix(secrets): keep the ssh user in a team's values-file identity (#879)

alice@host:team.git and bob@host:team.git can be different repos in each
user's home; they no longer share one values file. The scp and ssh:// forms
of one user, host, port and path still match; http(s) credentials are still
dropped.

* fix(env): overlay and remove env exec keys case-insensitively on Windows (#879)

Windows environment names are case-insensitive, so a declared api_url left an
inherited API_URL in place (Node keeps the first name of a case-folded pair)
and a removed secret survived in another case. On win32 a key now replaces or
removes every case variant; elsewhere nothing changes.

* fix(mcp): report a tracked file on a dry run and a withheld server already installed (#882)

Backports #880's merge 0d9f7fa7: a dry run (doctor, mcp list) names a
tracked file before any pull has listed it, mcp list reports withheld for a
server an earlier pull installed, and doctor's withheld note carries the
exclusion's own fix instead of the pull --force advice.

* fix(mcp): name a tracked MCP config before an unwritable .git/info/exclude (#882)

A tracked file needs `git rm --cached` whatever else is wrong, so
ensureExcludedFromGit checks gitTracks before the writability check, on a
pull and a dry run alike, and lists nothing for it.

* fix(mcp): take a project MCP config's exclude line back out once it holds no resolved value (#882)

A pull that lists a config in .git/info/exclude and then writes no value
into it (it does not parse, a member's server holds the team's name, the
write fails) removes the line it added. After a pull or `teamai mcp
remove`, a line whose configs are proven clean in every worktree, by the
proof uninstall uses (moved to mcp-reconcile.ts), is removed under the
lock; one not proven clean stays. A config listed before its write is
listed again after it, so a concurrent uninstall that dropped the line
between the check and the write does not leave the value unprotected.

* fix(secrets): tell an scp path in the ssh user's home from an ssh:// path from the root (#879)

git@host:acme/team is relative to the ssh user's home, ssh://git@host/acme/team
is absolute; on a plain ssh host they can be different repos, yet they shared
one values file. An scp path starting with neither / nor ~ is now keyed as
~/path, the path ssh://host/~/path names; host:/abs and ssh://host/abs still
match. The scp form without a user (host:path) is now read as ssh too.

* fix(mcp): judge a project MCP config by the manifest as it stood before the pull rewrote it (#882)

A pull whose manifest was lost before it ran recreates managed-mcp.json
while reconciling, so the clean-file proof read the new record and took
the exclude line out of a file still holding a teamai server that left
mcp.yaml with its variable unset. The proof now uses this worktree's
manifest as read before the reconcile.

* fix(mcp): log a rolled-back exclude line at debug level (#882)

A line this pull added and took back out, because it wrote no resolved
value into the file, was reported as removed although the member never
saw it added. Only removing a line an earlier run added stays at info.

* fix(mcp): keep a shared exclude line while another worktree's config holds a server (#882)

A pull or `teamai mcp remove` judged every linked worktree's MCP config
with today's definitions and values. Once a server's ${VAR} became a
literal, and the value was no longer set, a pull in worktree A took
worktree B's stale token-bearing entry for clean and removed the shared
/.mcp.json line, so `git add -A` in B staged the token.

These commands now release a line only when the current worktree's file
passes the full proof and every other worktree's file is missing or holds
no MCP server. `teamai uninstall` keeps its full proof in each worktree.

* test(mcp): build the other worktree from the real temp path so the test checks what it names (#882)

* fix(env): strip teamai env.sh exports from env exec with no config or an HTTP scope (#879)

Neither path applies team values, yet both passed on whatever a sourced
teamai env.sh exported, another team's credentials included. Remove every
inherited value that is not the member's own, as the unreadable-config
path does, and name the removed keys on stderr.

* fix(secrets): name a team's values file by the full SHA-256 digest (#879)

The file name is the boundary between teams, and 40 bits let a hostile
team search for a repo URL whose name matches another team's file and
read its values. Nothing has shipped, so no migration.

* fix(env): name the env.sh provenance marker by the full SHA-256 of its data home (#879)

* test(push): push publishes the secrets.yaml env add --secret leaves (#879)

* fix(uninstall): list no worktrees for a project root that no longer exists (#882)

buildRemovalPlan now lists every worktree to find teamai's exclude blocks,
outside the MCP cleanup's try. simple-git throws synchronously for a missing
directory, so a project uninstall whose root is gone crashed; main's #878 test
caught it after the merge.

* fix(secrets): keep the URL query and fragment in a team's values file name (#879)

* fix(mcp): treat an empty, unparsable or tool-less managed-mcp.json as no record (#882)

The clean-file proof took any managed-mcp.json on disk as teamai's record,
so an empty or truncated one let a pull, `mcp remove` or uninstall judge a
file still holding a stale secret-bearing entry clean and drop its exclude
line. A file now counts as recorded only when the manifest parses and holds
an entry for that tool's file. A project record teamai empties stays as []
so a file left with only the member's own servers can still be released.

* fix(env): compare exported, recorded and marked keys case-insensitively on Windows (#879)

* fix(uninstall): list no worktrees for a project root that no longer exists (#879)

* fix(mcp): keep the exclude line of a config this pull wrote when a later step fails (#882)

* fix(mcp): read a check-ignore error as unsafe unless ls-files proves the config untracked (#882)

* fix(mcp): report withheld only for targets delivery would write the server to (#882)

* docs(mcp): describe the .git/info/exclude block in the setup skill and the stricter git check (#882)

* fix(mcp): keep the exclude line of an entry a pull wrote with a resolved value after its definition turns literal (#882)

* fix(mcp): judge a nested repository's linked worktree config by its sibling's tool (#882)

* feat(mcp): record the project MCP configs a pull wrote a resolved value to in managed-mcp-files.json (#882)

* fix(mcp): keep protecting a config a pull wrote under a toolPaths mapping the team has since changed (#882)

* fix(mcp): keep a config's exclude line past the pull that rebuilt its lost record (#882)

* test(mcp): pin today's exclude rules for a missing, corrupt or locked managed-mcp-files.json (#882)

* docs(mcp): describe managed-mcp-files.json and the configs it keeps protected (#882)

* refactor(mcp): keep the #882 record edits off the lines #880 changes (#882)

* test(mcp): keep excluding a config whose entry a pull kept for a missing declared secret (#875, #882)

* fix(mcp): protect a config an older teamai wrote under a mapping an earlier teamai.yaml made (#882)

A teamai from before managed-mcp-files.json kept no record of the path it
wrote a resolved value to. Once the team changed that toolPaths mapping, no
pull visited the file. The first pull on this version now reads every
mcpProject path the team repo's history of teamai.yaml mapped, once per
worktree: a file under the project root that no current mapping or record
reaches, and that holds a resolved value, is listed in .git/info/exclude and
recorded. A git error leaves the read for the next pull; a shallow clone
reads the history it has.

* fix(mcp): keep a rebuilt record from persisting without its note of the file's other servers (#882)

When a pull rebuilt a lost managed-mcp.json and could not note the other
servers in the file (managed-mcp-files.json locked, an I/O error), it still
wrote the rebuilt record, so no later pull knew the record was rebuilt and a
stale server's line could go. The same manifest write now marks those records
unnoted: the file counts as having no record, so it keeps its line while it
holds a server, and the next pull notes them and clears the mark.

* fix(mcp): take back a managed-mcp-files.json record for a config the pull then did not write (#882)

A pull records a config before writing a resolved value to it. When the
write failed or did not happen (the file does not parse), the record stayed,
and once the mapping changed a config of the member's own at that path was
kept excluded while it held any server. The pull now takes back a record it
added for a file it did not write, as it does the file's exclude line; the
settle after records it again if the file holds a resolved value anyway.

* docs(mcp): describe the teamai.yaml history read, the unnoted rebuilt record and the record a failed write takes back (#882)

* refactor(mcp): keep the r3 edits off the lines #880 changes (#882)

* fix(mcp): also protect a config an older teamai wrote under a built-in default it has since changed (#882)

* fix(mcp): replace a symlinked Codex config instead of writing into the file it links to (#875, #882)

* fix(env): drop what a teamai env.sh exported from env exec when env.yaml or the values file fails (#875)

* fix(env): on Windows, match a declared secret to an env.yaml variable in any case (#875)

* fix(mcp): judge a project MCP config under a symlinked directory where the write lands (#882)

The appliers replace the file itself (tmp + rename) but follow its
directories. Every git check now judges realFilePath(file), the one
resolver the release keying already used: a directory linked out of any
repository no longer withholds the servers on git's "not a git
repository", and a tracked file there is named with both paths, with a
git rm --cached that works (git refuses the path through the link).

* docs(mcp): describe how a config under a symlinked directory is kept out of git (#882)

* refactor(mcp): keep realFilePath next to existingAncestor, without an import cycle (#882)

* refactor(mcp): key each worktree's targets with realFilePath, the one rule for where a write lands (#882)

* fix(mcp): judge a config under an earlier teamai.yaml mapping as a recorded file, not by today's records (#882)

* fix(mcp): have doctor check the configs earlier teamai.yaml mappings reach until a pull reads them (#882)

* docs(mcp): describe how a config under an earlier teamai.yaml mapping is judged, and doctor's check of it (#882)

* fix(mcp): record a config under an earlier teamai.yaml mapping that git tracks, and judge it once git no longer does (#882)

* fix(mcp): keep judging a recorded config for a tool the team moved while another tool still maps it (#882)

* docs(mcp): describe the tracked config an earlier mapping reached, and a moved tool's config another tool still maps (#882)

* fix(mcp): find a config an older teamai wrote under an earlier mapping another tool maps today, and judge it by that tool's records (#882)

* docs(mcp): describe the history read's configs another tool maps today (#882)

* fix(mcp): prove a shared config clean only while every tool that wrote a resolved value there has its record (#882)

* fix(env): on Windows, set and unset a key under the name the scope declares, in any case typed (#875)

* fix(mcp): judge a built-in location no mapping reaches today as an earlier-mapped file, and a shared config no pull recorded by every tool mapping it (#882)

* docs(mcp): describe the built-in location of a moved or dropped tool, and a shared config no pull recorded (#882)

* fix(env): on Windows, match a secret's declaration and stored value in any case (#875)

* fix(mcp): name the ignore rule that re-includes a config teamai just listed, instead of saying git tracks it (#882)

* fix(mcp): judge a moved tool's built-in location another tool maps for that tool too, and hold a config's line while no managed-mcp.json claims its servers (#882)

* docs(mcp): describe the no-manifest rule, a moved tool's built-in location another tool maps, and a re-including ignore rule (#882)

* fix(env): key a team's values by its repo URL when the configured remote is only an alias (#875)

* fix(env): on Windows, recognise a secret's env.yaml value under another case of its name in the environment (#875)

* fix(mcp): on Windows, keep an inherited value under another case of a team variable's name out of MCP servers (#875)

* fix(env): on Windows, replace and remove every stored entry under another case of the key (#875)

* fix(mcp): keep a tool's record as it was when its config does not parse, and take back each tool a write that did not happen recorded (#882)

* fix(mcp): mark the records a pull with no managed-mcp.json writes as unnoted until the servers no record claims are noted (#882)

* fix(mcp): have doctor judge a record marked unnoted like a missing managed-mcp.json (#882)

* fix(mcp): judge a config tools of different formats share in each of their formats before releasing its line (#882)

* fix(mcp): note the servers no record claims in the file of a tool whose record a pull writes first, as with no managed-mcp.json at all (#882)

* fix(mcp): take back a tool managed-mcp-files.json recorded before a write unless its own records hold a resolved value there (#882)

* fix(mcp): treat an installed tool's missing record as lost at every pull and in doctor, not only an empty managed-mcp.json (#882)

* fix(mcp): tighten to 0600 every project config the protection pass keeps out of git, not only the ones a pull writes (#879)

* fix(mcp): on Windows, fill a placeholder from the variable of the same name in another case (#875)

* fix(mcp): note the unclaimed servers under every format of a shared config, and pin an uninstalled tool's leftover config (#882)

* fix(mcp): count an uninstalled tool's missing record when no installed tool maps its file, and name a re-including .gitignore rule on a dry run (#882)

* fix(mcp): on Windows, match a placeholder to its secret in any case in mcp list and the missing-secret notice (#875)

* fix(mcp): scope a shared config's claims to the tools reading the same key, and settle its notes on every format's view (#882)

* fix(env): keep a file:// repo's .git suffix in its values file identity: team and team.git are two directories (#875)

* fix(env): on Windows, update and remove an env.yaml variable typed in another case (#875)

* fix(mcp): keep suspect an uninstalled tool managed-mcp-files.json lists as a writer, though an installed tool maps the file (#882)
2026-09-30 15:02:30 +08:00
Saul Moro ed95357fe8 fix(dry-run): let --dry-run reach recall maintenance and recall promote (#900) (#903)
* fix(dry-run): let --dry-run reach recall maintenance and recall promote (#900)

The root program declares --dry-run, and Commander gives it to the root
wherever it is written, so these two actions, which read only their own
options, never saw it: --prune --dry-run pruned and published,
--update-quality --dry-run wrote AI drafts, and promote --dry-run promoted and
published. Both actions now merge program.opts() as the other actions do, and
load with { dryRun }. --confidence-writeback, which ignored the flag, now
reports what it would write and publishes nothing.

* fix(dry-run): create the promotion target directory only when promoting (#900)

executePromotion created <repo>/<category>/ before its dry-run return, so
recall promote --dry-run left an empty directory in the team repo.

* refactor(promote): drop the ensureDir that writeFile already does (#900)
2026-09-29 21:17:22 +08:00
Saul Moro f8335ec40d fix(dry-run): stop dry-run and read-only commands persisting config migrations (#893) (#901)
* fix(dry-run): load mcp inject and mcp list without persisting a migration (#893)

mcp inject loaded its scope bare, so --dry-run still saved a pending role
migration, partition rename or self-mode bootstrap. It now forwards dryRun;
mcp list is read-only and loads with dryRun: true unconditionally, as status
and list do since #866.

* fix(dry-run): forward dryRun to the loader in roles init/add/remove/update (#893)

roles list is read-only and loads with dryRun: true unconditionally.

* fix(dry-run): forward dryRun to the loader in projects add/update/remove (#893)

projects list and projects members are read-only and load with dryRun: true.

* fix(dry-run): forward dryRun to the loader in tags add/remove (#893)

tags list is read-only and loads with dryRun: true.

* fix(dry-run): forward dryRun to the loader in source add/remove/add-http (#893)

source list and source browse are read-only and load with dryRun: true.
source list also reached a second bare load through loadLocalAgentConfig,
whose HTTP backfill reads only repo.kind and repo.url; it now loads with
dryRun: true, and the migration persists on the next command that writes.

* fix(dry-run): forward dryRun to the loader in remove and uninstall (#893)

* fix(dry-run): forward dryRun to the loader in packages install (#893)

* fix(dry-run): forward dryRun to the loader in import --from-iwiki/mr/claude/repo (#893)

--from-repo-list and --from-org reach the same load through importFromRepo.

* fix(dry-run): forward dryRun to the codebase loader; --lint and --status load read-only (#893)

* fix(dry-run): forward dryRun to the loader in models switch; models list loads read-only (#893)

* fix(dry-run): load read-only commands without persisting a migration (#893)

hooks list, members, exclude list, recall status and doctor load with
dryRun: true unconditionally, as status and list do since #866.

* docs(designs): name which commands take the dry-run detection path (#893)

* test(dry-run): pin that every loader on a dry-run or read-only path migrates nothing (#893)

LOAD_ONLY_COMMANDS gains the read-only commands and the previews that stay
clean on the legacy-role fixture. PREVIEWS covers the commands that fail past
the loader on that fixture: it asserts only config.yaml, with the same call
without --dry-run as the positive control that the load is on the path.

* fix(dry-run): forward dryRun through resolveMemberToolRoots for import --from-claude (#893)

scanCandidates resolves Claude's tool root before the loader import.ts already
fixed, through a bare load in resolveMemberToolRoots, so the dry run still
saved the migration. resolveConfigForDir and findUnreadableProjectConfig take
the same optional LoadOptions for the read-only callers below; hook and usage
callers pass nothing and behave as before.

* fix(dry-run): pass dryRun into loadLocalAgentConfig instead of forcing it (#893)

Forcing dryRun there made real hook runs print the [dry-run] migration
preview and stop persisting it. source list now passes it through
describeLocalAgent; every other caller is unchanged. source add-http forwards
{ dryRun: options.dryRun } like the rest of the file.

* fix(dry-run): load skill list/show, webhook list, stats and digest read-only (#893)

detectTeam takes optional LoadOptions; the Stop-hook share gate passes none.

* docs(designs): name read-only commands by example, not as every list subcommand (#893)

* test(dry-run): prove each row reached the load, and cover the remaining changed sites (#893)

Every dry-run row now asserts the loader logged its migration preview, so an
early return cannot pass. Adds source browse, codebase --lint, skill list/show,
webhook list, digest, projects update/remove, import --from-claude, and a
config-only table for members, projects members and stats.

* fix(dry-run): keep the pre-command migration from adopting a partition on a dry run (#893)

The preAction hook passes dryRun to maybeMigrate, but planMigration and
queueKeptInCheckout resolved the partition bare, so pull/push --dry-run
still renamed a pre-#546 partition. Both now take the option.

* fix(dry-run): load skill get/path read-only through the share gate (#893)

blockReason fell back to shareGate(), a bare load, when no team was passed;
only skill get, skill path and the catalog reach that fallback. The Stop
hook still asks shareGate directly and is unchanged.

* refactor(dry-run): give loadWebhookConfig the option instead of copying its branch (#893)

Also note in loadLocalAgentConfig that dryRun covers only its config.yaml load.

* revert(dry-run): leave stats out of this change (#893)

stats also loads bare per session through the dashboard scope helpers on the
pull/report path; a partial fix would claim more than it does. Tracked in the
follow-up issue with the other dry-run gaps.

* test(dry-run): cover the pre-command migration and skill get/path share (#893)

* refactor(dry-run): give shareGate the option instead of repeating it in blockReason (#893)

Also correct the test comment on when the pre-command migration runs.

* fix(dry-run): keep loadLocalAgentConfig from writing config.json under dryRun (#893)

The option reached only the config.yaml load. The legacy group-binding
cleanup, the binding-key canonicalization and the HTTP backfill still saved
config.json, so `teamai source list` could rewrite or create it. Under
dryRun they now stay in memory, and the cleanup prints a `[dry-run] Would
remove` preview instead of `Removed`.

* fix(dry-run): keep models list from saving a re-bound beta key (#893)

models list loaded the scope read-only but read the team keys through
loadTeamValues without the option, so a key a 0.26.0 beta stored under the
profile id alone was bound to its gateway and the values file rewritten.
It now binds the key in memory only; the next write command saves it.
2026-09-29 21:00:19 +08:00
Saul Moro 417e70427c fix(pull): keep a skill, rule or agent you edited instead of overwriting it (#822) (#865)
* fix(pull): keep a skill, rule or agent you edited instead of overwriting it (#822)

Pull records the sha256 of the bytes it writes at each skill, rule and
agent destination in the checkout's record (`delivered`), and the
pre-push sync records its writes too. On a full sync, a copy that no
longer matches the record is kept and named, with a warning when the
team version moved since; push repeats that warning. Tombstone cleanup
and the rules sweep keep an edited copy the same way. `--force` keeps
edits, `--dry-run` prints `Would keep`, and `doctor` does not fail on a
kept copy. Without a record (first pull on this version, a new
worktree) pull overwrites as before.

* fix(rules): drop a copy the rules reclaim removes from the delivered record (#822)

#815 removes an unedited copy of a team rule when none reaches the
directory, but the pull's delivered record kept its hash. Forget it there,
as the stale-rule sweep does, so the record matches what is on disk.
2026-09-29 18:25:23 +08:00
Saul Moro cd3a914112 fix(env): keep the user scope's env block when a project pulls (#878)
* fix(env): make doctor and uninstall act on their own scope's env block (#876)

A shell profile can carry a user-scope block and a project-scope block, but
doctor, uninstall and the active-profile resolver only ever read the first
`# [teamai:env:start]` pair. Doctor in a project then blamed the #661
backslash for the user scope's block, uninstall's plan missed a second block
of its own, and its cleanup step removed the first block whatever scope it
belonged to.

Add findEnvBlocks/findEnvBlockFor, which pick a block by the env.sh it
sources (the ownership test uninstall already used), and use them in all
three consumers. With only another scope's block present, doctor now says the
profile carries no block for this env.sh. The ownership test also accepts the
all-backslash spelling of the path, so a #661 block of this scope keeps its
diagnosis. The marker format and how pull writes the profile are unchanged.

* fix(env): keep the user scope's env block when a project pulls (#876)

Each pull replaced the first `# [teamai:env:start]` block in the shell
profile, whichever scope it belonged to, so a member with a user and a
project scope got only the last-pulled scope's env in new shells.

Injection now picks the block by the env.sh it sources. A scope replaces
its own block. Otherwise a project takes over another project's block,
never the user scope's (recognised by the env.sh the user-scope config
resolves to), and a user scope goes in right before the project block. The
user block therefore precedes the project block, so a project value wins on
a shared key, and the project block stays last-wins. Other blocks are left
alone; the marker format and the skip-when-unchanged write are unchanged.

* fix(env): address #876 review findings on env block ownership

- Drop the all-backslash spelling from candidateSpellings; the #661
  doctor tests now use a Windows-form data home instead.
- envBlockReferencesDataHome requires a path boundary before a raw
  spelling, so /home/me/.teamai/env.sh no longer claims a block that
  sources /data/home/me/.teamai/env.sh.
- resolveActiveShellProfile falls back to the first file along the
  sourcing chain that holds any teamai block when none holds this
  scope's, so a first user pull or a second project pull on the Git
  Bash chain lands next to the existing block instead of after the
  .bashrc source line.
- injectShellProfile uses one "user scope's block" predicate for both
  scopes; a block that sources no env.sh (pre-env.sh inline exports)
  is the user scope's, so a project pull no longer takes it over.
- Tests build well-formed blocks with EnvHandler.generateShellBlock.

Refs #876

* fix(env): match an env.sh path only up to its end (#876)

* fix(env): keep the user block when the user config does not parse (#876)
2026-09-29 11:09:05 +08:00
Saul Moro f19f9de23a ci(lint): enable type-aware linting with no-floating-promises and no-misused-promises (#836) (#863)
* ci(lint): enable type-aware linting with no-floating-promises and no-misused-promises

tsc does not report a promise nobody awaits or an async callback handed to an
API that ignores its result. In a CLI that is an error nobody sees. oxlint's
--type-aware mode checks both through oxlint-tsgolint.

- package.json: add oxlint-tsgolint 7.0.2003 (oxlint 1.85.0 needs
  >=7.0.2001) and pass --type-aware to `npm run lint`. The package ships
  prebuilt binaries as optional dependencies with no install script, so CI's
  `npm ci --ignore-scripts` needs no change.
- .oxlintrc.json: turn on typescript/no-floating-promises and
  typescript/no-misused-promises. Under src/__tests__/**, no-misused-promises
  skips checksVoidReturn.arguments: vi.spyOn(fse, ...) types the spy from
  fs-extra's last overload, the callback one that returns void, so every
  correct async mockImplementation is flagged.
- --type-aware also turns on oxlint's default type-aware rules. The ones
  with hits on main stay off, since each is its own evaluation under #836.
  The rest of the defaults have no hits and stay on.

Fixes, all behavior-neutral:
- dashboard.ts: the SSE debounce and the PID check become named async
  functions called with `void`; both already catch everything they throw.
  The HTTP handler becomes handleRequest, and createServer catches its
  rejection, logs it and answers 500.
- import-repo-list.ts: track in-flight imports in a Set, so the unused array
  splice returned is gone.
- scripts/mock-teamai-server.mjs: same handler wrapper as the dashboard, so a
  malformed JSON body answers 500 instead of an unhandled rejection.
- ai-client.test.ts: schedule the mock process events with queueMicrotask.

Refs #836.

* ci(lint): turn on typescript/no-useless-default-assignment

1 prod hit, type-only: syncTeamUpdatesToLocal's last parameter was
`placedRules: Record<string, string> | undefined = undefined`; it is now
`placedRules?: Record<string, string>`.

Refs #836.

* ci(lint): list only the type-aware rules that would otherwise run as off

consistent-return, no-unnecessary-boolean-literal-compare,
no-unnecessary-type-arguments, no-unnecessary-type-assertion,
no-unnecessary-type-conversion, no-unnecessary-type-parameters and
no-unsafe-type-assertion are not default rules, so listing them as off
changed nothing. The config keeps off only the nine default type-aware rules that
--type-aware would turn on.
`oxlint --type-aware --print-config` gives the same set of active rules
before and after.

Refs #836.
2026-09-29 11:07:58 +08:00
Saul Moro facc6104f6 fix: keep self-update off npm link checkouts, and report hook injection only on change (#871) (#874)
* fix(update): leave an npm link checkout alone instead of replacing it with the registry package

When the running CLI is not under node_modules (an npm link checkout),
resolveInstallPrefix returns null and doUpdate ran a prefix-less
npm install -g. That lands in npm's default global prefix, where npm link
put its symlink, so the next hook run swapped the checkout for the
published package. Skip the install and say how to update instead.

* fix(hooks): report OMP, OpenCode and Pi hook injection only when the file changes

These injectors rewrote their file and logged success on every pull, so
an up-to-date pull listed them as injected while Claude and Codex, which
already compare before writing, printed nothing. Use writeIfChanged and
log at debug level when the file is current.

* fix(update): refuse an unsupported install before prompting, and keep the skip in debug.log

Review follow-ups: resolve the install target before the prompt policy so
nobody confirms an update that is then skipped; persist the skip warning
because hooks discard stderr; word it for every null target, not only a
link; tell the user to pull before rebuilding; pass --registry in the
suggested install command.

* fix(hooks): report Hermes and OpenClaw hook injection only when something changes

Both are reached by the pull reconcile and still logged success on every
run. OpenClaw now writes its two files with writeIfChanged; Hermes also
reports whether its config.yaml entry and allowlist were written.

* fix(update): persist the vendored-install skip to debug.log too

Adversarial review follow-up: the Stop hook discards stderr, so the
vendored-layout refusal left no trace while the linked-checkout refusal
next to it is persisted.
2026-09-29 11:06:54 +08:00
Saul Moro 4284918b83 fix(tests): isolate Claude config dir from model tests (#890) 2026-09-29 09:50:47 +08:00
Saul Moro 11657bcfab fix(recall): read each agent's session variable so recall quality joins its session (#887) 2026-09-29 09:49:01 +08:00
Saul Moro f836db4230 fix(push): publish the env files env add leaves in a standalone clone (#881) (#885) 2026-09-29 09:48:00 +08:00
Saul Moro e2fd0abf11 fix(remove): remove a root MCP server named 'ambiguous' instead of refusing it silently (#862) (#864)
mcpRemovalTarget returned string | 'ambiguous' | null, so a root server
named 'ambiguous' came back as its own name and the caller read it as the
refusal sentinel: nothing removed, exit 1, no reason printed. Return a
tagged result and switch on its kind.
2026-09-28 22:31:07 +08:00
Saul Moro 53c096e02f fix(git): spawn git by its resolved path so macOS skips a PATH search per call (#869)
* fix(git): spawn git by its resolved path so macOS skips a PATH search per call

On macOS, a bare-name spawn of git costs a few milliseconds for every PATH
entry ahead of git's directory: 30-67 ms per call with a typical PATH, while
git itself takes about 8 ms. createGit now passes simple-git the absolute
path of the git that lookup would pick, resolved once per PATH value.

Refs #868

* fix(git): leave unreadable git candidates to the spawn's own lookup

A stat error other than ENOENT/ENOTDIR (a symlink loop) now keeps the bare
name instead of moving on to a later git. Narrow the CHANGELOG to the
simple-git paths and correct two comments.

Refs #868

* fix(git): keep the bare name when the first git on PATH does not start

A bare-name spawn moves past a PATH entry whose spawn fails, such as a
script whose interpreter is gone, while a spawn by path just fails.
gitBinary now spawns the candidate once (git --version) and keeps the
bare name unless it starts. POSIX-only gitBinary cases skip on win32.

* fix(git): look PATH up again when the resolved git is gone or not executable

A bare-name spawn walks PATH on every call, so a git removed or chmod'ed
mid-process is skipped. gitBinary's cache kept the stale absolute path.
One access(2) per call now drops a cached path that is no longer
executable.
2026-09-28 22:29:59 +08:00
Saul Moro 5fb316c7b8 fix(code-knowledge): keep tree-sitter grammars off V8's optimizing Wasm tier on Node 24 (#860) (#861)
On Node 24, V8's Turboshaft Wasm compiler exhausts its Zone memory while
tiering up tree-sitter grammar code and aborts the process with "Fatal
process out of memory: Zone" (nodejs/node#63421). The Swift grammar added
in #842 hits it right after its first parse, so `teamai codebase --extract`
on a repo with .swift files died with exit 133 and no fallback.

Set --wasm-tier-up-filter to an index no grammar has before the WASM
runtime starts, on Node 24 and later only. --liftoff-only has the same
effect but Node 24 ignores it when set at runtime. Node 20 and 22 keep
tier-up, which parses about 1.6x faster there.

Add Node 24 to the CI matrix so the regression test runs where it can fail.

Closes #860
Refs #842
2026-09-28 10:46:07 +08:00
Saul Moro 488c3074cf fix(learnings): publish what an older import --from-mr left in the checkout (#823) (#838)
* fix(learnings): publish what an older import --from-mr left in the checkout (#823)

Item 7. import --from-mr in 0.25.0 to 0.26.0-beta.3 wrote
learnings/<date>-<title>.md, with source_mr in its frontmatter, into the
learnings checkout and never committed it. Nothing published it. In single-repo mode it also kept
`git worktree remove` from removing the checkout an older teamai left in
.teamai/, so every pull and contribute stopped on CheckoutRefusedError.
publishQueuedLearnings now takes the sync lock first, and under it, before
listing the queue, queues every untracked file of exactly that shape
(directly under learnings/, date name, source_mr), in the active namespace
and with contribute's name, then deletes the original. It finds the one
checkout this repo registers for the branch (git worktree list), so the
shared checkout and the old .teamai/learnings-wt are both covered and
another repository's never is. A file the branch or the queue already has,
by source_mr or by content, is deleted instead, and the warning names what
has it. A dry run touches nothing.

Item 21. The branch side of that duplicate check was the checkout's own
tracked files. In single-repo mode the checkout is often the old
.teamai/learnings-wt, which nothing syncs any more, so a teammate's later
import of the same MR was missed and the remnant went out as a duplicate.
When there are remnants, the check now also fetches origin/teamai-learnings
(best effort) and reads what origin has that the checkout's commit lacks.

Item 20. pull --dry-run published the queue: publishQueuedLearnings
honoured dryRun only for the remnants. It now stops after listing the queue,
and pull prints "[dry-run] Would publish N queued learning(s)" instead of
publishing or warning.

Maintenance sweep. publishLearningsMaintenance staged all of learnings/,
so a confidence write-back or a prune swept any uncommitted file into its
commit. confidence write-back, prune and promote now return the files they
wrote or removed, and only those are staged (a removed file git never
tracked is left out, since naming it would fail the add). That exposed a
second bug:
simple-git lists a staged rename under `renamed`, not `staged`, so a
`prune --archive` with nothing else to stage counted as nothing to commit
and was never published. commitAndPushAt now counts renames.

#814 follow-ups. drainCheckoutQueue is gone: the preAction migration moves a
checkout's queue before contribute and import --from-mr. Retire-only now
says "Retired <legacy> to <backup>: this project's data already lives in
<partition>"; a linked worktree lands there too, so "Finished an
interrupted migration" was wrong for it. config.yaml.*.tmp, the temp an
interrupted config save leaves (#831), is ignored in the single-repo and
project-scope .gitignore, and the single-repo self-heal adds it.

Item 15. After a failed refresh, readableReportsWorktree called ensure
without the reports lock, so it could create the checkout while a writer
that had just taken the lock created it too. It now refreshes once more
under the lock and throws the cause if that fails as well.

Item 17. init replaced the team clone before saving the new config, so an
init that stopped in between (an unknown --role, a busy queue lock) left
the old team's config.yaml beside the new team's clone. Just before it
clones another owner's repo, init now settles the old install as the final
save would (queue set aside, indexes dropped) and moves its config.yaml to
config.yaml.previous. A failed init then leaves no config, and commands ask
for teamai init.

* fix(learnings): address review — literal pathspecs, carry settings after a failed clone

Maintenance now stages exactly the files it names: commitAndPushAt and the
removed-file ls-files lookup pass --literal-pathspecs, so a learning named
with [ or * no longer stages the stray files it matches as a pattern.

init reads the config it set aside when the rerun finds none, so an init
whose replacement clone failed no longer drops enabledAgents,
disabledAgents, toolRoots and inheritUserScope on the next run.

* fix(learnings): address review — HTTP maintenance, agent lists on a plain rerun

An HTTP install's learnings dir is no git checkout, so the removed-file
ls-files lookup threw after a prune had already deleted the file. It now
returns the same non-fatal failed publish commitAndPush gives.

init without --agent keeps the carried enabledAgents and disabledAgents,
so a rerun after a failed replacement clone no longer reactivates tools
uninstall --agent excluded.

* fix(learnings): address review — retry maintenance a busy lock or failed push kept local, keep remnants while origin is unreachable

* fix(learnings): address review — queue remnants when origin has no learnings branch, never let one bad maintenance record or remnant block the rest

* fix(learnings): address review — dedup remnants against origin's tree, keep a maintenance record a read failed on

* fix(learnings): address review — commit only the published paths, not the whole index (#823)

* fix(learnings): address review — keep a staged file across the push-retry rebase (#823)

The path-limited commit leaves a file someone else staged in the checkout,
and git refuses to rebase with anything staged, so a non-fast-forward push
failed every retry. Snapshot it with git stash create around the rebase, as
syncWorktree does, and re-apply it with --index so it stays staged.

* fix(learnings): address review — read the queue for remnant dedup under the queue lock and ownership check; move a stale config aside when init reuses a clone (#823)

* ci: re-run checks (flaky dry-run-load-path test, unrelated to this PR)

* fix(learnings): address review — keep staged files staged when the snapshot restore conflicts; never publish a hand edit as a recorded maintenance run (#823)

* fix(learnings): address review — resolve snapshot conflicts from the snapshot without a reset, keep a conflicting staged file unstaged, no hand-edit warning for a merged maintenance commit (#823)

* fix(learnings): address review — point the unpublished-edit warning at git status (#823)
2026-09-28 10:43:32 +08:00
Saul Moro 46ffa96f2c feat(init): let a member choose the git provider with --provider (#789) (#844)
* feat(init): let a member choose the git provider with --provider (#789)

A member of a team on self-hosted GitLab had to configure GITLAB_TOKEN
even when they only sync and never need the CLI to open merge requests.
`teamai init <repo> --provider <name>` now uses the named provider
instead of detecting one, and records it in the member's local config.
PR/MR creation and doctor's provider checks prefer it over the team's
teamai.yaml, which stays unchanged, so other members keep detection.
With `git`, push pushes the branch and says the MR must be opened by
hand, as it already does for a provider: git team repo.

* fix(init): address review — guard --provider gitlab and keep --provider git out of teamai.yaml

--provider gitlab on a host with no configured GitLab instance would send
the token to gitlab.com (the API base defaults there); stop with a hint to
set GITLAB_URL or use --provider git. A teamai.yaml that init creates now
records the provider detected from the URL instead of a member's git
override, matching the docs.

* fix(init): address review — do not record git as the team provider on an unconfigured GitLab

With --provider git, a teamai.yaml that init creates (empty team repo or
first self-mode init) recorded detectProvider(url), which skips the
self-hosted GitLab probe. On an unconfigured instance that wrote
`provider: git` and cost every teammate automatic merge requests. Init now
resolves the team provider as it would without the flag, including the
probe, and stops with a GITLAB_URL hint when the probe finds GitLab.

* fix(gitlab): address review — refuse a TEAMAI_GITLAB_HOST that disagrees with GITLAB_URL

Repos on TEAMAI_GITLAB_HOST were detected as GitLab while the API base,
token included, came from GITLAB_URL. Stop before any request when the
two name different hosts, and let gitlabWhoami surface the configuration
error instead of reporting a failed login.
2026-09-26 21:43:22 +08:00
Saul Moro c7e9f909e5 test(dry-run): stop git background maintenance from racing the tree snapshot (#845) 2026-09-26 21:15:45 +08:00
Saul Moro f7da1bb8b6 ci(lint): add oxlint and fail CI on any warning (#828) (#839)
* fix(tags,roles): honor --dry-run in tags subscribe, tags unsubscribe and roles set (#836)

The three commands saved the local config and reset lastPullRev even
under the global --dry-run, which is documented as "Preview mode, no
changes made". Their siblings (tags add/remove, roles init/add/remove/
update) already return early with a [dry-run] message.

The roles set preview names the additional roles it would save,
including none, because a real run replaces the existing list.

oxlint reported the unused options parameter in tagsSubscribe and
tagsUnsubscribe; rolesSet has the same bug but reads options.add.

* chore(lint): add oxlint with its default rules

Pinned to an exact version so a new default rule arrives in its own PR,
not as a CI failure on an unrelated one. The no-unused-vars options
keep oxlint's _ ignore patterns and add ignoreRestSiblings, which the
rest-omit in dashboard.ts relies on to keep config and roots out of
/api/workspaces.

* style(lint): apply oxlint safe fixes

Drop redundant escapes in regex character classes and template literals,
empty-object fallbacks in object spreads (spreading undefined adds
nothing), and anchored regexes that are plain startsWith/endsWith
checks. No behavior change.

* refactor(lint): remove unused imports and an unused catch binding

Applied with oxlint --fix-suggestions and reviewed by hand. Every
removed whole import is a library module with no import-time side
effects.

* refactor(lint): remove dead code reported by no-unused-vars

Each hit was checked against its callers and git history; none is
missing wiring (the two that were, tags subscribe/unsubscribe, are fixed
in the preceding commit). Removed: unused locals and functions, the
options parameter of tagsList, rolesList and generateDigest (read-only
commands), the never-read interactive option of importFromRepo, the
empty test/e2e.mjs left over from the E2E migration, and a try/catch
that only rethrew. new Array(n) becomes Array.from. No behavior change.

* test(lint): fix lint hits in tests

- contribute dry-run test asserted nothing; it now checks that the run
  leaves the repo/HOME tree unchanged (verified to fail when the dry-run
  early return is removed).
- Drop a no-op expect(result).not.toThrow on a string.
- Keep undefined in two optional-chain casts so a regression fails the
  assertion instead of throwing a TypeError.
- Remove unused locals, helpers and imports; new Array(n) becomes
  Array.from.

* refactor(lint): write control-character classes as \p{Cc}

no-control-regex flags literal control ranges. \p{Cc} names the same
set (C0, DEL, C1) and reads as what it means. Checked against the old
classes on every code point from U+0000 to U+10FFFF: manifest-schema and
agent-format match exactly, and contribute-check's normalization
pipeline produces the same output. The test assertion is now stricter
and checks every control character the sanitizer removes.

* refactor(lint): remove disable directives for rules that are not enabled

Four eslint-disable comments named rules this repo never ran
(no-await-in-loop, @typescript-eslint/no-explicit-any), so they
suppressed nothing.

* ci(lint): fail CI on any oxlint warning (#828)

npm run lint runs oxlint --deny-warnings and runs before the type check
in both GitHub Actions and Coding CI. The repo is at zero warnings, so
new code must stay clean. --report-unused-disable-directives also fails
on a disable comment that suppresses nothing, so a suppression cannot
outlive the code it was written for. CLAUDE.md, AGENTS.md,
CONTRIBUTING.md and the PR template list the command so contributors
and agents run it before opening a PR.

Closes #828

* chore(lint): pin oxlint 1.16.0, the newest release that accepts Node 20.0

oxlint 1.17.0 and later declare engines.node ^20.19.0 || >=22.12.0,
while the repo supports Node >=20. 1.16.0 declares >=8, supports
--deny-warnings and --report-unused-disable-directives, and reports 0
warnings on this branch.

* Revert "fix(tags,roles): honor --dry-run in tags subscribe, tags unsubscribe and roles set (#836)"

This reverts commit 224d459c99.

* refactor(lint): mark the unused options parameter of tags subscribe and unsubscribe

With the #837 dry-run fix reverted out of this PR, both functions no
longer read options. The underscore prefix keeps the signature and call
sites unchanged, so #837 can rebase onto it by renaming the parameter
back.

* chore(lint): restore oxlint 1.85.0

This reverts commit e2347efa. oxlint is a devDependency, so its Node
requirement (^20.19.0 || >=22.12.0) never reaches users installing
teamai-cli, and CI's node-version 20 resolves to the latest 20.x.
Staying on 1.85.0 keeps the #836 warning counts and the planned
type-aware follow-up on the same version.

* docs(contributing): note the Node version npm run lint needs

* docs(agents): note the Node version npm run lint needs
2026-09-26 19:56:03 +08:00
Saul Moro a47bb7eabc fix(cache,import): print cache and import review output in English (#836) (#840)
* fix(cache,import): print cache and import review output in English (#836)

CLI output must be English, but teamai import --cache-status and
--cache-gc printed Chinese headings, totals and lists, gcCache recorded
Chinese skip reasons that --cache-gc prints (and --json returns), and
the interactive import review printed a Chinese card, edit prompt and
write results.

Translate those user-facing strings and add tests that assert the
English text and no CJK characters in that output. The AI
classification prompt, the slug regex (which keeps CJK titles) and
code comments are unchanged.

* fix(cache): address review — print verbose cache debug output in English
2026-09-26 19:26:27 +08:00
Saul Moro be87a57534 fix(tags,roles): honor --dry-run and count namespaced skills in tags list (#836) (#837)
* fix(tags,roles): honor --dry-run in tags subscribe, tags unsubscribe and roles set (#836)

The three commands saved the local config and reset lastPullRev even
under the global --dry-run, which is documented as "Preview mode, no
changes made". Their siblings (tags add/remove, roles init/add/remove/
update) already return early with a [dry-run] message.

The roles set preview names the additional roles it would save,
including none, because a real run replaces the existing list.

oxlint reported the unused options parameter in tagsSubscribe and
tagsUnsubscribe; rolesSet has the same bug but reads options.add.

* fix(tags): count namespaced skills in tags list (#836)

The "N skill(s) have no tags and are always synced" hint counted the
top-level directories under skills/, so a namespace counted as one
skill and its skills were not counted at all. It also subtracted the
number of tags.yaml entries instead of checking which skills have tags.

Count the untagged skills among those pull delivers, using pull's own
resolver (resolveDesiredSkills with the member's role context). Pull
delivers every skill in the member's namespaces, all of them without
roles, whatever its tags; tags only add skills from elsewhere. So roles,
exclusions and a name held by both the root and a namespace are counted
the way pull counts them.

* fix(tags): address review — warn instead of aborting tags list on a malformed manifest (#836)

* fix(tags,roles): address review — keep the legacy role migration in memory under --dry-run (#836)

* fix(tags,roles): address review — preview the self-mode bootstrap and partition adoption under --dry-run (#836)

detectProjectConfig now takes the { dryRun } load option that roles set,
tags subscribe and tags unsubscribe already pass. Under it, a fresh
self-mode clone gets the config bootstrapSelfRepo would write, kept in
memory (previewSelfBootstrap), and a pre-#546 partition is read where it
is instead of being renamed. One test hashes HOME, the project (with .git)
and state.json around each command on three fixtures.

* fix(bootstrap): address review — make no provider auth call in the --dry-run bootstrap preview (#836)
2026-09-26 19:25:05 +08:00
Saul Moro 49675a9787 fix(push): stop reverting a teammate's update from HOME or a stale .teamai copy (#823) (#835)
* fix(push): stop reverting a teammate's update from HOME or a stale .teamai copy (#823)

Item 4, user scope: push compared HOME's rules and skills with the shared
lastPullRev only, because the per-checkout push bases of #819 were keyed for
project scope alone. After a push synced HOME's unedited copy to a teammate's
R2, the next push compared it with R1 and offered it back over the teammate's
R3. A user-scope pull now records HOME under checkoutKey(HOME) in the user
state.json, and push reads and extends it like a project checkout's. The
user-scope fast path still reads the shared fields, and an install with no
record yet keeps comparing with lastPullRev without the unrecorded-checkout
refusal: HOME is the scope's only checkout, so that revision is its own.

Item 19, inherited user scope: a project pull with inheritUserScope rewrites
HOME's skills, rules and agents under lastInheritedPullRev without moving the
user scope's push bases, so the next user-scope push offered a teammate's
newer update back the same way. That pull now adds its revision to HOME's
pushBaseRevs, creating the record from lastPullRev if there is none, and
leaves the record's rev, lastPullRev and the fast paths alone.

Item 10, single-repo: the active tree's .teamai/rules and .teamai/skills are
push sources, and on a branch behind the default branch they hold older team
versions nobody edited, which push listed as modified. The isPastVersionOf
guard that held only placed rules now covers every .teamai/rules copy, and
.teamai/skills gets the same guard: a skill is skipped with a warning when
every team file whose copy differs is an older version of it. A team file
missing locally is a teammate's addition when the member's branch never added
it, and the member's deletion otherwise; member-only files are ignored, as the
equality check already ignores them.

The checkout-base resolution that push and the agents scan each repeated
(key, record, checkoutBaseRevs, lastPullRev fallback) is now one exported
helper, resolveCheckoutBases, next to checkoutBaseRevs in pull.ts; pull uses
the same key for its record.

* fix(push): address review — record HOME's push base in an upgraded install (#823)

An upgraded user-scope install with no HOME record synced HOME from R1 to R2
on its first push but saved no base, because push recorded one only when the
bases came from a record. A teammate's R3 then made the next push compare the
R2 copy with lastPullRev R1 and offer it back. Push now creates HOME's record
from lastPullRev (userScopeRecord, shared with the inherited pull) and adds
the revision its sync reached. An unrecorded project checkout still records
nothing, since its fallback base may be another checkout's.

skill-data: contribute-member explains the stale .teamai copy warning and
how to publish an edit of such a copy.

* fix(push): address review — keep HOME's inherited base and a partial pull's delivered base (#823)

* fix(push): keep the push bases of skills a pull held, and match a replaced root rule at every base (#823)
2026-09-26 19:24:28 +08:00
Saul Moro 81aa8ea63e fix(learnings): find learnings despite a broken manifest and in the dashboard; show the MR import prompt (#823) (#834)
* fix(learnings): find learnings despite a broken manifest and in the dashboard; show the MR import prompt (#823)

Three places that decide which learnings are found, and the MR import
prompt.

Recall index rebuild (item 12). When recall had to build a missing index,
one unreadable roles.yaml or projects.yaml failed the whole build:
deliveredIndexSources and resolveActiveLearningsNamespaces threw, the error
went to log.debug, and recall said "No learnings available. Run `teamai
pull` first", which pull does not fix. Each now runs in its own try.
Learnings do not depend on the manifests and are always indexed: a broken
projects.yaml leaves only the shared root, never every namespace. Docs,
rules and skills get empty lists (undefined would index the whole trees),
and one warning names the cause; a broken projects.yaml, which both read,
gives one warning for both. The partial index is saved like any
other, so later recalls stay quiet and the next pull rebuilds it whole. A
skills collision with no index to keep skills from indexed none silently;
IndexedSkills' keep-indexed now carries the conflict line and recall shows
it when there are no indexed skills to keep (no index, or an older one).
Any other build failure is shown with its cause, not as "No learnings
available".

import --from-mr duplicate check (item 9). The scan listed only each root's
top level, so learnings under learnings/<ns>/, where #825 files MR
learnings, were never compared. Its result fed only a "marking as
superseded" warning, and nothing stored or read LearningDraft.supersedes,
so the claim was false. The scan (now findOverlappingLearnings) also walks
the active project namespaces, as the index does (safe single segments,
first root wins per relative path), and the warning becomes a
possible-duplicate notice naming the files. `import` ran the extraction as a task,
with the logger silenced, so importFromMR's warning never reached the
terminal; the extraction now runs before the tasks (item 18), where the
logger prints it. The namespaces are resolved first: a broken projects.yaml narrows the check to the
shared root with a warning, as recall does, so --dry-run and --output keep
working. LearningDraft.supersedes and SUPERSEDE_THRESHOLD are removed. The
CI extractor still reads only the root (it has no LocalConfig).

Dashboard knowledge report (item 16). Its fallback index, built when no
search index exists, passed no learningsNamespaces, so project learnings
were missing in both scopes; in user scope it read learningsRoots().read,
which keeps another repository's learnings checkout in the write root
(#808). It now uses indexableLearningsRoots in user scope, as project scope
already did, and resolves the active namespaces with the paths. A broken
projects.yaml leaves the shared root, with a warning when the report builds
its own index.

import --from-mr prompt (item 18). importFromMR, which asks "Accept
learning? [Y/n]" on a readline, ran inside the first listr2 task. In a
terminal the default renderer holds stdout back while a task runs, so only a
spinner showed and the prompt appeared after it was answered. The
extraction now runs before the task list, which starts at "Publish
learning" with the extraction's result as its context.

* fix(learnings): address review — write the partial recall index past the shrink guard (#823)

With a manifest recall cannot read, the rebuild leaves docs, rules and
skills out on purpose. Against an older-format index of a full corpus the
result is under 20% of it, so buildIndex's shrink guard kept the old file
and recall searched the entries its warning said were left out. The
degraded rebuild now passes `partial`, which skips the guard; every other
build keeps it.

* fix(learnings): address review — skip the older recall index when the partial one cannot be written (#823)

* fix(learnings): address review — show queued learnings in the dashboard's fallback index (#823)

The dashboard's temporary index now reads the contribution queue, as recall's
does; promotion and prune candidates keep to the published roots. The recall
rule tells the agent to relay the skipped-older-index warning.
2026-09-26 19:23:46 +08:00
Saul Moro b136c9c654 fix(pull): do not deliver an env, hook or MCP entry with a mistyped key (#822) (#833)
* fix(pull): do not deliver an env, hook or MCP entry with a mistyped key (#822)

Item 1. Env, hook and MCP entry schemas are plain z.object, which strips
unknown keys, so a mistyped scoping key (`role:` for `roles:`) vanished and
the entry reached every member. Each reader now reports the keys an entry
was written with that its schema does not know (known keys come from the
schema's own shape), and keepScopedEntry does not deliver such an entry and
warns once, naming the file, the entry and the key, the same path the
removed `projects:` key takes. doctor's per-entry-key check is retitled to
cover it. `env add`/`env remove` and `remove mcp` keep such a key when they
rewrite the file; `remove mcp` edits the YAML document instead of
re-serializing the parsed servers.

Item 4. recall ended every result with a Chinese line; it is English now.

Item 2 is not a bug: tags reaching a tagged skill in an inactive namespace
is the behavior #337 added and roles-tags-pull tests. The design doc's
Known gaps entry now says so.

Item 3 (pull --dry-run warnings) is left to #832.

* fix(env): warn when env add updates a variable pull does not deliver (#822)

Updating a variable that carries an unknown key keeps the key, so the
variable stays undelivered; env add now says so instead of only reporting
'Updated env variable'.

* fix(pull): keep installed MCP servers and hooks when their file has no known top-level key (#822)

A hooks or MCP file with `server:` for `servers:` parsed as empty and
removed every installed team server or hook for every member, silently.
Such a file now fails like one that does not parse, naming the keys found
and the key expected. An extra key beside a known one is still ignored.
2026-09-26 19:17:22 +08:00
Saul Moro 7c834ce428 fix(data-layout): let every self-mode worktree publish learnings and keep its queue (#808) (#814) 2026-09-26 00:29:00 +08:00
Saul Moro 87a606b727 fix(config): keep config.yaml readable while it is being saved (#823) (#831)
config.yaml was rewritten in place, so a command that read it mid-save saw
an empty file ("Invalid project config"), and a failed write left it
truncated. The three local config writers (saveLocalConfig,
saveLocalConfigForScope, the legacy role migration) now go through
writeFileAtomic: a sibling temp file renamed over the target, removed on
failure. The partition config.yaml already used it.

An existing config.yaml keeps its mode; a newly created one is 0600 (was
the umask default, usually 0644), as the partition config already is.

writeFileAtomic now writes a symlinked target at the end of its link chain
(temp file next to that file), so a symlinked config.yaml keeps its link
instead of becoming a regular file. A dangling link gets its missing target
(and directory) created, as the in-place write did; a link loop is refused
with an error and nothing is written.
2026-09-25 21:46:12 +08:00
Saul Moro e79db174c4 fix(push): stop offering a teammate's update back as a local edit (#823) (#827)
Three more ways push could list a copy the member never edited as modified,
ready to send a teammate's change back as the old version.

Single-repo mode (item 2). Push runs against a knowledge worktree whose team
root is <wt>/.teamai, a subdirectory of the git repo. The pre-push sync read
each base version with `git show <rev>:rules/x.md`, which git resolves from
the repo root, so it never found one, and every rule or skill a teammate
updated read as a local edit. The three reads now pass `./<path>`, which git
resolves from the working directory, as getFileContentWhenAdded and the agent
guard already did.

Placed agents (item 3). An agent placed with --role/--project is held when it
changed on the team since this machine's copy was current, and "current"
meant the version at the shared lastPullRev, which a pull in another checkout
moves past a copy a stale worktree still holds (the #812 revert, for agents).
The guard now reads this checkout's bases through checkoutBaseRevs, and falls
back to the shared lastPullRev for a checkout with no entry, as the pre-push
sync does. Push bases record where the sync moved rules and skills, not
agents, so the copy stays at the revision pull delivered: the guard holds an
agent that differs from its version at any base. Push records the team HEAD
as a base before the scan, and the file there is always the current one, so
the version the agent was added with is compared too whenever a base
predates it; otherwise a placement that landed after the last pull would go
back over a teammate's later edit. The hold message now says "this checkout".

Skill copy (item 5). The sync overwrote a local skill in place, so a copy
that failed partway left files from two revisions, matching no base, and the
next push listed the skill as modified. The update is now built in a hidden
sibling (the local copy, then the team version over it, so files only the
member has survive as before) and renamed into place; a failure leaves the
previous version whole. The stage carries the local modes, so cleanup makes
a read-only stage writable before removing it, and warns with the path if a
leftover cannot be removed; if the previous version cannot be renamed back,
the error names where it is.

Item 4 (user-scope push base) follows once #814 is merged.
2026-09-25 20:15:59 +08:00
Saul Moro 21cb76aa49 feat: one namespace model for every resource type (#707) (#816)
* chore: start one namespace model for every resource type (#707)

* refactor(pull): check agent and skill namespace collisions with one resolver (#707)

Add src/namespace-resolver.ts, the pure rule tickets 02-05 build on: an
active namespace item replaces the root item of the same name, and a name
twice in one place or in two active namespaces is a tagged conflict naming
both sources. The result depends only on the active order, not read order.

Agents and skills now run their duplicate checks through it and throw the
same messages. Agents still treat root + namespace as an error.

Add fast-check for the resolver's property tests.

* feat(env,hooks,mcp): scope env, hooks and MCP servers by namespace (#707)

env/<ns>/env.yaml, hooks/<ns>/hooks.yaml and mcp/<ns>/mcp.yaml are read where
<ns> is active in resources.env/hooks/mcp; a namespace entry replaces the root
entry of the same key, hook id or server name. A broken active file, a name
twice in one file or in two active namespaces stops that type for the run and
keeps what is installed. Per-entry projects: (and roles: on env) reach nobody;
roles: on hooks and MCP keeps filtering with a deprecation warning. Unknown
resources: keys warn instead of failing the manifest.

* feat(pull): let an active namespace item replace the root item for skills, agents, rules and claudemd (#707)

With a role or project configured, an item in an active namespace now
replaces the root item of the same name, whole:

- agents by stem: root + namespace is no longer a duplicate error, and a
  recorded (placed) agent replaces the root agent too
- skills by name, including a root skill received through a tag; an
  install removes the files of the version it replaces
- rules by first-level file name, in tool dirs and Hermes' SOUL.md block
- claudemd files by name in the managed block

Two active namespaces with one rule or claudemd name stop that type for
the run and keep what is installed. Push writes an edited overridden
skill, agent or rule back to its namespace, and the skills push scan
covers role and project namespaces. The placement record is withdrawn
by a same-name shared-root file only in legacy mode. Recall indexes the
skills pull delivers. doctor lists overrides, and in legacy mode repeated
names, as notes. Legacy mode is otherwise unchanged.

* fix(pull): deliver both namespace rules and claudemd files of one name (#707)

Rules and claudemd have no namespace-vs-namespace conflict: each
namespace rule keeps its own local path and each claudemd file its own
place in the block, so two active namespaces with one first-level name
are both delivered, as before. Only root suppression applies.

doctor override and legacy repeated-name notes now use the same wording
as the env, hooks and MCP ones. The usage guide and admin reference say
to keep overridable shared content at the root, with an example.

* fix(pull): stop only skills or agents on a namespace collision (#707)

Two active namespaces with one skill name or agent stem used to throw
and abort the whole scope, so rules, env, docs, cleanup and the search
index were skipped too. resolveDesiredSkills, resolveDesiredAgents,
scanRoleAwareSkills and filterAgentsByNamespaces now return a tagged
conflict. pull warns, leaves that type as installed (no install, no
inactive-namespace sweep) and syncs the rest. doctor reports the
collision as before; recall indexes no skills while it stands.

* feat(env,hooks,mcp): namespace flags, origins in status and doctor, docs (#707)

env add/remove take --role/--project; remove mcp searches every file and asks
for --role/--project when several define the name; push picks up
env/<ns>/env.yaml. status, list and doctor show where each entry comes from;
doctor lists overrides as notes and per-entry roles:/projects: as one
informational check. Usage guides and the admin skill reference describe the
namespace files; the #668 e2e moves onto them.

* test(env): show a broken env file stops env only (#707)

* fix(remove): remove an MCP server from the root file by default (#707)

remove mcp <name> follows push's convention: mcp/mcp.yaml when it defines the
name, else the one namespace file that does; --role/--project pick a namespace.
Only a name several namespace files (and not the root) define is refused.

* test(remove): expect namespace files in sorted order (#707)

* feat(docs): deliver a declared docs namespace only where it is active (#707)

A docs/<ns>/ that any role or project lists under resources.docs now
reaches only members with that namespace active; an undeclared
docs/<dir>/ stays shared. Leaving a namespace removes its local docs that
are byte-equal to the team copy and keeps edited ones with a line.
team-codebase is rejected as a docs namespace. The search index (pull,
recall, contribute) and doctor's "Team docs delivered" use the same set.

* feat(models): scope team model profiles by namespace and bind keys to their gateway (#707)

models/<ns>/models.yaml, declared under resources.models, replaces the root
profile with the same id while <ns> is active. A team API key is stored per
profile id and base_url origin, so pull never writes a key next to a gateway
on another origin; it prints the models switch line instead. Conflicts and
broken files stop model updates for the run. models list and doctor show
where each profile comes from; push validates every models file.

* docs(models): describe model profiles by namespace and key binding (#707)

* docs: list docs among the axes declared by hand (#707)

* docs: describe one namespace model for every resource type (#707)

Rewrite the multi-project design doc's precedence section for
namespace-over-root, add the per-type conflict, failure and legacy-mode
rules, and replace the per-entry key rows in the product overview. The
JSON doctor notes now also carry namespace notes.

* docs(changelog): replace per-entry scoping with namespace files (#707)

Drop the beta-only per-entry projects: entry, add the namespace axes,
the override, the per-type failure policy, the roles: deprecation on
hooks and MCP, and the upgrade-every-member-first note.

* docs(changelog): say the model key binding re-keys once and affects only betas (#707)

* refactor(namespaces): one warn-once registry instead of the quiet flag (#707)

Namespace fallback warnings, entry notices and unknown resources: keys
now go through utils/warn-once, reset once per pull, so the quiet option
threaded through eleven signatures is gone. The one-line wrappers
resolveTeamEnv, resolveTeamMcpServers and resolveTeamProfiles are removed;
every caller uses resolveEntries/resolveEntriesFor with the type's reader.

* refactor(namespaces): shared entry-file helpers, no unsafe casts in new code (#707)

- listEntryFiles/entryFileAbsolutePath replace the per-type file listers
  in env, mcp and models and the repeated path joins.
- gatewaySuffix replaces three spellings of the gateway suffix.
- LATER_RESOURCE_TYPES and friends are named for what they are:
  HAND_DECLARED_RESOURCE_TYPES, HandDeclaredNamespacesShape.
- Error messages use instanceof Error; manifest role/project ids are
  narrowed instead of cast; mapResources builds a typed object; pull
  writes env through an EnvHandler instance instead of a cast.
- status keys counts by entry type; rules localNameFor reuses
  deliversEveryNamespace, whose false answer is now documented.

* refactor(desired): move the desired-set resolvers out of pull.ts (#707)

Commands must stay thin, and recall, contribute and doctor imported
pull.js only to learn what a member receives. The skills, agents, rules
and claudemd resolvers, RolePullContext and the index sources now live in
src/resources/desired.ts; root suppression, the override note and the
repeated-name grouping live in namespace-resolver.

- A skill or agent conflict is a tagged DeliveryConflict carrying the
  resolver's NamespaceConflict, rendered once by describeDeliveryConflict
  (wording unchanged); DesiredItems names the result union.
- DesiredItems keeps each override, so doctor no longer rebuilds skill
  and agent overrides by hand.
- recall and contribute share deliveredIndexSources; pull indexes through
  the same indexedSkills instead of a second copy.
- doctor: one unresolvableCheck for skills, agents and docs; the docs
  check reports an unreadable manifest instead of returning nothing; the
  namespace notes catch only the team-repo reads.
- docs withdrawal reuses utils pruneEmptyDirs.

* fix(pull): name both files in a skill or agent conflict (#707)

Story 8 asks for a message naming both files. A skill or agent conflict
named only the namespaces, and an agent defined twice inside one
namespace (agents/a/x.md next to agents/a/x.yaml) read as 'found in
active namespaces "a" and "a"'. The duplicate case now names its one
place, and both cases list the two files.

* fix(rules): only a delivered namespace rule replaces the root rule, in every tool (#707)

- The rules override ran before the tag filter, so a namespace rule the
  member's tag subscription excludes still suppressed the root rule and
  the member received neither. The tag filter now runs first.
- JoyCode, OMP, Pi and Copilot share their rule directory with the
  member's own rules, so the stale sweep deletes nothing there unless a
  tombstone names it: the root rule a namespace rule replaced stayed
  installed and both versions loaded (story 6). pullAllRules now removes
  such a copy while it is byte-equal to its render, as agents do.

* fix(recall): keep the indexed skills while a skills conflict holds them (#707)

On a skills conflict pull keeps the installed skills, but the index was
rebuilt with none, so recall returned none of the skills the member still
has. The index now keeps the skills entries it already held, and pull
does the same when resolving the skills fails.

* fix(hooks,mcp): fail hooks inject and mcp inject when team entries do not resolve (#707)

reconcileTeamHooksForConfig returned [] when the team hooks could not be
resolved, the same value as a team without hooks, so hooks inject printed
'Hooks injected into all AI tool settings' and exited 0 over a broken
hooks/<ns>/hooks.yaml. It now returns { ok: false }, and hooks inject
exits 1 after the warning that names the file. mcp inject said 'Already
up to date.' in the same case; the MCP reconcile now marks the result
unresolved and mcp inject exits 1.

* fix(models): bind a beta API key to the gateway it was sent to, once (#707)

- A key a 0.26.0 beta stored under team:<id> counted for whatever origin
  the root profile had now, so a root profile moved to another host got
  the old key written next to it. The first pull or models command that
  reads such a key now binds it to the origin TeamAI last wrote into the
  agents switched to that profile (the root's current origin when it is
  among them), else to the root's current origin, and never re-reads the
  unbound key. Pull then leaves a moved agent alone and asks for the
  switch.
- The 'switch to set a key' and 'no longer active' lines are written to
  debug.log too: SessionStart pulls run silent.
- Legacy mode, which reads no namespace, says a profile 'was removed'.
- A models command whose profiles do not resolve reports it with
  log.error and exit code 1, as env list and mcp list do, instead of
  throwing.

* fix(entries): keep 0.25 files that repeat a name under different roles: working (#707)

0.25.0 let hooks.yaml and mcp.yaml repeat a hook or server name under
different roles:, delivering every copy that passed the role filter (MCP
kept the last). The namespace resolver treated that as a duplicate, so a
member holding both roles, or a role-less member in a team with
projects.yaml, stopped receiving hooks or MCP entirely. During the
roles: deprecation window such a repeat is delivered as in 0.25; a name
repeated without roles: on every copy is still a duplicate.

Also restores the test that the role filter runs before
requireTeamScripts, so the transparency print lists only what will run.

* feat(doctor): say where each entry type's entries come from (#707)

The spec asks doctor, like status and the list commands, to show each
entry's namespace; doctor listed overrides only. For env, hooks, MCP and
models, a namespace contributing any entry now adds a note counting the
entries by origin, 'env: 3 received here (2 root, 1 checkout)', from the
describeOrigins that status uses.

* test(pull): env and hooks conflicts between two namespaces, and builtin: in a namespace file (#707)

Seam 1 asks every type to show, through pull, that two active namespaces
defining one name keep the installed state and name both files. Env and
hooks were covered only through doctor and the handler; so was the
warning for builtin: in a namespace hooks file.

* docs(skill-data): hooks and MCP edits are published with git, not teamai push (#707)

manage-admin.md told admins to publish hooks/MCP file edits with
teamai push, which sweeps only rules/, env/ and .codebuddy-plugin/, so
an agent following it would push nothing. It now says to commit and
push the file with git, as the usage guide does.

* refactor(models): read switched agents without a cast (#707)

* docs(changelog): doctor counts entries per namespace; inject fails on unresolved entries (#707)

* refactor(pull): drop imports the resolver move left unused (#707)

* fix(hooks): install the built-in hooks when the team hooks do not resolve (#707)

A first init or bootstrap whose team hooks did not resolve (a broken
namespace file, a clash, a duplicate id) installed no built-in hook, so
the session-start pull that heals the member never ran. Installed team
hooks are still kept; the built-in hooks are now installed where missing,
with the root file's builtin: overrides whenever hooks/hooks.yaml parses,
and with their defaults only in a tool with no teamai hook when it does
not. init and bootstrap say that the team hooks were not installed.

* fix(manifest): keep an unknown resources: key when roles and projects save (#707)

zod stripped the key this CLI only warns about, so a projects or roles
command run on this version deleted a newer CLI's type from the team
repo for everyone.

* fix(docs): withdraw a copy the team edited after delivery, not only an unchanged one (#707)

Withdrawing an inactive docs namespace compared the local copy with the
current team file only. A doc the team changed after the member received
it was then kept forever with a false 'you edited it' line. A copy equal
to an earlier team commit is what the mirror delivered, so it goes too.

* fix(rules): withdraw a replaced root rule edited in the same push, name a kept copy (#707)

In the JoyCode, OMP, Pi and Copilot rule dirs, a replaced root rule's
copy was removed only while it matched the current root rule. When the
admin edited the root rule and added its namespace override in one push,
the member's unedited copy stayed loaded beside the override, silently.
It is now also compared with the render at the last pull, and a copy
that is kept is named with the fix.

* fix(doctor): fail a check when team hooks or model profiles do not resolve (#707)

teamai status counts such a type as 0 and says to run doctor, but doctor
had failing checks only for env and MCP, so a duplicate hook id or a
two-namespace clash showed nothing there.

* fix(env): warn when --role names a namespace nothing declares (#707)

env add/remove --role <ns> wrote env/<ns>/env.yaml for a namespace no
role or project lists under resources.env, so the variable reached
nobody and nothing said so. The same applies to remove mcp --role.

* fix(env): find a changed namespace env file whose name is not ASCII on push (#707)

git ls-files quotes such a path by default, so it never matched the name
on disk and push skipped the change.

* fix(entries): an active env, hooks, MCP or models file that cannot be read stops the type (#707)

The readers folded every read error into 'file does not exist', so an
unreadable namespace file silently delivered the root entry in place of
its override. Only ENOENT is absence now; any other error is a broken
file, like one that does not parse.

* fix(entries): match env, hooks, MCP and models namespace dirs case-folded, as docs does (#707)

A declared namespace was joined onto the path as written, so with
env: [checkout] and a directory env/Checkout/, macOS and Windows members
got the override and Linux members the root value.

* fix(doctor): split the legacy claudemd paths on '/', not path.sep (#707)

listFilesRecursive always joins with '/', so on Windows every path was
one segment and a claudemd/<ns>/x.md beside claudemd/x.md was never
reported.

* fix(recall): index the rules pull delivers, not the whole rules/ tree (#707)

A namespace rule replaces the root rule of its name, but recall, contribute
and pull indexed every file under rules/: the replaced root rule and the rules
of inactive namespaces came back from recall. Index the resolved rule set, as
docs and skills already do.

* fix(entries): write a namespace file into the directory pull reads it from (#707)

Pull matches a declared namespace to its directory case-folded, but --role and
--project returned the spelling typed. On a case-sensitive filesystem
`env add --project checkout` created env/checkout/, which shadowed
env/Checkout/ and dropped its variables from delivery.

* fix(remove): remove no MCP server by a bare name while an MCP file does not parse (#707)

The team scan skips a file that does not parse. With mcp/mcp.yaml broken,
`remove mcp db` took the one readable checkout/db as the target and removed
it. Refuse and name the file unless the readable root defines the name.

* docs: rules in recall, namespace writes and remove mcp on a broken file (#707)

* fix(remove): say the MCP file --role or --project names does not parse, not that the name is missing (#707)

The team scan skips a file that does not parse, so `remove mcp db --project
checkout` with a broken mcp/checkout/mcp.yaml reported "Not found". Name the
file and remove nothing.

* fix(skills): remove a leftover of another skill version only when it matches that version (#707)

Install removed any installed file at a path another team version of the skill
has, by path alone. A file a member added under that name, e.g. README.md
beside a namespace they never had, was deleted on every pull. Remove it only
when it is byte for byte that version's file; keep any other and name it.

* fix(env): edit no env file that does not parse, and no --project target after a failed refresh (#707)

env add and env remove read the target through parseEnvYaml, which answers an
empty list for a file that does not parse, then wrote that back: every
variable the file had was replaced. They now refuse and name the file.

--project resolves through manifest/projects.yaml. After a failed pull that
copy may be stale and name a namespace the project no longer uses, whose file
push would publish, so --project now changes nothing then. The root file and
--role do not depend on the manifest and still only warn.

* fix(pull): let no unusable namespace item replace the root one (#707)

A skill directory without SKILL.md replaced the root skill of its name:
install overlaid it and removed the installed SKILL.md as the other version's
leftover, while pull still counted the skill as synced. Such a directory is
no longer a skill; pull names it and keeps delivering the root one.

An agent file that does not parse delivers nothing, yet it still replaced the
root agent, and cleanup removed the unchanged root copy because no active
destination held that stem. The root agent now stays while its replacement
cannot be read or parsed.

* fix(push): take no namespace directory without SKILL.md for a member's skill (#707)

Pull stopped delivering such a directory in 6c7df8e3, but the push scan still
mapped the skill name to it. An unedited root skill then showed as modified,
and push wrote it into skills/<ns>/<name>/, deleting what was there and making
it a namespace skill that replaces the root one.

Also keep only a replaced root agent while its replacement does not parse, so
an unchanged copy from an inactive namespace is still removed, and put
renderedForTool's doc comment back on it.
2026-09-25 17:59:19 +08:00
Saul Moro f558b94614 fix(push): keep a teammate's update when pushing from a stale worktree (#812) (#819)
Before scanning, push syncs each rule and skill the member never edited to
the team repo's version, and "never edited" meant equal to the version at the
project's shared lastPullRev. state.json is shared by every worktree, so a
pull in another checkout moved that revision past the copy a stale worktree
still held: the unedited copy read as an edit, and push offered it as
modified, ready to send the teammate's change back as the old version.

Push now compares with the revision this checkout last synced, from its
lastPullByWorkspace entry (checkoutKey is exported from pull.ts), and falls
back to the shared lastPullRev for a checkout with no entry. For that entry to
survive, a pull at a new team revision no longer drops the other checkouts'
records: a checkout recorded at an older revision already misses the fast
path. When the pull finds lastPullRev cleared, it resets the other records
to an empty rev (FORCED_FULL_SYNC_REV), which matches no revision, so a forced
full sync reaches every checkout, single-repo mode included, while each record
keeps its push bases. Each full sync keeps only the records
of checkouts `git worktree list` still reports, and keeps them all when the
list comes back empty, so a removed or re-created worktree's entry does not
pile up.

The sync itself moves the unedited copies to the team repo's revision, so
push then adds that revision to the entry's pushBaseRevs (newest first, the
20 newest kept) and the next push compares with them; otherwise a copy synced
to R2 read as an edit against R1 once a teammate published R3. Push leaves the
entry's rev alone, since the pull fast path reads it and the checkout still
lacks that revision's docs and agents; the next pull rewrites the entry
without pushBaseRevs. The sync accepts a copy at any of pushBaseRevs or rev (a
skill only when all its files are at one of them), so a copy it left alone as
edited is synced again once the member undoes the edit, back to whichever
version a sync gave it. The base is recorded even when the sync stops
partway, which now warns, since the copies it did not reach still match an
older base; a revision push cannot save stops the push before the scan. A
checkout with no entry (a new worktree, or one last pulled by an older CLI)
can only sync against the shared lastPullRev, which may be another
checkout's or cleared: when the scan lists a team rule or skill as modified,
push stops before creating a branch and asks for a pull there, warning that
the pull replaces those files. A rule this machine placed does not count, and
config-only pushes and new resources go through.
2026-09-25 12:04:16 +08:00
Saul Moro c7723d652b fix(hooks): keep a removed worktree's hook events in its project (#810) (#824)
A hook resolves its scope from the payload's cwd, and resolveConfigForDir
answers the user scope for a directory that no longer exists. So once a
session's worktree was removed, its remaining events (tool_use, SessionEnd,
Stop) and skill uses were recorded under the user scope, which then counted
the session and reported the skills to its team, or were dropped when there
was no user scope. The project lost the session's last snapshot.

resolveHookConfig (dashboard-collector.ts) is the one resolver for the
hook dispatcher and the legacy dashboard-report, track and track-slash
entry points. For an existing (or absent) cwd it is resolveConfigForDir, as
before, and reads nothing else. For a cwd that is gone it reads this
session's last event that recorded a dataHomeKey, once per process, preferring
the events recorded at that same cwd (a detached Stop can run after the
session moved on to another repo), and
resolves the config at that event's projectAnchor, the main checkout,
which still exists (for a bare repo, whose anchor is the git directory,
at one of its worktrees that still exists). It uses that config only when it is still the scope the
recorded dataHomeKey names, so a worktree's own legacy .teamai never becomes
the main checkout's scope. If that config exists but cannot be read, the
event is dropped rather than given to the user scope (#748). With nothing
to match (no events, events from before #809 without an anchor), it is
today's answer.

The dispatcher's track and track-slash handlers now use the dispatcher's
config instead of resolving their own, and eventProjectAnchor gives an
event whose cwd is gone the session's last anchor, as process_exit does.
The legacy track-slash looks skills up under the resolved scope's tool
roots before the cwd's. The share reminder's gates (contribute-check on
Stop, pending-hint on the next prompt) ask about hookScopeDir, the
directory resolveHookConfig resolves from, so a removed worktree's
session gets the project's reminder settings, not the user scope's. The
legacy `teamai contribute-check` command gates the same way. The hook
session id has one implementation, deriveDispatchSessionId in
utils/session-id.ts, shared by the dispatcher and the event writers.
2026-09-25 11:51:40 +08:00
Saul Moro 4a65e3f676 fix(import): publish the learning import --from-mr extracts (#823) (#825)
`import --from-mr` wrote its learning into the teamai-learnings worktree,
then pushed with autoPushViaMR, which commits `.` in repo.localPath: the
knowledge clone, another checkout. That found nothing to commit, so the
learning stayed untracked on this machine and never reached the team,
while the command still reported the push step as done.

The draft now goes into the contribution queue, and a "Publish learning"
step calls publishQueuedLearnings, the path `teamai contribute` uses: it
commits and pushes the queue on teamai-learnings and drops an entry once
it is on origin. When publishing fails the learning stays queued, the step
says so, and the next `teamai pull` publishes it. As in contribute, the
queued file takes contribute's name (a random suffix keeps two learnings
with the same title and day apart), the recall index is rebuilt after the
publish attempt, the supersede check also reads the queue, and a read-only
(HTTP) source is refused up front instead of queueing a learning nothing
can publish; --dry-run and --output still work there. "Push changes via
MR" is left for the teamwiki update it was also for.

The learning also lands where contribute puts it: resolveLearningsSubdir
(now exported) picks learnings/<namespace>/ when exactly one active project
declares a learnings namespace, else the shared root. It used to go to the
root, where every project's members recall it.

Also, from the same follow-up issue:
- wiki slug: the main checkout's root takes its repo's name too, so one
  opened through a differently named symlink writes the same evidence as
  its worktrees. Subdirectories keep their own name.
- repo labels: a path is not qualified into a label a remote-form key
  already has (github.com/acme/api vs /x/acme/api), so the two no longer
  merge in `stats --by-repo`. The fallback is the repo's directory, so a
  bare repo's keys still share one row.
- local-agent tests use a session id unique per run: the hint markers are
  machine-wide files in os.tmpdir() keyed by session id, and overlapping
  runs deleted each other's.
2026-09-25 11:50:20 +08:00
Saul Moro c73d22147d fix(report): each scope reports its own dashboard sessions once, against its own snapshots (#785, #786) (#791)
* fix(report): each scope reports only the dashboard sessions recorded in it (#785)

Every scope read one machine-wide events.jsonl and picked its sessions out
by cwd prefix. The user scope excluded nothing, so a user-scope pull
reported every project's sessions (and, through the shared reported
snapshots, took them from the project's own report); Copilot sends no cwd,
so a project never reported its Copilot sessions; and a raw cwd under a
symlink or /tmp never matched the realpath'd projectRoot.

The hook now stamps each event's dataHome with the data home of the scope
the dispatcher resolved (the key the per-scope usage file already uses), and a
report keeps only its own scope's events, comparing realpath'd keys. A
project also owns its in-repo .teamai key, where hooks record until
migration moves it to a partition. Events written before this carry no
dataHome: a project keeps those whose realpath'd cwd is under its root, the
user scope never reports them. The log stays machine-wide for the
dashboard UI, stats --by-repo, session save and the contribute check.

Removes the excludeProjectRoots option, which pull only ever passed as []
(the user target exists only when no project config resolved), and the
projectRoot option now carried by selfConfig. The usage guide documents how
to remove by hand a skill an earlier release pushed into stats/<user>.yaml
from another project.

* fix(report): address pre-review findings (#785)

- Events record `dataHomeKey`, a hash of the realpath'd data home, instead
  of the path. A Copilot event persisted a workspace path through its data
  home (the raw root for a non-git project, the path-derived partition name
  otherwise), breaking the path-free Copilot contract from #666.
- A data home that no longer exists (an in-repo .teamai removed after
  migration) keys through its parent's realpath, so it still matches the key
  recorded while it existed.
- A non-git project's root is realpath'd before older events' cwd is matched
  against it, as the cwd already was.
- A key that is not a string (a hand-edited log) counts as absent instead of
  throwing and skipping the whole report.
- The legacy `dashboard-report` command's stamping is asserted.
- CHANGELOG and the comment say teamai does not record Copilot's cwd, not that
  Copilot sends none.

* docs(report): place the stats cleanup under usage reporting (#785)

The manual `stats/<user>.yaml` cleanup sat under single-repo mode, but the
pre-#748 leak hit every team with a git-kind repo, so it moves to "Usage
reporting" and notes where an `http` team repo keeps the file. The guide
also says the scope key is per event: hooks that run outside the project
(a worktree removed before the session ends) report to the scope they ran in.

* docs(report): name where unattributed sessions go (#785)

The CHANGELOG now says a session in a directory that resolves to no project
(a non-git project's subdirectory, a submodule or nested clone) is the user
scope's, as for skill usage. The usage guide drops the line on http team
repos: pull does not report usage to them, so no stats file there needs
cleaning.

* fix(report): each scope keeps its own reported dashboard snapshots (#786)

The report sends per-session deltas against reported-*.json snapshots that
every scope shared. A session whose events belong to two scopes (a cd into
another project mid-session) was then reported by the first scope, and the
second compared its own part with the first scope's totals and sent nothing.

Each scope now keeps its snapshots in <dataHome>/dashboard/, and the user
scope, whose data home holds the shared files, in user-reported-*.json. The
first time a scope needs one it copies the shared file, so the first report
after the upgrade sends nothing already reported; after that it reads only
its own. The user scope moves too, unlike the ticket proposed: had it kept
writing the shared file, a project seeding later would copy the user scope's
part of a split session and report nothing for its own. The shared file is
no longer written, except by an earlier release after a rollback, which only
a scope not yet seeded reads.

* fix(report): report each dashboard session once, from the scope it started in (#785, #786)

A Stop carries the whole transcript's totals (prompts, tokens,
interventions, request cost). Filtered per event, a session that moved into
another scope mid-session was reported whole again by the scope holding the
later Stop: 3 user-scope prompts then 2 in P reported 3 to the user team and
5 to P. Each session is now decided once, by its first keyed event, and
reported whole by that scope. This replaces #786's "a split session reaches
both teams with its part"; per-scope snapshots stay, so a session ID another
scope already reported (Copilot's PID fallback) still counts as new.

Unkeyed sessions from before the upgrade are decided by their first cwd. The
user scope now takes those whose directory still exists and resolves to it
(resolveConfigForDir, the dispatcher's rule) instead of dropping its whole
backlog; no cwd, or one removed since, is still no scope's.

The Copilot test also runs a payload without cwd from a hook in the project.

* fix(stats): read the scope's own dashboard filter and snapshots (#785, #786)

`teamai stats` (#771) still called filterEventsByScope with the old
{ projectRoot, excludeProjectRoots } options, synchronously, after #795
made it async and keyed by the scope config, so main no longer type-checks
and stats-scope fails. It also subtracted the shared reported-*.json, which
no scope writes since #786.

stats now filters with the config it resolved and subtracts that scope's
own snapshots (readReportedInterventions / readReportedPromptTokens, the
report's readers), so what it shows matches what pull reports. The user
scope leaves a project's older sessions out, as the report does (#785); the
stats-scope case that pinned "no exclusion in the user scope" now expects that.

* fix(stats): address CI review (#785)

A session ID now names one run up to its session_end or process_exit.
A PID-fallback ID (Copilot) comes back for a later run, maybe in another
scope, and the log keeps the ended run below the compaction threshold, so
grouping by ID alone gave the later run to the first run's scope. Each run
is still decided whole by its first keyed event.

Events written by main since #795 record the data home as a path
(`dataHome`); the report now keys them the way the writer derives
`dataHomeKey`, so pending Copilot sessions (no cwd) are not dropped.

* fix(stats): address CI review (#785)

A later run of a reused session ID (Copilot's PID fallback) was decided
on its own but returned under the same ID, so aggregation and the
per-scope snapshots merged two runs in one scope back into one session.
The filter now returns a later run as `<id>@<first event timestamp>`;
the first run keeps the bare ID, so existing snapshots still match.

An unkeyed event's cwd under a project root counted even when the
directory was gone (realpath fell back to the raw path). It now counts
only while it exists, as the docs and the user-scope rule already say.

* fix(stats): address CI review (#785)

Run identity no longer depends on which earlier runs compaction kept:
every run is `<id>@<first event timestamp>`, so a reused PID-fallback ID
is a new session even when the scope's snapshot still names the run
compaction dropped. Snapshot entries keyed by the bare ID (written by
earlier builds) are adopted by the first run of that ID in the log, so
the upgrade re-sends nothing; the next snapshot holds only run IDs.

An unkeyed event's cwd is now owned by the scope resolveConfigForDir
resolves it to, for projects as for the user scope, so a nested clone
under a project is no longer reported by both. The lexical root matcher
and its string-level tests go; the cases move to real repositories.

* fix(stats): address CI review (#785)

adoptBareKeys() read a legacy bare `pid-N` snapshot entry as the first
run's, but only in memory: the success writes merge into the file, and
with nothing new to report nothing was written, so the bare entry stayed.
Once compaction dropped that run, the next run reusing `pid-N` read it and
was suppressed. The report now writes each snapshot as soon as a bare entry
is retired, under the run ID only, even when there is no delta.

* fix(stats): address CI review (#785)

A bare snapshot entry is given only to a run an earlier release recorded
(its first event has no dataHomeKey). Only earlier releases wrote bare
entries, and a seeded one may be another scope's run under a reused
PID-fallback ID, so a run this release recorded takes none. A marker of
the seed time would miss the common case: a scope seeds at its first
report, usually the pull its first session's SessionStart triggers.

A second end of a run with nothing recorded since the first (the
dashboard monitor's process_exit after SessionEnd) joins the run it
closed instead of opening a terminal-only run counted as a session.

* fix(stats): address CI review (#785)

A scope's first snapshot is seeded only with the shared entries of its
own runs in the log, under their run IDs, and none for a run recorded
with a dataHome path: that release already kept per-scope snapshots, so
a shared entry under the same ID is another scope's. An unmatched entry
is dropped instead of copied, so a later reuse of the ID cannot inherit
it.

The dashboard monitor records processExitAfter, the last event it
observed, and the scope filter closes only that run. A delayed exit
appended after the next run of the same ID began no longer ends it and
splits it in two; an exit whose run compaction dropped is ignored.

* fix(stats): address CI review (#785)

An earlier release summed every run of a reused ID under its bare
snapshot entry, but only the first retained run adopted it, so the next
one was reported again. Each of those runs in the log but the last is
now taken as reported at its own totals and the last takes the entry,
in the report, in teamai stats and in the seed from the shared file.
The last run is undercounted by at most the other runs' share, once.

A session_start on a fallback ID from another monitorPid than its open
run's begins a new run, so a run that crashed with no dashboard running
no longer takes the next invocation, maybe another scope's. A tool's own
ID is not split: Claude fires SessionStart again on resume, in a new
process, and its Stop carries the whole transcript.

* fix(stats): address CI review (#785)

An end splits runs only on a fallback ID (pid-…). A tool's own session
ID is one session whatever ends it records: claude --resume continues it
in a new process, and its Stop carries the whole transcript, so a second
run counted it again, maybe in another scope.

* fix(stats): address CI review (#785)

A tool's own session ID is keyed by the ID itself again, as on main,
not by its first event's timestamp, so a session resumed after
compaction dropped its events still reads what its scope reported.
Only PID-fallback runs carry the timestamp.

A bare fallback entry is the sum of the runs of its ID in the log at
the earlier release's last report, and compaction keeps or drops an
ID's runs together. Those runs now consume it in log order, each up to
its own totals, so a later run that release never reported is sent
instead of taking the whole entry. The prompt-token snapshot decides
which runs it covered; interventions and daily follow it, and the first
run always takes a share.

* fix(stats): address CI review (#785)

Seeding a scope from the shared snapshot splits the whole log into runs,
lets every scope's runs of a bare ID consume its entry in log order, and
keeps the shares of the scope's own runs. The shared file summed every
scope's runs, so one scope consuming it alone could spend another
scope's baseline and suppress its own pending run.

The scope that first reports a tool's own session ID records itself in
~/.teamai/dashboard/session-owners.jsonl (the ID and its data home key,
no path), and a recorded session stays that scope's wherever it is
resumed, after compaction dropped its events too.

A dashboard started before processExitAfter existed reads the log and
appends its exit in one pass, so an unannotated exit less than one PID
check after the open fallback run began belongs to the run closed before
it instead of closing the next invocation.

* fix(stats): address CI review (#785)

A run taking its share of an earlier release's summed daily snapshot
keeps its own success and correction flags: the sum's are no single
run's (a successful run and an interrupted one sum to unsuccessful), so
an adopted run changed sessionsSucceeded without sessionsEnded.

An unannotated process_exit from a dashboard started before
processExitAfter no longer ends the open fallback run when more events
of that ID follow before the next start: a dead process records nothing
more, so it was observed before that run and belongs to the run closed
before it. This replaces the 15 s window, which a delayed callback or a
skewed clock could miss.

* fix(stats): address CI review (#785)

session-owners.jsonl is first written from the per-scope snapshots an
earlier release left: a tool's own ID in the user scope's or a
partition's prompt-token snapshot is that scope's, so a session main
reported in P, compacted and resumed in Q, stays P's instead of being
reported again to Q. An ID the shared snapshot also holds is left out:
main copied the shared file into every scope, so it names no owner, and
every scope already has its baseline. The file is created exclusively,
so a concurrent report in another scope reads the one written first.

* fix(stats): address CI review (#785)

Owner migration reconciles every per-scope baseline of an ID: a tool's
own ID in any of a scope's three snapshots is the scope's that holds its
greatest total (prompts, then tokens). A session main split per event
holds only part of it elsewhere, and a scope may have reported past the
shared total it was seeded with, so neither the first holder nor
leaving shared-held IDs out was right; a session reported with no
prompts, only its intervention count, is found too.

Besides the user scope and the partitions, it reads a project whose data
home is in its workspace that a session still in the log leads to, and
each report records the IDs of its own snapshots that no owner claims
yet, for such a project the log no longer leads to.

* fix(stats): address CI review (#785)

Owner migration assigns no owner when the greatest total ties across
scopes: main copied the shared snapshot into every scope, so equal
totals show only that copy, and each scope already holds the baseline.
A report records an ID of its own snapshots only when they show it
reported it (absent from the shared snapshot, or past its total there),
so a copy no longer claims it either.

A crashed fallback run a start from another process supersedes counts
as the run closed before it, so a late unannotated exit of it no longer
closes the new run.

* test(stats): pin a pre-upgrade exit reported before the next run's first prompt (#785)

The run split is recomputed from the whole log on every report, so once
the next run's first prompt follows the unannotated exit, the exit is
the earlier run's and the next run keeps its ID: the second pull
reports only its delta, not another session.

* fix(stats): give a compacted session resumed elsewhere to the project its transcript started in (#785)

Once compaction dropped every event of a project whose data home is in
its workspace, nothing outside it pointed to it, so a resume of its
Claude session in another project reported the transcript there again.
The transcript itself records where the session started: a Claude
transcript keeps its first cwd when resumed from another project (the
resume appends to the same file), as a Codex rollout keeps its
session_meta. Hooks now record transcriptPath on UserPromptSubmit and
SessionEnd as well as Stop (not SessionStart, whose path on such a
resume names a file that never exists; never Copilot's). A tool's own
session with no owner is the scope's that its origin resolves to, when
that scope's snapshots already hold it; otherwise it is decided as
before.

* fix(stats): address CI review (#785)

A session main split across scopes per event is credited once with
every part it reported: for each scope its `dataHome` names, the
shortest prefix of its events whose metrics reach its snapshot, and the
owner takes the metrics of their union as reported when they exceed its
own entry. Parts counted before any Stop carried the transcript's total
are no longer sent again by the owner, and cumulative Stops are not
credited twice.

A Copilot session with an explicit ID is traced to where it started by
Copilot's own session log, found by the session ID (TeamAI stores no
path of it, #666): its session.start context names the directory.

Compaction keeps a session whose tool process is still running, so a
run that an exit from a dashboard before processExitAfter marked stopped
keeps its start and its ID.

* fix(stats): address CI review (#785)

Owner migration takes a scope's entry as evidence only when its
snapshots show it reported the ID: the shared snapshots (interventions
included) hold none of it, or the scope is past their total. Main copied
the shared file into every scope it ran in, so a copy, even the only
one, names no owner, and the per-report recording follows the same rule.

A session main split across scopes whose events are gone is credited
from the parts' snapshots: a part whose daily entry shows a Stop holds
the transcript's cumulative total, so the greatest counts once; a part
with no Stop counted its own prompts, which add; intervention counts
add, tokens take the greatest. The credit rides on the owner's line in
session-owners.jsonl (numbers only) and is applied once as its baseline.

* fix(stats): keep a tool's own sessions in the first snapshot, parse legacy entries (#785)

Seeding a scope's snapshot from the shared one kept only the runs still
in the log, so a session reported before #795 and compacted before the
scope's first pull was sent again in full when resumed. Only fallback
entries need that filter, against a reused PID; a tool's own session ID
is one session, so its entry is copied whole, as main did.

Splitting a bare entry across runs read the snapshot entry as typed,
and one without `tokens` (hand-edited or truncated) threw and skipped
the whole report; the prompt-token and intervention shares now parse it
as the owner-migration path already does.

* fix(stats): report a resumed Codex rollout after compaction dropped the earlier one (#785)

A Codex build that writes a new rollout per resume restarts its
transcript counters, and the session summed only the rollouts still in
the log. Once compaction dropped rollout A, a resumed rollout B with
smaller counters was compared against A's reported total and reported
nothing until it passed it; routing the session back to the scope it
started in made that loss reach the resume in another scope too.

The prompt-token snapshot now keeps each rollout's reported prompts and
tokens under a hash of its path (no path stored), and a rollout that is
gone keeps its reported totals in the session's sum, so B is reported in
full. A session's prompts also sum its rollouts' Stop counts, which
restart per rollout like the tokens. An entry from before is compared
as a whole once, then kept per rollout.

* refactor(stats): move dashboard scope attribution and session owners out of team-push (#785)

No behavior change. src/dashboard-scope.ts holds which scope reports a
dashboard session (the log split into runs, each given whole to the
scope it started in, and the transcript origin); src/session-owners.ts
holds the machine-level owners index, its seeding from earlier
snapshots, and the snapshot files it reads. team-push.ts keeps the
snapshot adoption, deltas and push, and one reportedBaselines() now
serves both the report and `teamai stats`, which repeated the adoption
sequence.

* test(stats): real CLI resume of a compacted session from a workspace-data project (#785)

A non-git project W keeps its data home in its workspace. W reports a
Claude session; with no owners index and every W event compacted, the
session is resumed in git project Q through the real hook dispatcher,
appending to W's transcript. Q reports only its own session and W the
resumed turn; on a build without the transcript origin, Q reports both.
The fixture gains a second project and hooks sent as the installed
ones send them.

* fix(stats): credit a split session's Stop-derived interventions once (#785)

Interruptions and tool rejections come from Stops, which carry the
transcript's cumulative counts, so a compacted split session's credit
takes the greatest part, as it does for tokens; summing them made the
next cumulative Stop report nothing. Corrections are counted per prompt
in each part's own events, so they still add.

* fix(stats): place a compacted split session's parts by its transcript (#785)

A split session's credit added a part with no Stop to the greatest
cumulative Stop, which already counts that part when it came before the
Stop: after 3 prompts in P and a cumulative Stop of 5 in Q it credited
8, and the next Stop of 6 reported nothing. The credit now keeps each
part (scope key, prompts, whether it ended in a Stop; numbers only), and
the owner places them by the session's transcript, which keeps every
prompt in order with the directory it was typed in: the Stop covers the
first prompts, and only the part's prompts after those add. With no
transcript to place them, they all add, as before.

The Stop scan's human-turn test is now isHumanPromptEntry, shared by
both, so the two count prompts alike.

* fix(stats): keep a dropped Codex rollout's totals for daily and interventions, and migrate whole entries (#785)

An entry from before rollouts were kept is one total. An earlier
release rewrote every session in the log on each report, so it covers
the rollouts begun by the time its file was last written, read before
this report writes it: those still in the log consume it in order, what
is left is the dropped rollouts', kept as one prior rollout, and a
rollout begun later is new. Rollout B after a compacted A is no longer
compared against A's total and lost.

A rollout also keeps its Stop's interruptions and rejections, and its
dropped totals now reach the intervention and daily sums too, not just
prompts and tokens: daily prompt turns and intervention counts of a
resumed rollout were compared against the dropped one's.

* fix(stats): keep every metric of a dropped Codex rollout, with or without tokens (#785)

A Codex session is now kept per rollout whenever its Stops name a
rollout, not only once a Stop carries a token record, so a tokenless
resumed rollout is not compared against the dropped one's totals. Each
rollout also keeps its corrections (a correction goes to the rollout of
its prompt), its active time (each gap to the rollout of the event it
ends at) and its request costs, and a dropped rollout adds them to the
intervention and daily sums, with its cache tokens from its tokens.

The prompt-token snapshot, which holds the rollouts, is written with any
delta, so a rollout whose rejections alone moved keeps its new totals.

* fix(stats): sum a Codex session's rollout costs and keep a dropped rollout's failure (#785)

The daily snapshot took the request costs of the latest rollout only,
so with rollout A still in the log a rollout B was compared against A's
costs and clamped; a Codex session's daily costs now sum its rollouts.

Each rollout also records whether it failed (an error, an interruption
or a correction). A dropped rollout that failed keeps the session
unsuccessful, and one with a correction keeps it corrected, so a clean
later rollout does not turn it into a success.

* fix(stats): keep modern Codex rollouts, and their submit-counted prompts, per rollout (#785)

A Codex session whose tokens come from the thread-level counter
(tokenScope session) was not split into rollouts, so its prompts,
interventions, active time, costs and failure were compared against a
dropped rollout's. It is now kept per rollout like the others; the
counter already spans the rollouts, so no rollout holds tokens of its
own and the session total stays that counter's.

A Codex Stop may count no prompts, so a rollout's prompts are its
Stop's count or else its own submits: a dropped rollout's submit-counted
prompts are no longer lost.

* fix(stats): no legacy tokens on a spanning Codex counter; teamai stats writes no seed (#785)

A whole entry an earlier release left became a prior rollout carrying
its tokens, which were then added to a thread-level counter that already
holds them: rollout B's counter at 530 after A's 500 re-sent 500. A
session whose counter spans its rollouts now takes no tokens from a
dropped or prior rollout.

`teamai stats` only reads, but seeding a scope's first snapshot wrote it
with the current time, which a later report reads as the time an entry
from before covers, taking a rollout begun earlier as reported. A read
that does not persist now writes no seed, and a written seed keeps the
shared file's time.

* fix(stats): read a legacy daily entry's session cost fields as its day's costs (#785)

parseDailySnapshot() dropped the top-level pricedRequests, costMicros,
cache tokens and priceVersion a daily entry from before per-day costs
held, so an entry from before rollouts were kept lost its cost in the
prior rollout, and a later rollout's cost was compared against it and
omitted. They are now read as the session day's request costs, as
computeDailyStatsDelta already reads them.

* fix(stats): keep every Codex variant per rollout, and an older Stop's request cost (#785)

Rollout tracking recognized only `codex`, not `codex-internal` or
`tcodex`, which write the same rollouts; it now uses isCodexTool(). A
rollout's cost was read from requestDaily only, so an older Stop's
requestMetrics left the rollout without cost, and the daily snapshot,
which sums rollouts, omitted it; it is now that Stop's day's cost, as
outside rollouts.

* fix(stats): keep a Codex rollout's latest Stop by timestamp (#785)

A rollout's prompts, interventions and request costs took the last Stop
appended, though background Stop handlers may append an older scan
after a newer one, which then replaced the newer totals. They now keep
the latest Stop by its timestamp, as the rollout's tokens already do.

* fix(stats): an entry from before covers a running Codex rollout only as far as it had got (#785)

Migrating a whole entry from before rollouts were kept consumed it with
each covered rollout's current totals, so a rollout begun before the
entry was written but grown since had its later prompts taken as
reported: an entry of 6 (A's 5, B's 1) with B now at 3 reported nothing.
It now consumes it with each rollout's totals as of the entry's write,
the metrics of the events up to then; what a rollout has done since is
new.

* fix(stats): credit a split session counter by counter; read an old entry's cutoff before its push (#785)

Both credit paths applied only when the parts' prompts exceeded the
owner's, so a part that reported more active time, tokens or costs with
no more prompts was sent again by the owner. The owner's entry is now
raised counter by counter to at least the credit.

An earlier release wrote its snapshot after the push, so events that
arrived during the push predate the snapshot's time without being in
it. The team stats file in the scope's reports checkout was written
after that report read the log and before the push; the earlier of the
two times is now the cutoff an entry from before covers.
2026-09-25 10:36:37 +08:00
Saul Moro 5e5b86d9e9 fix(stats): count every worktree of a repo as that repo (#809) (#813) 2026-09-25 01:39:23 +08:00
Saul Moro a725574b34 fix(usage): cap usage.jsonl in scopes that never report (#788) (#790) 2026-09-24 23:43:42 +08:00
Saul Moro 9f81ae6751 fix(pull): deliver team resources to a worktree added after the last pull (#807) (#811)
state.json lives in the shared project partition, so a new worktree matched
the revision another checkout recorded and took the "Already synced" fast
path, leaving it without the team's skills, rules, agents and docs. The
shared tool targets also made two checkouts with different tool directories
force a full sync on each other on every pull.

Record the revision and targets per checkout in lastPullByWorkspace, keyed by
the checkout path plus the identity of its .git entry so a worktree re-created
at the same path does not inherit the old record. The fast path still requires
the shared lastPullRev, which exclude, tags, roles, projects, init and
bootstrap clear to force a full sync, and a pull that records a new revision
drops the other checkouts' records so every checkout does its own full sync.
2026-09-24 22:22:38 +08:00
Saul Moro ec56a67c1b fix(recall): search nothing in a project whose config cannot be read (#796) (#798)
* fix(recall): search nothing in a project whose config cannot be read (#796)

Detection skips a project config it cannot read and returns what loads
next: a legacy .teamai/ behind a broken partition, which may name another
team, or the user scope. recall searched that knowledge, recorded recalled
counts for it, and `recall --check` answered for it; with nothing behind
the broken file it printed NOT_RELEVANT, so the recall subagent told the
member the team had no knowledge and nobody learned the config was broken.

recall() now listens for the unreadable config before anything else,
searches and records nothing, prints the problem with
BROKEN_CONFIG_ADVICE and exits 1, `--check` included. A silent caller
records it in debug.log only, the rule pull follows since #784. The
teamai-recall agent relays that line instead of skipping the precheck.

* fix(recall): address pre-review findings (#796)

- The relayed line ends with "move it aside and run `teamai init`", and
  the main conversation may not have loaded the teamai skill that asks
  for consent first. The recall agent now tells it to show the line to
  the user and not act on it without their consent.
- Tools that run `teamai recall` directly (the Bash method of the recall
  rule, deployed to every tool) get the same instruction.
- CHANGELOG: only the subagent a pull from this release deploys relays
  the line; a project that broke before the upgrade keeps the old one
  until a pull succeeds there.
- The legacy-team test also asserts no votes land in that team's repo.

* fix(recall): address pre-review findings (#796)

- CHANGELOG: the entry covers `teamai recall <query>` and `--check`. The
  recall subcommands (enable, disable, status, feedback, maintenance,
  promote) still resolve their scope as before; that is a follow-up.

* fix(recall): address CI review (#796)

Reject a missing query before resolving the project, so a bare
`teamai recall` runs no detection (and no self-mode bootstrap). An empty
`--check` still resolves first: it must refuse rather than print
NOT_RELEVANT in a project whose config cannot be read.

Drop the silent branch: `recall` has no --silent flag and no caller
passes `silent`, so it was a contract nothing could invoke.

* test(recall): follow #787's per-scope votes (#796)

#787 moved recalled counts from the shared ~/.teamai/votes/ into each
scope's votes directory, so the #796 tests look for any votes directory in
the sandbox. #787's broken-project recall test expected a search to run;
#796 searches nothing there, which it now asserts, while its checks that no
scope received a vote stay.
2026-09-24 20:54:28 +08:00
Saul Moro 1fd400ecfa fix(pull): sync nothing in a project whose config cannot be read (#784) (#792)
* fix(pull): sync nothing in a project whose config cannot be read (#784)

Detection skips a project config it cannot read and returns what loads
next: a legacy .teamai/ behind a broken partition, which may name another
team, or the user scope. pull() deployed and reported for that team, and
the session-start hook did so on every session (reports-wt/ and
learnings-wt/ appeared in the legacy .teamai/).

pull() now listens for the unreadable config, syncs no scope, prints the
problem with BROKEN_CONFIG_ADVICE and exits 1. A silent pull (the
session-start hook, or a pre-dispatch hook running `teamai pull --silent`)
records it in debug.log only. Agent-root seeding and the package hint
refuse the same way, so a session start there does nothing. Hooks and
usage already follow this rule since #748. The message trimming detectTeam
used moves to config.ts as describeUnreadableConfig so both share it.

* fix(pull): address pre-review findings (#784)

- The session-start handler returns when the dispatcher resolved no
  config for the hook's cwd, which is what an unreadable project config
  resolves to since #748. That one guard replaces the unreadable-config
  sinks added to seedProjectAgentRoot and the package-hint context, and
  follows the #769 contract that handlers read their scope from the
  dispatcher. The handler tests that exercise cwd routing now pass a
  resolved scope; a new one pins that nothing runs without one.
- CHANGELOG and usage guide (en, zh-CN): a session start there runs no
  pull; only `teamai pull --silent` from a pre-dispatch hook writes the
  reason to debug.log.
- skill-data troubleshooting: what `Nothing was synced` means, and that
  moving the config aside and re-running init needs the user's consent.

* fix(pull): address CI review findings (#784)

- The session-start pull is registered with `requiresConfig` instead of
  returning early inside the handler: the dispatcher drops it wherever no
  config resolves, which covers an unreadable project config (#748), and
  spawns no detached pass for it. Where no teamai config exists at all
  it did nothing on main either (no scope to pull, no project root to
  seed, no config for a package hint). The #748 registry test and the
  docs no longer list it as machine-level work.
- `teamai pull --silent` exits 1 on the refusal too. Pre-dispatch hooks
  run it as `… 2>/dev/null || true` (`; exit 0` on Windows), so hosts
  still see success.
- The dispatch-scope test asserts the pull is skipped and resets the
  pull mock it queues.
- skill-serving design doc: `teamai pull` now reports an unreadable
  project config too. skill-data troubleshooting: `teamai doctor` can
  pass there.

* fix(hooks): still pull at session start where teamai is not set up (#784)

A null config from the dispatcher means either "no teamai here" or "the
project config cannot be read". Gating the session-start pull on
`requiresConfig` stopped it in both; only the second must stop it. The
handler now asks `findUnreadableProjectConfig` for the hook's cwd when
no config resolved (a cwd that no longer exists holds none) and runs
nothing when it reports a file. Everywhere else it runs as on main, so
the docs list it as machine-level work again.
2026-09-24 20:07:24 +08:00
Saul Moro dc233e4489 fix(votes): keep votes with the scope they were cast in (#787) (#793)
* fix(votes): keep votes with the scope they were cast in (#787)

Every scope recorded into one ~/.teamai/votes/<user>.yaml, so a vote cast in
one project (recall feedback, a recall search, a Stop whose push failed) was
pushed to the team of whichever scope synced next: the leak usage.jsonl had
before #758.

Votes now live in the data home of the scope that resolves for the session:
<dataHome>/votes/ for a project, ~/.teamai/user-votes/ for the user scope.
The Stop hook uses the config the dispatcher resolved; the pull report, recall
search, `recall feedback` and the vote view read only that scope's votes, and
the CLI readers resolve it with resolveConfigForDir, so an unreadable project
config falls back to no other scope: `recall feedback` exits 1 and the vote
view names the broken file.

The shared ~/.teamai/votes/ is never read. Its V2 `votes` map is the last
merged remote snapshot of whichever team synced, not this scope's history, and
seeding a scope from it would let `recall feedback --negative` push a
decrement and merged timestamps derived from another team. The remote
votes/<user>.yaml format is unchanged.

* fix(votes): address pre-review findings (#787)

- recall search: in a project whose config cannot be read, detection falls
  back to another scope; record no recalled count there, so the vote cannot
  reach that scope's team. Which scope the search itself uses stays #796's.
- recall feedback: with no project config and an empty or invalid user
  config, name the file and the fix (requireInit's error) instead of
  "not set up here".
- CHANGELOG: note the recall search case; the shared directory is never
  read or pushed "by this release" (an earlier release still pushes it).
- Design doc: getUserVotesDir() is the exception to "getters unchanged".

* fix(votes): address second pre-review round (#787)

- recall feedback: name an unusable user config through
  throwMissingOrInvalid (now exported) instead of re-running requireInit,
  which loaded the config twice and printed its parse error twice.
- recall search: a project detection that throws is treated like an
  unreadable config, so no recalled count lands in the fallback scope.
- The recall-search test now asserts the search ran and that neither the
  shared directory nor the broken project's votes/ was written; it fails on
  origin/main too.
- Design doc: re-wrap the edited paragraph.

* fix(votes): record recall votes from a deleted cwd in the user scope (#787)

The previous commit treated any throw from recall's project detection as an
unreadable project, including a cwd that no longer exists. Such a cwd holds
no project and resolves to the user scope everywhere else
(resolveConfigForDir, detectTeam), so its recalled counts belong there.

Tests pin that case and the single parse-error line for an invalid user
config in recall feedback.

* fix(votes): address CI review findings (#787)

- getVotesDir: a historical project-scoped ~/.teamai/config.yaml with no
  projectRoot (schema-valid, not backfilled) made getDataHome throw, so
  recall, feedback and the Stop hook recorded no vote. It lives in
  ~/.teamai, as recall and viz already treat it, so its votes go to the
  user scope's user-votes/.
- git-native-memory design doc: the local votes path is user-votes/.

* fix(votes): address CI review (#787)

- recall feedback --negative counts the upvotes the scope's own team
  already holds (its reports checkout's votes/<user>.yaml plus the
  deltas not yet pushed). A scope's file starts empty on upgrade, so a
  doc upvoted before it was rejected as not found, or as having no
  upvotes once a later recall counted it. The shared ~/.teamai/votes and
  other scopes' teams are never read.
- The adoption judge (#723, merged meanwhile) recorded into and synced
  from the user scope's votes in every scope, which pushed the user
  scope's pending votes to the project's team. It uses the scope's votes
  like the Stop handler.
- votes-scope tests: Stop transcripts prove adoption with a Read of the
  recalled file (#723), and the update.js mock keeps the real lock.
2026-09-24 20:06:35 +08:00
Saul Moro fd0e913814 fix(stats): await the async dashboard scope filter (#795) (#806)
* fix(stats): await the async dashboard scope filter (#795)

#795 made filterEventsByScope async and changed its argument from a
projectRoot/excludeProjectRoots filter to the scope config, while #771
still called it synchronously with the old filter. On main, tsc fails in
stats.ts and every stats-scope test throws "events is not iterable";
`teamai stats` crashes once there are dashboard events.

stats now awaits the filter and passes the scope config, the call pull
makes, which is what #771 set out to do: show what the report sends.
The cwd-based project-root resolution is gone with the old argument.

The user-scope test followed pull's rule before #795 (keep events that
carry no dataHome); it now follows the current one: the user scope
never reports them, so it does not count them.

* fix(stats): subtract the scope's own reported snapshots (#786)

Since #795 each scope reports against its own reported-*.json under its
data home, and the shared ~/.teamai/dashboard files are no longer
written. stats still subtracted the shared files, so after the upgrade
every session reported since counted twice in the headline.

stats reads the snapshots through team-push's readers with the scope
config, including the one-time seed from the shared file.
2026-09-24 19:50:37 +08:00
Saul Moro 8cee7ab23e fix(migrate): keep the legacy .teamai/ while the partition config cannot be read (#797) (#799)
* fix(migrate): keep the legacy .teamai/ while the partition config cannot be read (#797)

planMigration took a partition config.yaml that merely existed as a built
partition and planned a retire-only cleanup, so the first init/pull/push after
the partition file broke renamed the legacy directory to .teamai.bak although
it held the only config that still loaded. Only a partition config that
detection's own reader (readConfigFrom) accepts now counts as built; one that
exists but cannot be read plans nothing, and the next write command after the
fix retires the legacy dir as before. The re-check under the sync lock in
runMigration uses the same rule, so a broken file that appears between planning
and locking skips instead of retiring.

* fix(migrate): address pre-review findings (#797)

- Warn with the file and the first line of the reason when an unreadable
  partition config holds the migration back. The skip was silent, so a member
  had no signal which file to fix, including under --dry-run. The text says what
  happens next and hedges for a file caught mid-write.
- Stand down when the partition dir exists without a config.yaml. Keeping the
  legacy dir made "move it aside and run teamai init" (BROKEN_CONFIG_ADVICE)
  reach the full copy, which removes the partition dir before renaming the
  staged copy in and so deleted its pending learnings, env and clone. The
  re-check under the lock stands down on an existing partition dir too.
- Decide built / unreadable / absent in one helper that uses detection's own
  onUnreadable report, so the plan and the re-check cannot drift.
- Tests: a real YAML syntax error, a partition config that is not scope:
  project, the moved-aside sequence, a fresh project still planning 'full', and
  the re-check retiring when a readable partition appears after planning.
- Design doc and CHANGELOG: list every cause detection reports and the guard.

* fix(migrate): address CI review (#797)

The upgrade note in both usage guides promised an unconditional migration;
it now says a partition whose config.yaml cannot be read, or is missing,
keeps .teamai/ and what the member does next. The full copy's re-check
under the sync lock logged only at debug level when it stood down; it now
gives the same actionable warning as the planner.
2026-09-24 19:20:06 +08:00
Saul Moro 352cfc4ccc fix(report): each scope keeps its own reported dashboard snapshots (#786) (#795)
* fix(report): each scope reports only the dashboard sessions recorded in it (#785)

Every scope read one machine-wide events.jsonl and picked its sessions out
by cwd prefix. The user scope excluded nothing, so a user-scope pull
reported every project's sessions (and, through the shared reported
snapshots, took them from the project's own report); Copilot sends no cwd,
so a project never reported its Copilot sessions; and a raw cwd under a
symlink or /tmp never matched the realpath'd projectRoot.

The hook now stamps each event's dataHome with the data home of the scope
the dispatcher resolved (the key the per-scope usage file already uses), and a
report keeps only its own scope's events, comparing realpath'd keys. A
project also owns its in-repo .teamai key, where hooks record until
migration moves it to a partition. Events written before this carry no
dataHome: a project keeps those whose realpath'd cwd is under its root, the
user scope never reports them. The log stays machine-wide for the
dashboard UI, stats --by-repo, session save and the contribute check.

Removes the excludeProjectRoots option, which pull only ever passed as []
(the user target exists only when no project config resolved), and the
projectRoot option now carried by selfConfig. The usage guide documents how
to remove by hand a skill an earlier release pushed into stats/<user>.yaml
from another project.

* fix(report): each scope keeps its own reported dashboard snapshots (#786)

The report sends per-session deltas against reported-*.json snapshots that
every scope shared. A session whose events belong to two scopes (a cd into
another project mid-session) was then reported by the first scope, and the
second compared its own part with the first scope's totals and sent nothing.

Each scope now keeps its snapshots in <dataHome>/dashboard/, and the user
scope, whose data home holds the shared files, in user-reported-*.json. The
first time a scope needs one it copies the shared file, so the first report
after the upgrade sends nothing already reported; after that it reads only
its own. The user scope moves too, unlike the ticket proposed: had it kept
writing the shared file, a project seeding later would copy the user scope's
part of a split session and report nothing for its own. The shared file is
no longer written, except by an earlier release after a rollback, which only
a scope not yet seeded reads.
2026-09-24 19:18:09 +08:00
Saul Moro 2ab697d062 fix(usage): discard pre-upgrade usage and stop falling back past an unreadable config (#748) (#758)
* fix(usage): discard pre-upgrade usage and stop falling back past an unreadable config (#748)

Follow-up to #753, from its review.

- resolveConfigForDir returns null when any project config was reported
  unreadable, even if a lower-priority one (a legacy .teamai/ behind a broken
  partition) loads: that one may name another team.
- The user scope's usage.jsonl is the old shared file. #753 only emptied it on
  a machine's first user-scope init, so a machine that already had a user
  scope reported every project's pre-upgrade usage to it. The first access
  after the upgrade now discards what an earlier release left there and
  writes ~/.teamai/usage-per-scope. One process discards, under acquireLock;
  concurrent hooks wait for the marker, so none deletes what another recorded.

* fix(review): parse the empty usage file in its test; scope the fallback wording to team hooks (#748)

- "handles empty file" wrote no marker, so the discard removed the file and
  the read passed on a missing file. A first read now settles the file as
  the scope's own, and the test asserts the file survives.
- The session-start pull still resolves its project on its own, so the
  "never falls back to a lower-priority config" rule is stated for team
  hooks and skill usage only (CHANGELOG, usage guide en/zh-CN).

* fix(usage): keep the user scope's usage in its own file, safe across a rollback (#748)

The usage-per-scope marker could not tell a pre-upgrade event from one an
earlier release appends after a rollback, so a reinstall reported those to
the user-scope team. The user scope now records in ~/.teamai/user-usage.jsonl,
which no earlier release writes; ~/.teamai/usage.jsonl is removed, never read.
Drops the marker, its lock and the bounded wait.

* fix(usage): a failed removal of the shared usage file does not stop the user scope (#748)

Also names the user scope's own file where comments and the design diagram
still described every scope's usage as <dataHome>/usage.jsonl.

* docs(changelog): drop the claim that teamai doctor reports an unreadable project config

resolveDoctorContext falls back past an unreadable project config the way
detection does, so doctor diagnoses the config it falls back to and says
nothing about the broken one (#752).

* fix(usage): leave the shared usage file in place instead of removing it on every access (#748)

getUsagePath deleted ~/.teamai/usage.jsonl on every user-scope call,
including each hook append and the read-only `teamai stats`. The user scope
never reads that file, which is what keeps its events off the team; the
delete added a side effect to a path getter and a failure path to guard.
2026-09-24 15:04:18 +08:00
Saul Moro cc2772114f fix(hooks): run every handler in the scope the dispatcher resolved (#752) (#769)
* fix(hooks): run every handler in the scope the dispatcher resolved (#752)

hook-dispatch resolves the scope from the payload cwd, then chdirs there.
The correction keywords, votes-sync and webhook-dispatch read their config
again through autoDetectInit() on the process cwd, so when that chdir failed
(a deleted worktree) and the host started the hook inside another project,
they used that project: its keywords, its member for the session's votes,
and its webhooks.

HookHandler.execute now receives the config the dispatcher resolved, and the
three handlers read their scope from it.

* fix(hooks): start the background pass when the hook's cwd is gone (#752)

The detached child was spawned in the payload cwd. On macOS and Linux a cwd
that no longer exists (a deleted worktree) fails the spawn, the error is
swallowed, and no background handler runs: no session-start pull, webhook or
update check. Start it in the temp dir instead, as the Windows WMI launch
already does. The child resolves its scope from the payload, and its handlers
now use that config, so the directory it starts in names no project.
2026-09-24 10:13:54 +08:00
Saul Moro 6a53b6f925 fix(lock): never rename over a lock that was just released or is still being written (#760) (#761)
* fix(lock): never rename over a lock that was just released or is still being written (#760)

A stale verdict lets the reclaimer rename over the lock, and two live states
read as stale: a file that vanished (released, and possibly re-created by a
third process before the rename) and an empty file (its owner opened it but
has not written it yet). lockState now tells live, stale and missing apart:
a missing lock gets one more exclusive create, and an empty one reads as
held until it has stayed empty for 5s (its owner died before writing it).

16 processes x 2000 attempts on one lock: 4-18 overlapping holders per run
before, 0 after, with acquisitions in the same range.

* fix(review): EPERM owner is live; cover the second pass; complete lock users (#760)

- process.kill(pid, 0) throwing EPERM means the owner is alive under another
  user (e.g. `sudo teamai`); it read as stale and was renamed over.
- Test the missing branch under the reclaim sentinel, where the rename lives.
- Comment and design doc state exactly what reads as stale (an unreadable
  lock still does); CHANGELOG lists every lock user.

* fix(review): publish the lock with its content in place (#760)

An empty lock expired after 5s, so a creator stalled between open and write
(SIGSTOP, sleep, slow I/O) could be taken over while it still held the lock.
exclusiveCreate now writes the payload to a private temp file and hard-links
it to the lock name (link fails with EEXIST like O_EXCL), so the lock never
exists empty. Without hard links it falls back to O_EXCL, where the 5s grace
still applies. The same create backs the reclaim sentinel.

* fix(review): O_EXCL fallback verifies its payload; unreadable lock is held (#760)

- Without hard links the lock is created with O_EXCL and sits empty until
  written. A creator stalled past the 5s grace could be replaced and still
  report success; it now reads the lock back and holds it only if its own
  payload is there.
- A lock that exists but cannot be read (EACCES: another user's 0600 lock)
  read as stale and was renamed over; it now reads as held.
- Use fse.link like every other lock operation, and mock it in update.test.

* fix(review): a wx creator holds only a lock it wrote inside the grace (#760)

- Without hard links, a reclaimer that read the lock empty could rename over
  it after the stalled creator had written and checked it. The creator now
  keeps the lock only if it wrote it within half the grace, so no reclaimer
  can have judged it stale; otherwise it gives it up.
- An unreadable lock logs a warning naming the file.
- Migration skips lock artifacts (<lock>.*.tmp, .sentinel, .new-*), which a
  contending pull creates and removes during the copy.
- Stale comments; the stall tests restore their spies in afterEach.

* fix(review): reclaim only a lock whose owner is provably dead (#760)

Each timing rule for locks that name no owner (empty, partly written) left
an ordering where two processes held the lock: a partial payload read as
stale at once, an older teamai stalled past the grace, a slow fallback
creator yielded but kept blocking. Only ESRCH now makes a lock stale; a lock
that names no owner, cannot be read, or has an EPERM owner is held, and a
warning names it so a crash leftover can be removed by hand. Drops the 5s
grace, the 2.5s creator limit and the read-back.

- parseLockContent: a bare legacy PID is valid JSON (a number) and was
  returned as unparseable, which only worked while unparseable meant stale.
- Migration skips only the real lock artifact formats, not every <lock>.*.

* test(lock): exercise the live-holder verdict; drop grace leftovers (#760)

- "returns false when a live process holds the lock" failed on the temp write
  before any verdict ran; link now rejects with EEXIST on the lock path.
- Remove Date.now/utimes staging the removed grace no longer reads, merge the
  duplicate empty-lock tests, restore the stall spies in finally.
- Docstrings and design doc: only a lock that names no owner or cannot be read
  is warned about; the sentinel-steal residual is stated as it is.
2026-09-24 07:36:20 +08:00
Saul Moro 72305c68b8 fix(skills): one share gate, actionable refusals, and a louder stub deploy (#747)
* fix(skills): one share gate, actionable refusals, and a louder stub deploy

Follow-ups from the review of #699:

- The Stop-hook reminder and `teamai skill get share` ask one gate
  (`shareGate`, through `contributeHintAllowed`). The hook skipped the
  unreadable-project-config check, and the legacy `teamai contribute-check`
  command, still called by hooks written before the dispatcher, checked
  nothing, so both nudged towards a command that refused.
- The gate reads only a config load failure as "cannot be loaded"; any other
  fault propagates (the hook withholds the reminder and logs it at debug).
- A `config` refusal says what failed (the file and position for a parse
  error) instead of pointing at `teamai doctor`, which cannot see a broken
  config. `skill show` now refuses through the same helper, so its hint moves
  from stdout to stderr like `skill get` and `skill path`.
- `pull` warns when the discovery stub cannot be deployed (it was an empty
  catch on the fast path and a debug line on a full sync), and so does the
  legacy prune.
- Error text no longer claims a reason was logged when none was: an empty
  config is named as empty, and `init` points at ~/.teamai/debug.log, where
  every path that deploys nothing now records why.
- `core` routes a bare `/teamai` right after a friction reminder to `share`,
  as the stub already said.
- The command drift guard rejects an unknown subcommand inside a group
  (`teamai skill gett core` passed before).
- The contribute-check e2e asserts the reminder's real text again; the usage
  guides (EN, zh-CN) and the design doc cover the config refusal, the gate and
  the reminder routing.

* test(learnings): retry temp-dir cleanup that races a detached git gc

A push into the bare origin can leave `git gc --auto` writing to
objects/pack after the test returns; the single rmdir in afterEach then
fails with ENOTEMPTY (seen on CI, Node 22 ubuntu, #747).

* fix(skills): gate skill show before its lookups, name the failing field

Review of #747:

- `skill show share` under a broken project config searched the user
  config's team repo and agents, which detection falls back to, and printed
  a `share` found there. It now asks the gate first and refuses on a config
  block before any lookup. With an empty user config it refuses instead of
  ending in a stack trace.
- A config that parses but fails validation reported the Zod JSON dump,
  whose first line is `[`, so the refusal said `config.yaml: [.`. Every
  config loader now reports each issue as `field: reason` on one line.
- The docs and skills that describe the share reminder or the refusal say
  it is withheld on a read-only source and while the config cannot be
  loaded, and that a validation failure names the field: product-overview
  and usage-guide (EN, zh-CN), designs/skill-serving.md, core/SKILL.md,
  contribute-member, setup-admin, join-member and manage-admin.

* fix(skills): no share reminder where teamai is not set up

`contributeHintAllowed` fell open with no config at all, so a caller other
than the dispatcher (the legacy `teamai contribute-check`) still nudged in
projects that never set up teamai, which have no team to share with
(#748). It now returns false there. Serving the skill stays fail-open.

* fix(contribute-check): gate the legacy reminder on the session's cwd

`teamai contribute-check --stdin` asked the share gate about the directory
the hook process started in, while the session analysis used the payload
cwd. Started outside the project, it could read the user config and nudge
where `teamai skill get share` refuses (a project config that does not
load). It now moves to the payload cwd first, as hook-dispatch does.

* fix(pull): keep a debug.log record when the stub cannot be deployed

The previous commit turned both deploy catches into `log.warn`, which is
muted in silent mode and never reaches debug.log, and a SessionStart pull
runs detached with its output discarded. So the automatic pull, the one
that deploys the stub for most members, lost the only persistent record
it had. Both catches now warn and write the same line to debug.log.

* fix(skills): skill show and list never answer for the fallback team

Known issues left by #747:

- `skill show <name>` and `skill list` on a config that exists but does
  not load ended in a Node stack trace, and under a broken project config
  they searched the user config detection falls back to: another team's
  repo and agents. Both now ask `detectTeam`, the one place that tells
  "this team", "no team" and "cannot tell, and why" apart (`shareGate` is
  built on it). Without a usable team, `show` answers from the package
  alone and `list` prints only the packaged catalog; both say what failed
  on stderr and exit 1.
- A teamai.yaml that exists but fails validation was reported as "not
  found. Check your repo path". It is now named as invalid, empty or
  unreadable, like the local config.

* fix(skills): the gate reads the session's directory, and no project config is skipped

Codex review of 5793758:

- A project-location config that is not `scope: project` (or omits
  `scope`, which defaults to user) was skipped without a word, so the gate
  read past it to the user config. It is now reported as unusable, unless
  it is the user config itself, as when running from HOME.
- The legacy `contribute-check` changed into the payload cwd and, if that
  failed, asked the gate about the directory the process started in. It
  now passes the payload cwd to the gate (`detectTeam(cwd)`), and a cwd
  that no longer exists holds no project config, so only the user config
  is asked, as #753 does.

* fix(logger): record warnings in debug.log

`log.warn` wrote to the console only and was muted in silent mode, so a
detached SessionStart pull, whose output is discarded, lost every warning:
the stub deploy failure and the legacy prune among them. Warnings now reach
debug.log like debug and error lines. `warnStubNotDeployed` drops the
second `log.debug` call, which printed the line twice under --verbose.

* fix(skills): the dispatcher gate reads the payload cwd; a symlink is not HOME

Codex review of b0583f5:

- The dispatcher's `contribute-check` and `pending-hint` handlers asked
  the gate about the process's directory, trusting hook-dispatch's
  `chdir`; when that failed, the launcher's config decided. They now pass
  `resolveHookCwd(stdin)`, as the legacy command does.
- The HOME exception for a non-project scope compared the config file's
  real path, so a project config symlinked to ~/.teamai/config.yaml passed
  for the user config. It is now decided by the project's location: its
  root is HOME.

* fix(skills): only a missing cwd falls back to the user config; load it once

Codex review of 15a5b5d:

- `detectTeam` read any failure to see the payload cwd as "deleted", so a
  cwd it could not open (no permission, a path through a file) fell back
  to the user config and could allow the reminder. Only ENOENT does now;
  anything else is `unusable` and withholds it.
- `skill show share` and `skill list` loaded the config twice, through the
  gate and then the team lookup, and reported a broken one twice. Both
  detect the team once and hand it to the gate.

* fix(logger): a file-only record instead of persisting every warning

b0583f5 made every `log.warn` append to debug.log, wider than the two
failures it was for, and it wrote unrelated subprocess errors to disk.
`log.warn` is console-only again; `log.persist` writes one line to
debug.log and never to the console. The stub deploy catches and the
legacy prune catch use both, so a detached SessionStart pull keeps the
record and --verbose prints it once.
2026-09-23 22:47:00 +08:00
Saul Moro 5502d8ec10 fix: keep TeamAI out of projects that never set it up (#748) (#753)
* fix(hooks): run team hooks only where TeamAI is set up (#748)

Hooks of a project-scope install live in HOME, so they fire in every
project on the machine. With no config for the hook's cwd they ran anyway:
the Stop share nudge (even with recall off), the TodoWrite recall nudge,
and local capture of sessions and skill usage that another project's
report later pushed to its team.

Handlers that need a team now declare requiresConfig; the dispatcher drops
them when neither a project nor a user config resolves. Only machine-level
work runs there: update check, session-start pull, local agent, package
pending hint. The legacy track paths skip recording the same way.

* fix(stats): keep skill usage in the scope that recorded it (#748)

Every scope appended to one ~/.teamai/usage.jsonl, so whichever project
pulled next reported every project's skills to its own team.

Usage now goes to <dataHome>/usage.jsonl of the scope that resolves for
the session's directory (detectProjectConfig(cwd) ?? loadLocalConfig(),
as the dispatcher does). Each report reads and truncates only its own
file; teamai stats shows the current scope. A machine's first user scope
starts with an empty file: what it held cannot be attributed.

* fix(review): one scope resolver for hooks, usage and stats (#748)

- resolveConfigForDir (config.ts) is the single resolution the dispatcher,
  the usage writers and readers, and teamai stats use. A host that sends no
  cwd (OpenClaw) resolves from the process cwd, and an unreadable project
  config yields null instead of falling back to the user scope.
- readUsageEvents / truncateUsageAfterReport require a scope; no path reads
  the old shared file by default. readKnownSkills reads its scope's file.
- Real-dispatch tests: no-config session leaves no trace, cwd-less host,
  unreadable project config.
- Docs: machine-level handler list, data-directory-layout note.

* fix(review): scope wording in CHANGELOG and readKnownSkills (#748)

* fix(review): gate legacy contribute-check and tolerate a missing cwd (#748)

The hidden 'teamai contribute-check' command, still called by hooks of
older installs, had no config gate, so it nudged in projects without
teamai. It now resolves the scope like the dispatcher and skips when
none resolves.

resolveConfigForDir no longer throws for a directory that does not
exist (simple-git refuses it); a hook naming a deleted worktree falls
back to the user scope.
2026-09-23 20:36:07 +08:00
Saul Moro 95cea46182 fix(manifest): reject namespace strings that are not safe path segments (#710)
* fix(manifest): reject namespace strings that are not safe path segments

A resource namespace becomes a directory component (skills/<ns>/,
agents/<ns>/, learnings/<ns>/) exactly as a project id does, but only the
project id was refined. Both manifests accepted a namespace like
'../../evil', and roles.yaml had the same hole.

Guarded at the manifest boundary, which is where projects.ts already claims
it is enforced and the only place these strings enter the process. The
existing isSafeNamespaceSegment guards in contribute.ts and
resources/agents.ts stay as defence in depth.

The doc comment pointed at the wrong layer: an id read from a hand-edited
config.yaml resolves through getProjectOrThrow, so it can only ever name a
project the manifest already validated. Corrected to say so.

No fixture or e2e manifest in the repo ships a namespace containing '/' or
'..', so nothing that parses today stops parsing.

* docs(changelog): note the manifest namespace guard

* fix(manifest): guard namespaces against traversal only, and say what is wrong

Review findings on #710.

The first cut reused SAFE_ID (^[A-Za-z0-9._-]+$) for resource namespaces. That
allowlist is right for a project id, which is also typed on the command line and
split on commas, but for a namespace it rejects far more than traversal: a team
whose skills live under a non-ASCII directory, or one with a space in the name,
would have stopped parsing although the directory is perfectly safe. The
namespace guard now tests what actually matters -- no path separator, no `:`
(drive-relative on Windows), no control character, and not `.` or `..` -- while
the project id keeps its narrower spelling.

Both live in src/manifest-schema.ts, which is what roles.yaml and projects.yaml
genuinely share; roles.ts no longer reaches into projects.ts for the schema.

Running the CLI against a manifest with `../../evil` showed the second half: the
zod failure escaped as a raw ZodError, so `teamai pull` printed a validation
object instead of a sentence. parseManifest now reports it the way the
hand-written checks beside it do, naming the entry:

  Invalid projects manifest: projects.0.resources.skills.1: resource namespace
  must be a single path segment (no '/', '\', ':' or control characters, and
  not '.' or '..')

Docs: the namespace rule is stated where each manifest is documented, in both
usage guides and in the multi-project design doc.

* fix(manifest): reject DEL and C1 controls in a namespace too

Review P2 on #710: the guard rejected only U+0000-U+001F while the error
message and the docs promise every control character, so U+007F and the
C1 range U+0080-U+009F still parsed. Range extended and the three ranges
covered in the projects and roles fixtures.

* fix(manifest): reject the Win32 dot/space aliases of . and ..

Review P1 on #710: the guard tested for the exact strings '.' and '..',
so '.. ', '.. .' and '...' passed. Win32 strips trailing spaces and
periods from a path component, so each of those reaches the filesystem
as '..' and escapes the namespace directory it was supposed to name.

A segment of nothing but dots and spaces is '.' or '..' in disguise and
is refused as such; 'a..' keeps parsing, since it stays inside its
parent. The project id, whose allowlist already excluded spaces, refuses
any run of dots for the same reason.

* fix(manifest): keep the project id rule untouched, scope the role docs

Review on #710.

The dot/space fix reached further than it needed to: tightening the id
to reject every run of dots also rejected '...', a working POSIX
directory name the id rule has always accepted, so a manifest that
parses today would have stopped. The id is back to the exact '.'/'..'
check it had before this PR, with a test that says so. Only the
namespace rule moves.

The role docs claimed every namespace under resources: follows the rule,
but roles.yaml's learnings: is accepted for backward compatibility and
ignored at runtime -- it names no directory, so holding an old manifest
to the rule would reject it over a field nothing reads. Both usage
guides and the design doc now name the fields that do take effect.

* docs(manifest): state the id and namespace rules separately

Review P2 on #710: after the id was left on its old rule, the docs still
described one rule for both, so they claimed a project id rejects any
name made only of dots and spaces while '...' parses. Each rule now
stands on its own in both usage guides and the design doc, and the
projects.ts comment says why the id is not held to the namespace rule.

* fix(manifest): reject a namespace with a trailing '.' or space

Review P1 on #710: refusing only names made entirely of dots and spaces
left the aliasing half open. Win32 strips trailing periods and spaces
from every path component, so 'frontend.', 'frontend ' and 'frontend..'
all resolve to 'frontend' -- one namespace reading and writing another's
directory, which is the isolation a namespace exists to provide.

The rule is now the trailing character itself, which covers the escape
('.. ' arriving as '..') and the aliasing in one test, and '.' and '..'
fall out of it. A dot inside a name ('alpha.v2') is untouched.

* fix(pull): a roles manifest that does not parse must not widen delivery

Review P1 on #710. resolveResourceNamespaces caught every failure from
loadRolesManifest and carried on with no role filter, which for a member
with no active project means an unfiltered sync: making the schema
stricter would have turned 'skills: [../../evil]' into 'deliver every
namespace', the opposite of what the guard is for.

The catch was covering two cases at once, because loadRolesManifest
throws both when the file is absent and when it is invalid. Only the
first is the legacy, unfiltered case, so it now throws a typed
RolesManifestMissingError and the catch reacts to that alone. An invalid
manifest propagates and pull fails the scope with the entry named --
exactly what an invalid projects manifest already does.

Verified against the real CLI: with 'skills: [evil/nested]' pushed to the
team repo, pull reports the failed 'Skills to deliver can be resolved'
check and the three delivered skills are left untouched; restoring the
manifest syncs them again.

* fix(manifest): narrow 'absent' to ENOENT, refuse Windows device names

Review on #710, two of the three findings; the third was a stale read of
the PR description, which the e2e section had already been rewritten to
match and which is now updated before the push rather than after.

readFileSafe returns null for every read failure and for an empty file,
so an unreadable roles.yaml was indistinguishable from one that was never
written -- and 'never written' is the one case allowed to relax role
filtering. projects.yaml had the same hole, where a null manifest means
'this team is not partitioned'. Both loaders now read the file directly:
ENOENT is absence, and a permission error, a directory or an empty file
is an error that fails the pull.

Windows opens a device for CON, NUL, AUX, PRN, COM0-9 and LPT0-9 in every
directory, extension or not, so a namespace spelled that way cannot be
the directory the manifest names. A name that merely starts like one
(console, community) is untouched, and the project id stays out of this
rule as it stays out of the others: it is a working POSIX name the id
rule has always accepted.

* fix(manifest): narrow every roles fallback, drop COM0/LPT0 from the device set

Review on #710.

The fail-closed change covered resolveResourceNamespaces but not the
other callers that fall back when the loader throws, so a malformed
manifest still reached an unfiltered sync by another route: bootstrap.ts
left the member role-less while auto-selecting the sole role,
resources/skills.ts and push.ts guessed the namespaces from the role ids,
and config.ts skipped the legacy migration and left the role unset. Each
now reacts to RolesManifestMissingError alone. roles-cmd.ts keeps its
broad catches on purpose: those commands report the error to the person
running them instead of deciding what to deliver.

Windows reserves COM1-COM9 and LPT1-LPT9, not COM0/LPT0, so the guard was
rejecting two ordinary directory names for no safety gain. Both are now
covered by the test that pins 'console' and 'community' as valid.

* fix(manifest): prove absence before trusting it, add the superscript devices

Review on #710.

ENOENT is not proof that a manifest is absent: a committed symlink whose
target is missing reads exactly the same way, and absence is the one
answer that lets a caller relax its filtering. The path is now lstat-ed
before absence is believed, so a dangling link is an error like any other
unreadable file.

Windows reads the superscript forms of 1, 2 and 3 as device numbers, so
COM and LPT followed by one of those join the ASCII-digit set.

The third finding, that resolveResourceNamespaces returns before reading
roles.yaml, is not a fail-open and is left as it is: that branch is
reached only when the member has no role, and a role-less member gets the
same unfiltered sync from a perfectly valid manifest, since every role
namespace below is gated on primaryRole. Reading the manifest there would
only add a new way for their pull to fail. The reasoning now sits in the
code beside the early return.

* fix(init): a broken roles manifest must stop init, a skipped prompt must not

Review on #710.

Both init paths swallowed every role-selection failure and carried on
without a role. A role-less config matches every role when hooks are
reconciled, so a manifest that does not parse installed exactly the hooks
it restricts.

Narrowing the catch to RolesManifestMissingError alone was too much: the
same block also absorbs a person skipping the role prompt, and a
non-interactive run reaches it, so init would have started failing for
anyone who does not pick a role. That case is now its own type,
NoRoleSelectedError, and the two lenient cases are named while a parse
failure or an unknown --role propagates. Both catch blocks read the same.

Also: ENOENT proves nothing about absence when the DIRECTORY is a
dangling link -- readFile and lstat on the file both report ENOENT -- so
the path's components are walked, and the first link that leads nowhere
is reported instead of being read as 'no manifest'.

Verified in init.test.ts (malformed aborts and writes nothing, absent
still initializes role-less) and against the real CLI in single-repo
mode: no manifest exits 0 with the role unset, '../../evil' exits 1, and
a valid manifest with --role sets primaryRole: frontend.

* chore(ci): re-run review against the rebased head

No code change. The Codex review workflow re-reviews on push, and the
PR body now carries the real-CLI matrix run on the rebased head.

* fix(manifest): expand '~' when reading a manifest, guard role ids used as fallback namespaces

Review findings on #710 after the rebase.

readManifestFile replaced readFileSafe/readFileIfExists, which expanded a
home-relative repo.localPath. Without the expansion a documented
`~/.teamai/...` path is searched under the current directory, read as
absent, and roles.yaml absence relaxes the filtering. The path is expanded
before both the read and the dangling-link walk.

When roles.yaml is absent, skills.ts and push.ts fall back to the role ids
as namespaces. A role id is an unrestricted string, so a value such as
'../../outside' reached path.join, and SkillsHandler.removeItem could
recurse outside the team repo. Both fallbacks now pass the ids through the
namespace guard and fail with the rule's message.

* fix(config): expand '~' in repo.localPath at the config boundary

A home-relative repo.localPath reached simple-git, the manifest readers and
every resource path unexpanded, so `teamai pull` failed with
'Cannot use simple-git on a directory that does not exist'. The schema now
expands it once, at parse time, so no consumer has to. expandHome moves to
utils/home.ts (fs.ts re-exports it) so types.ts can import it without
pulling in the fs helpers.

* fix(manifest): reject the CONIN$ and CONOUT$ console devices too

Review finding on #710. Windows opens the console for these names in any
directory, extension or not, the way it does for CON, so a namespace spelled
that way cannot be the directory the manifest means.

* fix(push): guard the role id silent mode uses as a namespace

Review finding on #710. With a valid manifest that maps the role to several
skill namespaces, silent push assigned primaryRole as the namespace without
the check the fallback path already has. It now goes through the same
guard, so 'frontend.' or 'CON' fail the push instead of becoming a path.

* fix(manifest): reject namespaces that differ only by case, classify the guard as breaking

Review findings on #710. Two namespaces of one resource type that differ
only by case (or Unicode normalization) name a single directory on the
default Windows and macOS filesystems, so a role scoped to 'frontend' would
read 'Frontend' too. Each manifest is checked when it loads; the pull path
checks roles.yaml against projects.yaml as well, since both share skills/,
knowledge/ and agents/.

The namespace guard makes a manifest that parsed before fail every pull, so
the changelog entry moves under Breaking Changes.

* fix(pull): check roles.yaml against projects.yaml for role-less members too

Review finding on #710. The cross-manifest case-alias check ran only when
the member had a role, so a project-only member pulling skills/Common with a
role's skills/common in the same repo was not stopped. roles.yaml is now read
whenever a projects manifest is in play; an absent one stays absent, a broken
one fails the pull as it does for a member with a role.

* fix(manifest): fold case the way filesystems do when comparing namespaces

Review finding on #710. The alias key was normalize('NFC').toLowerCase(),
which is not case folding: 'σ'/'ς' and 's'/'ſ' stayed distinct although
case-insensitive filesystems give each pair one directory. The key now
upper- then lowercases each code point on its own, which folds both pairs
and sidesteps the context-sensitive final-sigma rule. It errs toward
joining ('ß'/'ss', 'ı'/'i'), which can only reject a pair.

* fix(config): keep the config loadable when the roles manifest is broken

Review finding on #710. migrateLegacyRoleConfig rethrew a manifest parse
error, which loadLocalConfig caught and turned into null, so every command
reported "teamai is not initialized" — pull included, leaving the member no
way to fetch the fixed manifest. The migration now skips with a warning and
returns the config unmigrated.

That alone would widen delivery: a role-less member who would have been
migrated to 'hai' reached resolveResourceNamespaces' unfiltered early return
without roles.yaml being read. roles.yaml is now read for every member before
that return, so a broken one fails the pull (absent still means unfiltered).

* fix(status): report a resource type it cannot scan instead of crashing

Found running the real CLI on #710. scanLocalForPush resolves namespaces
through the roles manifest (agents via resolveResourceNamespaces, skills
when it falls back to role ids), and a manifest that does not parse now
throws there instead of being read as "no filter". status let that escape
as a stack trace after printing half its report. Status is where a member
looks to find out why pull failed, so it now warns with the error for that
type and lists the rest, as it already does for git status.

* fix(status): do not report "(none)" when a resource type could not be scanned

Review finding on #710. With every successful scan empty and one type
failing, status printed the warning and then "(none)", which reads as a
complete clean result. It now says "(none in the types that could be
scanned)" in that case.

* fix(config): a role the manifest could not resolve matches no role-scoped entry

Review finding on #710. When the legacy role migration cannot read the
roles manifest, the config stayed plainly role-less, and resolveMembership
reads role-less as "every role": hooks, MCP servers and env variables scoped
to roles reached a member the manifest would have made 'hai'. The pull
refused the manifest for skills, but those reconcilers still ran.

The migration now marks the in-memory config roleUnresolved, a runtime-only
field like dataHome that serializeLocalConfig drops and the schema strips on
load. activeRoleIds returns [] for it, so role-scoped entries reach nobody,
unscoped ones apply as before, and the reconcilers remove role-scoped entries
already installed. The next load decides the role again.
2026-09-23 19:22:54 +08:00
Saul Moro ca6e51251f feat(skill): serve builtin skill content from the CLI, deploy a discovery stub (#699)
* feat(skill): serve packaged skill content from the CLI

Add `teamai skill get <names...> [--full] [--all]` and `teamai skill path
[name]`, so an agent can read built-in skill content that always matches the
installed CLI version instead of a copy deployed into its skills directory.

`get` prints SKILL.md byte for byte, frontmatter included, with {SKILL_DIR}
resolved to the absolute packaged directory so documented script invocations
run as-is. `--full` appends references/ and templates/, walked recursively and
sorted by relative path, because our references nest one level deeper than the
flat layout agent-browser assumes.

Content goes to stdout and every diagnostic to stderr, so the output stays
byte-exact when piped. An unknown flag warns and continues; an unknown name is
fatal, since acting on the wrong skill is worse than a retry.

`skill list` gains the served catalog and `--json`; `skill show` resolves
packaged skills before the installed-agent fallback, which is what keeps it
working once the deployed unit becomes a stub. Legacy directory names resolve
as aliases.

Refs #678

* refactor(skills): move content to skill-data and deploy a single stub

Agents now receive one file: `skills/teamai/SKILL.md`, a discovery stub of
about 2 KB whose description carries the triggers of every workflow and whose
body holds the commands that load them. The workflow content moves to
skill-data/{core,share,wiki}, which is never deployed and is printed by
`teamai skill get`.

Before this, `deployBuiltinSkills` copied three whole trees — 176 KB — into
every installed agent on every pull, so the text an agent read could disagree
with the CLI it documented until the member ran a pull, and a machine with ten
agents held ten copies. skills/ keeps its meaning ("everything here is
deployed"), which is what lets BUILTIN_SKILL_NAMES collapse to one name.

The stub is copied verbatim: no ensureSkillFrontmatter on the way out, so a
deployed copy that differs from the packaged one is a bug rather than a
variant. Recall no longer gates deployment, since the stub routes to every
workflow; the run-time gate for share lands with the pruning pass.

Uninstall learns the legacy directory names, which it would otherwise leave
behind on every machine that upgraded.

"skill-data" is added to package.json files, with a test that asserts it
through `npm pack`: without that entry every test still passes against the
repo and `skill get` serves nothing once installed from the registry.

Refs #678

* fix(skills): repair stale commands, broken refs and frontmatter

An audit of the three builtin skills found 60 defects. This fixes the ones that
survive the move to skill-data, and splits the two skills that were carrying
more than one job.

Stale CLI surface. The wiki skill advertised `teamai extract graph`, a command
that has never existed. The hand-written "ground truth" cheat sheet in the
teamai skill omitted 19 real commands while telling the agent that anything
missing from it could be checked with `--help` — which fails for the flags
`--help` hides. The cheat sheet is replaced by
`skill-data/core/references/commands.md`, rendered from the CLI's own command
table, with hidden flags marked as such. Two tests guard it: one regenerates
the file and diffs, the other resolves every `teamai …` string written anywhere
in skill-data against the command table and fails on an unknown command or
flag. That second test is the one that would have caught e151d43, 1ca43ac,
8bb0548 and 2ddb546 before they shipped; it carries a case proving it catches
`teamai extract graph`.

Paths. Everything the skills told an agent to read or execute assumed the skill
sat in the agent's own directory: `python3 scripts/scan_repo.py` from a cwd that
is the target repo, references cited by bare filename in two different
conventions, methodology paths handed to sub-agents inside input packets. All of
them now go through {SKILL_DIR}, which `skill get` resolves. The README template
nobody referenced is wired into the step that writes the knowledge-base README.

Frontmatter. None of the three skills declared allowed-tools, so the first
command of every flow hit a permission prompt. The wiki skill kept its trigger
words and prerequisites inside the description text; both move into the body.

Splits. `core` keeps what a daily user needs and `setup` takes day 0 and the
repo lifecycle, so the common path no longer carries ~500 lines of repo
creation. The wiki skill's phase procedures move into references/phases/, taking
its SKILL.md from 38.7 KB — larger than agent-browser's entire core — to 17 KB
with an index that says when to load each phase.

One contradiction is resolved in the author's text: the share skill mandated
that every generated document be written in Chinese, against global rule 1
("reply in the user's language") and this repo's own English rule. It now
follows rule 1.

Refs #678

* feat(pull): prune legacy builtin skill directories, gate recall at run time

Upgrading the CLI used to leave the pre-stub trees in place: cleanup skips
builtin names, and nothing else knew about them, so `team-wiki-codebase` and
`teamai-share-learnings` would sit in every agent directory on the machine
forever. Deployment now removes them first, in both the configured skills path
and Codex's shared `.agents/skills`. Unconditional, because those trees were
overwritten on every pull, so no local edit ever survived in them.

Recall moves from deploy time to run time. Before, `skipRecall` decided whether
the share skill reached the agent at all; with one stub routing to everything,
there is no directory to withhold, so `teamai skill get share` checks instead
and says what to enable. `--all` is exempt: an inventory dump is not an attempt
to run the workflow. With no team config to consult the gate fails open — a
fresh machine reading the docs gets the content rather than a refusal it cannot
act on.

deployBuiltinSkills drops its `skipRecall` option rather than keeping one that
no longer decides anything, and recall-toggle stops deleting a skill directory
it no longer owns.

Refs #678

* docs: align the nudge and the guides with CLI-served skills

`/teamai-share-learnings` was never a slash command of its own — it existed
because the directory was installed. The Stop-hook nudge now names `/teamai` and
carries `teamai skill get share` literally, so an agent can act on it without
having to infer the intent from the conversation. The five READMEs and both
usage guides follow.

Both guides gain the `skill get` / `skill path` commands and a short section on
why built-in skills are served rather than copied. `docs/designs/skill-serving.md`
records the contracts that are easy to break later: byte-for-byte output,
{SKILL_DIR} substitution, stdout/stderr discipline, recursive `--full`, the
run-time recall gate, the three drift guards, and when to retire
LEGACY_BUILTIN_SKILL_NAMES and the long-name aliases.

AGENTS.md and CLAUDE.md gain the rule that keeps this from rotting: skill-data
is treated like documentation, a behaviour change updates the affected skill,
commands.md is regenerated rather than edited, and new workflows go under
skill-data instead of into the stub.

Refs #678

* fix(skills): apply standards review findings

`teamai init` printed "Built-in skills (e.g. team-wiki-codebase) are ready to
use in your IDE now" seconds after deployment deleted that very directory. The
message now names the teamai skill and how it loads its workflows. Two comments
carrying the same stale name follow.

The five READMEs said different things: only the English one named the share
workflow and its command. All five now do.

`collectSupplementaryFiles` hand-rolled a recursive walk that
`listFilesRecursive` already does, including the ignore list that skips `.pyc`
and `__pycache__` next to the wiki's Python scripts. It calls the helper
instead. `listServableSkills` drops its fallback to `skills/`: a package without
`skill-data/` is broken, and serving the stub as if it were the content hides
that from the one error message built to report it.

Tests drop six non-null assertions for a helper that throws, per the repo's rule
against moving a compile-time error to run time.

AGENTS.md and CLAUDE.md record the exemption the branch created: skill content
printed by `skill get` keeps the language its author wrote it in, while the
command's own prompts, errors and listings stay English.

Refs #678

* fix(skills): apply spec review findings

Upgrading left the old references in place. Releases before the stub deployed
`skills/teamai/` with six reference files beside SKILL.md, and `teamai` is not a
legacy name to prune, so copying one file over that directory kept ~39 KB of
pre-stub instructions next to the new stub for good. Deployment now clears
everything the deployed unit does not contain before writing it, and a test
seeds the old layout to prove it.

Pruning reached neither reporting-only teams nor the Codex shared directory in
any test. The prune now runs before the reporting-only return, so a team that
switched to reporting-only still loses the stale trees, and the Codex
`.agents/skills` path is covered by a test. Excluded agents stay untouched, as
the enabledAgents whitelist documents.

The served content still routed to skills that no longer exist: five mentions of
`teamai-share-learnings` and one `/team-wiki-codebase --update`, which is the
rule this branch itself added being broken on arrival. One second-hop path
inside a sub-agent input packet was still relative.

The share skill's document template, frontmatter table and tag taxonomy move to
`references/doc-template.md`, taking the always-read body from 3 701 to 2 471
bytes. Two tests now assert what nothing guarded: every served skill's
frontmatter name matches its directory and declares allowed-tools.

A new e2e file runs the built CLI the way an agent does: every listed skill is
servable and byte-identical bar the resolved placeholder, the wiki scripts run
from the directory `skill path` prints, an unknown name exits 1 with empty
stdout, a hallucinated flag warns and still serves, and `--full` appends the
nested references in sorted order.

Refs #678

* fix(skills): prune only directories the CLI owned, gate every content path on recall

- LEGACY_BUILTIN_SKILL_NAMES drops teamai-workflow and teamai-import: they were
  reserved in the old guard set but never packaged, so a directory by either
  name is the user's own skill. Test: user-created skills with those names
  survive pull.
- The recall gate now covers skill get --all (blocked skill skipped, named on
  stderr), skill path (refused) and skill list --json (blockedByRecall, path
  null). skill list reads the flag from the catalog instead of re-checking.
- skill get [names...]: the positional is optional so --all is reachable from
  the real CLI; Commander used to fail with 'missing required argument'.
  Covered by the skill-serving e2e.
- commands-reference renders Commander's variadic marker (<names...>); the
  snapshot is regenerated.

* test(skills): drive the recall gate through the real CLI, guard the stub description budget

- skill-serving e2e: a HOME with a team whose recall is off; skill get share,
  --all, skill path share and skill list --json each withhold share, and all
  serve it after recall enable. The earlier HOME has no team config and fails
  open, so the gate was never exercised through dist/index.js.
- skill-content test: the stub description stays within 1024 characters.
- skill-commands-exist also scans the deployed stub.
- Docs and PR lead with content versioned with the CLI; the size numbers are
  measured (stub description 0.8 KB, body 1.3 KB; --full 32/36/115 KB).
- Content audit against origin/main: every file has a counterpart. Fixes:
  {SKILL_DIR} defined in core/setup/share where the references are listed, the
  wiki overview draws the served layout, team-wiki-codebase kept as a trigger
  word in the stub description.

* fix(skills): gate skill show on recall, prune Codex's shared dir only from Codex

Review follow-up on #699.

- `skill show <served skill>` refuses a recall-blocked skill with the same
  message and exit code as `skill get` / `skill path`; it printed the
  directory those two withhold.
- A skill resolved from skill-data/ is classified `[builtin]` directly.
  BUILTIN_SKILL_NAMES only knows the deployed stub, so `skill show core`
  reported `[local-only]` beside a package path.
- pruneLegacyBuiltinSkills reaches `.agents/skills` only on Codex's own
  pass. Another enabled tool's pass deleted Codex's legacy copies while
  Codex was excluded, against the enabledAgents guarantee.
- The share skill and its references are written in English; the generated
  document still follows the session's language. The AGENTS.md exception
  for Chinese skill-data output is dropped.

* fix(skills): one resolver for served skills, legacy names kept out of push, wiki in English

Review follow-up on #699.

- resolveServableSkill is the only way to obtain a PackagedSkill outside
  skill-content.ts; it returns `blocked` instead of the skill, so `get`,
  `path`, `list` and `show` inherit the recall gate by construction.
- push never offers `team-wiki-codebase` / `teamai-share-learnings` as new
  user skills: between the upgrade and the first pull they are still on
  disk (isCliOwnedSkillName).
- `recall disable` removes the legacy `teamai-share-learnings` directory
  again (LEGACY_RECALL_SKILL_NAMES), skipping excluded agents.
- `skill list` prints the packaged catalog before `teamai init`, with a
  hint for the team half, instead of failing on the team listing.
- skill-data/wiki (SKILL.md, 14 references, 2 scripts) translated to
  English. Generated document names follow one glossary; validate_kb.py
  still recognises headings of knowledge bases built by the previous
  release, matched by code point so the source stays ASCII.

* fix(skills): prune only files the CLI packaged, let local skills win by name

Review follow-up on #699.

- PACKAGED_SKILL_FILES lists every file a release ever wrote under skills/,
  as the union of `git ls-tree -r <tag> -- skills/` over all 91 tags. The
  prune removes those paths and the directories they leave empty; a file a
  member added is kept, its directory with it, and pull says which and why.
  The stub directory loses its six known references by name instead of
  "everything that is not SKILL.md". Python bytecode of a script we shipped
  counts as ours, so a __pycache__ does not strand the tree.
- locateSkill searches the team repo, then installed agents, then the
  package. A directory a member created under `codebase`, `default`,
  `learning` or `share` is the skill they asked about, and the recall gate
  does not apply to it.
- A guard test fails when a file ships under skills/ without being recorded
  in the manifest, which a later migration would otherwise leave behind.

* fix(skills): close the last recall bypass, uninstall Codex's shared stub

Review follow-up on #699.

- `skill path` takes a name, always. The argument-less form printed the
  `skill-data/` root, and `<root>/share/SKILL.md` is readable from there —
  the content the gate withholds one command over.
- uninstall discovers skills in Codex's shared `.agents/skills` root, where
  resolveSkillDestination puts the stub whenever the skill already lives
  there. Without it, uninstall reported success and left it behind. Codex
  only, as the legacy prune already does.
- core/SKILL.md said team sharing is enabled by default; getRecallSharing
  defaults it to false. It now says recall is off by default and names
  `teamai recall enable`.

* fix(skills): uninstall by the same ownership rule as pull, quote {SKILL_DIR}

Review follow-up on #699.

- uninstall removed a CLI-owned skill directory whole, undoing one command
  over the guarantee pull makes. It now removes the PACKAGED_SKILL_FILES
  paths through the same removeOwnedFiles, keeps a directory holding a file
  the member added, says which one, and tells the confirmation prompt so it
  no longer promises a directory it will keep. A team-repo skill is synced
  whole and still goes whole.
- Served shell commands quote the placeholder: `python3 "{SKILL_DIR}/..."`.
  Unquoted, an install path with a space ("Program Files", "Application
  Support", a Windows path through Bash) splits into two arguments and the
  documented invocation fails. A test fails on an unquoted occurrence after
  any command word, in SKILL.md or any reference.
- wiki/references/overview.md said the methodology, scripts and agent specs
  are deployed into agent directories. They are not: only the stub is, and
  the rest is served from the installed CLI.

* fix(skills): deploy the stub in reporting-only mode, drop the stale list alias

Review follow-up on #699.

- Reporting-only HTTP pull pruned the legacy trees and deployed nothing, so
  a member on an HTTP team came out of the upgrade with no built-in entry
  point at all. The skip predates CLI-served content: it existed because the
  only deployable unit then needed a team repo. The stub does not — its
  workflows are printed by the installed binary, and `skill get wiki` is a
  local knowledge-base generator that never touches a repo. The stub now
  deploys in every mode, and `reportingOnly` goes with the branch it gated:
  nothing else read it.
- `teamai skill list` called itself an alias for `teamai list skills
  --source all`. It has not been one since it started printing the CLI-served
  catalog underneath. Both descriptions, the generated command reference and
  both usage guides now say what it does.

Refs #678

* fix(skills): carry the TGit provider guide into the served setup skill

#724 landed `skills/teamai/references/provider-tgit.md` and repointed
setup-admin.md and join-member.md at it. Rebasing onto that left the new
file in a tree this branch no longer deploys, and the pointers in bare
`provider-tgit.md` form the served skills do not use.

- move it to `skill-data/setup/references/`, beside the two files that
  cite it, so `teamai skill get setup --full` serves it
- rewrite every pointer to it as `{SKILL_DIR}/references/provider-tgit.md`
- list it in the setup skill's reference table
- add `references/provider-tgit.md` to PACKAGED_SKILL_FILES, so the prune
  removes it from members who pulled a release that shipped it

* fix(skills): make the blocked catalog entry unrepresentable, drop unsafe casts

Review findings from the standards axis, plus the doc half of the prune
count.

- `SkillCatalogEntry` allowed `{blockedByRecall: true, path: '/…'}`, an
  invariant `skillCatalog` then upheld by hand. Split it on
  `blockedByRecall`, so the withheld directory is a type error rather than
  a review catch. Both variants keep the `path` key, so the
  `skill list --json` shape is unchanged.
- `command.commands as Command[]` stripped commander's `readonly` in three
  places. `for…of` and `.find` need no cast.
- `docs/designs/skill-serving.md` still said the prune removes six
  `teamai/references/*.md`; provider-tgit.md makes it seven.

* fix(skills): keep publishing a skill reachable when recall is off

Publishing a skill is `teamai push --skill`, which never consulted recall
(`src/push.ts` names it nowhere). On main the flow shipped in the teamai
skill, ungated. Moving `contribute-member.md` under `share` put it behind
the recall gate, so with recall off — a new team's default — the core
routing table sent the agent to `teamai skill get share`, which exits 1
and tells it to enable recall. Wrong advice for a flow recall does not
touch, and no other path to the instructions.

Move the file to `core`, the skill that already owns `push`, and split the
routing row so publishing and session learnings stop sharing one
destination. The gate itself is right and stays: learnings do need recall.

`share/SKILL.md` already called this "a different flow"; now it points at
`teamai skill get core --full` instead of at its own references.

PACKAGED_SKILL_FILES is unchanged: the legacy path a pre-stub release
wrote is still `teamai/references/contribute-member.md`.

* fix(skills): back up what the prune removes, so no edit is a one-way door

Review finding: `removeOwnedFiles` proves ownership by pathname and
deletes without reading the file, so a member's edit goes with it.

For a path the current package still ships that changes nothing: the old
deployment overwrote it with `overwrite: true` on the same three triggers,
so the edit died either way, at the same moment. The case the objection
gets right is a path a retired release shipped and the package no longer
does — the overwrite never reached it, so the edit did survive, and the
prune is the first thing to remove it.

Copy every pruned file to `~/.teamai/removed-skills/<date>/<tool>/<skill>/`
before removing it. Outside every agent directory, so nothing reads it back
as a skill.

Verifying contents against a hash of each released version was the other
way out, and it is worse: anything not byte-identical is then kept, so one
CRLF checkout on Windows — a platform this project supports — leaves the
whole 176 KB in place and reports success. Backing up gives the same
guarantee without betting the migration on byte equality.

Uninstall keeps deleting outright: there the member asked for the files to
go.

* fix(skills): let no backup failure authorise a delete, give each root its own

Two holes in the backup the previous commit added, both reported in review.

The copy's failure was swallowed at debug level and the delete went ahead
regardless, so a full disk or a read-only home turned the migration back
into the data loss the backup exists to prevent — and the log still named
a backup directory that held nothing. A file whose copy fails is now kept,
counted, and named at warn level; `removeOwnedFiles` returns what happened
instead of a bare boolean, and only a run that copied something names the
directory.

The backup path was `<date>/<tool>/<skill>` with `overwrite: true`, so the
second copy of a name silently replaced the first. Codex prunes the same
skill from `.codex/skills` and the shared `.agents/skills`, and two pulls
share a date. The path now carries a per-run id and the skill root, and the
copy refuses to overwrite rather than clobbering a copy it cannot replace.

Tests cover both: a file where the backup tree must start makes every copy
fail, and the two Codex roots land in separate directories. Each fails
against the previous commit.

* fix(skills): stop at a symlinked root, archive only what is retired

Three review findings, all in the prune.

A symlinked skill directory was walked through. `readdir` follows the link,
every path under it matches a packaged name, and the delete lands in
someone else's checkout. Ownership now stops at the link: the root is
lstat'd, a symlink is refused, and link and target are left alone.

The stub directory was pruned against the full historical file list, which
includes the SKILL.md written one line later. Deployment runs on every
session start, unchanged revision included, so that archived an identical
copy per session forever. Only paths this release no longer ships are
archived now.

Backups were written under the tool's base directory, which under project
scope is the repo root, so they landed in the working tree outside the
generated .teamai/.gitignore. They go to the machine's home.

Also: `skill show <packaged>` resolved the team before the package, so it
failed on a machine that never ran `teamai init` for content that needs no
team. Packaged names resolve first and print without the team-dependent
fields.

* fix(skills): stop at the first link above a skill dir, report a half prune

Findings from a self-review run before pushing, plus the two from the last
review round.

The symlink guard was one level too low. It lstat'd the skill directory, so
the common shape — `~/.claude/skills` itself linked at a dotfiles checkout —
walked straight through: every directory under the link is real. The guard
now walks each component below the tool's base directory and stops at the
first link, which covers the prune and the stub write with one check.
Components at or above the base are not checked: a home directory under a
link is ordinary, and refusing there would disable deployment on those
machines.

The symlink branch borrowed the foreign-files message, so a member was told
"delete the rest yourself" about a directory nothing had touched. Following
that destroys what the guard just protected. It has its own sentence now, in
pull and in uninstall.

`remove()` was not fail-closed the way the backup is: a read-only parent
left the tree half-pruned under a debug line, and a `walkFiles` that threw
returned success. Both are recorded in `notRemoved` and reported.

The backup path gained the base directory: `inheritUserScope` deploys the
user base and then the project base in one process, same tool, same root,
same skill name, and `errorOnExist` turned that collision into files the
second pass could neither archive nor prune.

Docs corrected against the code: the archive path, the tag count (98, not
91, and `teamai-wiki` is excluded), the version line, and the size table.

* fix(skills): route the nudge and skill publishing where they land, classify legacy names as ours

The Stop-hook hint said "run /teamai", but bare /teamai prints the menu and
stops, so following the primary suggestion never reached the share workflow.
It now names an invocation the core skill routes to share, with the
`teamai skill get share` fallback kept. The four docs that quote the hint follow.

The setup skill sent "publish one skill" to `teamai skill get share`, which
handles session learnings and is refused when recall is off (the default);
reusable-skill publishing lives in core's contribute-member reference and needs
no recall. The routing row and the two references that repeated it now point
there.

classifySkill checked BUILTIN_SKILL_NAMES alone, so until the first pull pruned
them, team-wiki-codebase and teamai-share-learnings showed as [local-only]. It
now uses isCliOwnedSkillName, the rule push and uninstall already apply.

* fix(skills): pre-push review — gate the nudge on recall, keep bytecode out of the tarball, report a failed uninstall delete

A review of the whole branch against #678, #730 and the design doc, run before
pushing. What it found and what changed:

- The Stop-hook share reminder was gated on the hint switch alone; recall is off
  by default and `teamai skill get share` refuses then, so the reminder pointed
  at a command that said no. It is withheld while recall is off, the same gate
  the workflow has; the served text about when the prompt appears now matches.
- `npm pack` swept `skill-data/wiki/scripts/__pycache__` into the tarball once
  the e2e suite had run the scripts. Excluded in package.json "files", asserted
  absent in the tarball test, and the e2e run sets PYTHONDONTWRITEBYTECODE.
- The `share` description still offered to publish reusable skills, the flow its
  own body sends to `core`; the sentence is gone.
- `skill show <unknown>` before `teamai init` threw the init error as a stack
  trace; it prints the not-found line and exits 1.
- The stub pre-approved every `teamai` command from the always-loaded unit;
  narrowed to `Bash(teamai skill:*)`, which is all it asks for (#678).
- Six routing lines loaded `core --full` to reach one reference; they name the
  file under `$(teamai skill path core)/references/` instead.
- `uninstall` reported a failed delete as "holds files TeamAI did not put there;
  the packaged files were removed", both false and the error unprinted. It names
  the file and the error; a test makes the stub directory read-only.
- The stub directory archived under `<tool>/.claude-skills-teamai/teamai/` while
  the legacy trees used `<tool>/.claude-skills/<skill>/`; one layout now.
- CHANGELOG entry; dead `isRecallEnabled` import; wiki heading still naming
  `team-wiki-codebase`; JSDoc on the wrong declaration; stale byte counts; the
  usage guides gain the recall refusal and the archive location; the design doc
  records the `--json` deviation, the legacy-name classification rule, the
  uninstall symlink scope and the fail-open wording.

* fix(skills): English-only served content, one link guard for every caller, withhold share from read-only sources

The reviewer flagged Chinese in skills/ and skill-data/ a third time. Both
reach the agent as CLI output, so the stub's trigger keywords, the paired
sample invocations and the Chinese name for TGit go; the agent translates for the user. A test
fails on CJK anywhere under either root.

Pre-push review of the whole branch, and what changed:

- uninstall walked through a linked ~/.claude/skills and deleted the
  packaged files inside the member's dotfiles checkout; pull refused the same
  layout. removeOwnedFiles now owns the guard, so pull, deploy and uninstall
  apply one check: the skills root and the skill directory. A linked
  ~/.claude (stow, chezmoi) is no longer refused, since every other resource
  writes through it and refusing left those machines on the pre-stub trees.
- share was served to read-only HTTP teams, where its last step
  (teamai contribute) always fails; reportingOnly used to skip it. The
  serving gate carries a reason (recall | read-only) with its own message,
  and `skill list --json` reports it as `blockedBy`.
- The bytecode rule claimed any file under any __pycache__; it now claims
  only the .pyc of a shipped script.
- recall disable pruned the shared .agents/skills root for an uninstalled
  Codex; it has deployment's install gate now.
- The source-team guard lost the legacy names when BUILTIN_SKILL_NAMES
  narrowed, so a source removal could delete a legacy tree wholesale.
- Routing: the admin wrap-up and the stub still sent "share what I learned"
  to share without saying it needs recall, and the stub filed "share this
  with my team" (the publish-a-skill phrase) under share. recall enable is
  described as the per-machine override it is, next to the team key.
- {SKILL_DIR} definitions now say how a reference file opened on its own
  spells the directory, since serving resolves the definition too.
- skill show: packaged resolve only when init fails, aligned label, a
  served skill is "served by the CLI, not installed".
- Docs: uninstall removes the archive with ~/.teamai; zh said the whole
  directory is kept; the product overview lacked the recall gate; the design
  doc's release, tag and byte figures were stale; CHANGELOG notes the
  language change of generated documents.

* fix(skills): check every path component below the base for a link, in uninstall too

The previous commit narrowed the guard to the skills root and the skill
directory, so a link at ~/.config or ~/.config/opencode was walked through:
the prune could delete, and deploy write, inside a dotfiles checkout. The
full walk from the tool's base directory is back, and removeOwnedFiles now
requires the base, so uninstall applies it too; each skill directory in the
uninstall plan carries the base its skills root hangs off.

A member whose whole ~/.claude is a link keeps the pre-stub trees and gets
the warning naming the path, as before the previous commit. Deleting through
a link is the one thing the prune must never do.

* fix(skills): withhold the share hint on read-only sources, route legacy names through the gate

contributeHintAllowed checked recall only. The dispatcher already drops this
gitOnly handler for HTTP teams, but the gate now says so itself, so the
reminder never points at a `share` that refuses as read-only wherever it runs.

`skill show teamai-share-learnings` searched the agent directories before
the package, so a legacy tree a pull had not pruned yet was shown with its
path while the gate refused `share`. A legacy built-in name now skips the
agent search and goes to the packaged skill and its gate; ordinary names and
aliases such as `share` keep a member's own directory first.

* fix(skills): deploy and prune built-ins where the tool keeps its skills

deployBuiltinSkills joined baseDir with the configured skills path, while
team-skill sync resolves the directory through skillsDirForTool: OpenClaw's
workspace, and HERMES_HOME for Hermes. Those agents got the stub in a
directory they never read and had their legacy trees pruned from the wrong
place. Deploy, the legacy prune, recall disable and uninstall now resolve
the same directory; the link guard starts at the tool's base directory when
the skills directory sits under it, else at that directory's parent.

* fix(skills): check an external skills root for a link, keep config-load logs off stdout, drop inert allowed-tools

- A skills directory outside the tool's base (HERMES_HOME, an OpenClaw
  workspace) had the guard start at the root itself, so a linked root was
  never checked. It starts one level above now, and a linked HERMES_HOME is
  refused like a linked ~/.claude.
- The share gate loads the config, which can migrate it and report that with
  log.info on stdout: an upgrading machine got that line in `skill get`
  output and in `skill list --json`. Config loading reports on stderr for
  that call; setStderrOnly returns the previous mode so it can be restored.
- `allowed-tools` in the served skills was printed as command output and
  never processed as skill metadata, so it granted nothing. Removed, and the
  test now fails if one comes back. Only the stub's line pre-approves.

* fix(skills): walk the link guard from the scope root, so a linked COPILOT_HOME is refused

skillsGuardBase started at the tool's base directory, which for Copilot in
user scope is COPILOT_HOME, so the walk never checked whether COPILOT_HOME
itself was a link, and pull and uninstall wrote and pruned through it. The
guard now starts at the scope root (home, or the project root), where a link
at or above is ordinary, in deploy, the legacy prune and uninstall alike; a
root configured outside it still has the walk start just above that root.

* fix(skills): keep generated documents in Simplified Chinese, fail on a broken config, quote skill paths

- share and wiki had moved generated learnings and knowledge-base documents
  from always Chinese to the session language. Serving the instructions from
  the CLI does not need that, so both say "Simplified Chinese" again, in an
  English instruction; the CHANGELOG entry follows.
- skill show and skill list treated every autoDetectInit failure as "not
  initialized" and pointed at `teamai init`. requireInit now throws a tagged
  NotInitializedError; only that falls back to the packaged catalog, and a
  malformed or unreadable config propagates.
- `$(teamai skill path …)/…` is word-split in a shell command like an
  unquoted {SKILL_DIR}; all ten occurrences are double-quoted and the quoting
  test covers the form.

* fix(config): raise NotInitializedError only when the config file is missing

loadLocalConfig returns null both for a missing file and for one that fails
to parse, validate or migrate (it logs the reason). requireInit turned every
null into NotInitializedError, so skill show and skill list still fell back
to the packaged catalog and a `teamai init` hint on a broken config. Only an
absent file is NotInitializedError now; an existing one that could not be
used is an error naming its path, in requireInit and the user branch of
requireInitForScope. Covered through the real loader and the built binary.

* fix(skills): prove ownership by content, not path alone; retire the second Codex copy

- The legacy prune and uninstall removed any file at a path a release had
  packaged, so a member's edit, a skill of their own under an old name, or a
  root TeamAI never managed (toolPaths or HERMES_HOME moved) lost its files.
  A file is ours now only at a packaged path and with content a release
  shipped there: PACKAGED_SKILL_DIGESTS records the sha256 of every blob over
  all 99 tags through v0.25.0 and main before the stub, 37 versions across 21
  paths. A skill-root SKILL.md is compared by its body, since releases before
  0.17 shipped no frontmatter and the deploy of the day repaired it on disk.
  The current stub is ours by the packaged copy. Anything else stays.
- Codex reads .codex/skills and the shared .agents/skills, and the stub goes
  to the shared one when a copy lives there; the copy an earlier release left
  in the other root kept its old SKILL.md and references. It is retired by
  the same ownership rule, archived first, and named when kept.
- Tests mock the digest table with a stand-in for shipped content, and a test
  keeps the stand-in on the same paths as the real table.

* chore(skills): carry main's skill edits into the served copies after the rebase

#713 and #736 edited skills/teamai/references/*.md, which this branch moved to
skill-data/setup/references/. Two hunks did not follow the move:
- join-member.md: TGIT_TOKEN is REST-API-only and cannot clone (#713).
- setup-admin.md: the /teamai share entry publishes a reusable skill; a
  session's learnings are automatic (#736), in English as the served text is.
#739's partial config mock is restored in skip-uninstalled-tools.test.ts.

* fix(skills): deploy before pruning, block share on an unloadable config, drop hidden commands from the reference

- Legacy trees were pruned before the stub was written, so a refused or
  failed stub (a link, a read-only directory) left the agent with nothing
  to discover. They go only once the stub deployed for that agent.
- The share gate failed open on any config error. Only a machine with no
  config (NotInitializedError) is served; a config that exists but cannot be
  loaded blocks with its own reason, `blockedBy: "config"`.
- The KB template told agents to run `code-to-knowledge --update`, which
  does not exist; it names `teamai codebase --extract … --incremental`.
- The generated command reference listed hidden hook plumbing (`track`,
  `contribute-check`, `todowrite-hint`, …). It renders what `--help` lists.
- removeEmptyDirs swallowed every rmdir error, so a directory that stayed
  could be reported removed. Only "still holds something" is expected; any
  other failure is reported.

* fix(skills): whole-file ownership, stub before its references, no side effects before the link guard

Review of 327f9cd:
- SKILL.md was compared by its body, so a member who changed only its
  frontmatter lost the file. Every release from 0.16.1 (the first whose
  deploy repaired frontmatter) shipped complete frontmatter, so what is on
  disk is what was shipped: digests are whole files now (42 versions over
  100 tags and main). A link is never ours; bytecode is ours only beside a
  script proven ours by content, decided before anything is removed.
- The stub dir's retired references were pruned before SKILL.md was copied;
  a failed copy left the old skill pointing at files that were gone. The
  stub is written first.
- The Codex destination was resolved with the reconciliation that deletes a
  duplicate, before the link guard ran. It is resolved side-effect free; the
  other copy is handled under the guard by retireOtherCodexCopy, whose
  report now names a failed backup or delete as such.
- A broken project config was skipped by detection, so the share gate
  answered with the user config. findUnreadableProjectConfig reports it via
  an optional sink on detection (no caller changes), and the gate blocks.
  The Stop-hook reminder is withheld on an unloadable config too.
- init announced the stub as ready when nothing was deployed; hook-dispatch
  is hidden (hook plumbing), and the reference says it lists public commands;
  the design doc no longer says teamai-workflow/teamai-import are removed.

* fix(config): report a broken higher-priority project config even when a fallback loads

findUnreadableProjectConfig dropped a recorded error whenever detection
went on to find a later candidate: a broken partition config followed by a
valid legacy .teamai/ config returned null, and the share gate answered with
the fallback's team. It now reports the first unreadable file regardless.
An existing config file that is empty or cannot be read is reported to the
sink too, instead of returning without a word.
2026-09-23 16:15:38 +08:00
Saul Moro ff47714902 fix(doctor): probe the Copilot hooks file where inject writes it (#732) (#733)
In a non-self project scope, resolveDoctorContext forces the hook paths to
the hook scope ('user', per resolveHookScope) so settings-based hooks are
probed where reconcileHooksToAllTools writes them. The standalone Copilot
hooks file is written by reconcileTeamHooksForConfig at the config's own
scope instead, so the doctor ended up joining the userScope relative path
(hooks/teamai.json) onto <projectRoot> and reported Copilot missing right
after a successful `hooks inject`.

buildHookChecks now takes both maps and picks per hook kind: a standalone
`hooks` file from the config-scoped paths, `settings` from the hook-scoped
ones. Same rule `hooks list` already applies.

Hypothesis confirmed: scope mismatch between the two path maps, introduced
when #695 moved the doctor's hook paths to the hook scope for Qoder CN.
2026-09-23 15:22:00 +08:00
Saul MoroandSaul Moro d80d5a8678 fix(push): namespace new rules and agents from --role/--project (#649) (#698)
* fix(push): namespace new rules and agents from --role/--project (#649)

`--role`/`--project` only ever placed new skills, so a rule pushed with
`--project front-app` landed at `rules/<name>.md` and a new agent at
`agents/<name>.yaml` — both of which `pull` ships to every member. The
flag also collapsed into the project's `skills` namespace, which is the
wrong directory for a rule: a rule is namespaced on the `knowledge` axis,
and the manifest allows the two to differ.

Each pushable type now resolves from its own axis (skills → `skills`,
rules → `knowledge`, agents → `agents`), and the destination is printed
rather than chosen silently. Where the named project declares no
namespace for a type being pushed, the command fails and names it
instead of writing to the shared root.

Only new resources already at the shared root are placed; anything the
scanner namespaced keeps its path (#654), and an open PR's recorded
destination still wins so a force-push never moves a resource.

Also resolves a root-level local rule against its namespaced team copy,
so a rule that was placed on an earlier push is not re-pushed to the
shared root once it merges.

* fix(push): record rule placement in state instead of matching by basename

RulesHandler.scanLocalForPush matched a root-level local rule against the
sole active rules/<ns>/<name>.md by basename. A namespaced team rule is
pulled into a namespaced local directory, so another member's unrelated
root rule with the same name would have been read as a modification of the
team rule and overwritten it under --all.

push now records where it placed each root-level rule (state.placedRules,
name -> team path). The scanner redirects a root-level local rule only when
that record exists and its team file is still present; otherwise the rule
is new. The ambiguity warning goes with the basename index.

Review: https://github.com/Tencent/teamai-cli/pull/698#issuecomment-5770251793

* fix(push): stop on unreadable roles manifest, sync placed rules, resolve in dry-run

Three review findings on top of the #649 placement fix.

A roles manifest that exists but cannot answer — unparseable, or missing the
configured role — no longer falls back to an empty namespace list, which sent a
new rule or agent to the shared root and therefore to the whole team. Only an
absent manifest keeps the pre-manifest fallback, so `loadRolesManifest` now
throws a tagged `RolesManifestNotFoundError` to tell the two apart.

`syncTeamUpdatesToLocal` follows the same `placedRules` record the scanner does,
so a root-authored rule placed under `rules/<ns>/` takes part in the three-way
sync. Without it a teammate's newer version was never synced down and the stale
root copy was pushed over it.

`--dry-run` now runs the recorded-destination and placement steps before it
exits, so it reports where every new resource goes and fails on the same
unresolvable project axis the real command refuses.

* test(push): cover #649 placement with the real CLI across agents and providers

Drives the built dist/index.js against real git remotes, a fake `gh` and a fake
GitLab API, and asserts on the branch content that reached the remote.

The four agents × three providers cover the placement itself; the remaining
cases cover what review round 2 raised — the pre-push sync following
placedRules, an unreadable roles manifest stopping the push, and --dry-run
resolving the same destinations. Reverting any of those three fixes turns
exactly its case red and leaves the rest green.

* fix(push): keep a placed resource maintainable and removable by its author

Three review findings on the placement this PR added.

A placement record only ever meant "push put this here", and the two sides that
read one disagreed about how much it was worth. `placedResourcePath` is now the
single resolver: it validates the record (inside the resource root, namespaced,
no traversal, named after the resource) and both the push scanner and the
pre-push sync go through it, so they cannot drift apart again. The record also
takes precedence over a shared-root file that appears later with the same
basename — mapping the author's copy onto somebody else's rule would push their
content over it.

Agents gained the analogue, `placedAgents`. `AgentsHandler.scanLocalForPush`
only accepts a team source whose namespace is ACTIVE here, so an agent
published with --role/--project into a namespace this directory never
activated was skipped as "no active source" on the author's very next edit:
they could create the agent and then never maintain it.

`teamai remove rules <name>` resolves the same record through the new
`publishedNameFor` hook. The author's copy stays at the rules root, so the name
they type is the bare one, and remove answered "not found" about a rule it had
recorded publishing. It now reports which name it resolved to, deletes the
namespaced team file, and takes the author's root copy with it — left behind,
that copy re-publishes the rule on the next push.

* fix(push): narrow what a placement record grants, and when it is written

Four review findings, each about the record rather than the placement.

`remove` consulted it only after a bare-name match failed, but the LOCAL scan
contributes the bare name whenever the author's own copy has edits — so
`remove rules my-rule` deleted that copy, reported success, and left
`rules/<ns>/my-rule.md` published. The record is now resolved first.

A record is written only for a resource push actually placed: `new`, and
namespaced by this run. Recording a `modified` agent meant a namespace that
happened to be active at edit time became standing permission to keep editing
that agent long after the role or project granting it was dropped.

Records are persisted per group, right after that group reaches the remote,
instead of after every group completes. A failing later group returned early
and took the earlier group's mapping with it, so a resource that WAS pushed
came back misclassified once its PR merged.

And the roles manifest is held to the same rule as `--role` and the projects
manifest: a namespace is one path segment. `foo/bar` wrote an agent below the
depth pull looks at, and read back as namespace `foo` for a rule.

* fix(push): let --role/--project decide which team agent a local edit belongs to

`AgentsHandler.scanLocalForPush` picked the team file to edit by activity
alone, and it runs before the destination is resolved. So `agents/other-ns/vr.yaml`
— an agent this directory never activates — made `push --project front-app`
report "no active source" and push nothing, even though the same stem is
allowed to exist in several namespaces and the flag had named a different one.
The scan now takes the requested namespace, through a new optional
`ScanForPushOptions`; with one named, sources in other namespaces are other
agents, and an absent one means this agent is new there. A shared-root copy
still blocks, and now says why: both would be active at once, which is the
collision pull reports.

`RulesHandler.removeItem` swept the bare basename unconditionally, so
`remove rules fe/foo` deleted an unrelated personal .claude/rules/foo.md. The
bare copy is only ours to delete when this machine's placement record says the
two are the same rule.

`teamai push --help` said both flags target skills.

* fix(push): never place a new resource onto one that is already there

Placement rewrote a new root-level resource to the resolved namespace without
looking at what was at that path. An unrelated local `foo.md` — which the
scanner rightly calls new, since no record maps it anywhere — landed on
`rules/<ns>/foo.md` and replaced somebody else's rule, silently, in a run they
never reviewed. Push now stops and names the file. The same guard covers the
`--role`/`--project` skills override, for new skills only: a modified one is
meant to land on its own directory.

`loadRolesManifest` read through `readFileSafe`, which answers null for every
failure, so a manifest that exists but cannot be read arrived looking exactly
like a missing one — and a missing one is the pre-manifest layout, which sends
new rules and agents to the shared root. The two are told apart now.

`RulesHandler.removeItem` tombstoned only the name it was given. Removing
through a placement record means the author's source is named `<name>` while
the published file is `<ns>/<name>`, and the local sweep skips excluded tools,
so a root copy could outlive the removal there and come back on the next push.
Both names are tombstoned when the record vouches for the bare one.

* fix(push): resolve a placed agent on removal, and collide on either extension

`remove` asked every handler for the published name, but only rules answered.
So `teamai remove agents vr` matched the bare stem and deleted every `vr` in
every namespace — other people's agents included — while the namespaced
`placedAgents` record, keyed by a path the removal never named, survived.
`AgentsHandler.publishedNameFor` resolves it now; the existing sweep already
narrows a `<ns>/<stem>` to one file, since the root directory is one of the
directories it probes. The bare stem is tombstoned alongside the published one
and swept from the tool directories, the same way rules are.

The placement collision check tested the proposed path alone. `pull` reads a
legacy `<stem>.md` as the same agent as `<stem>.yaml`, so a new `.md` landing
beside an existing `.yaml` passed the check and left two copies answering to
one name. Agents are now checked under both canonical extensions.

* fix(remove): make the placed-agent resolution actually reach the command

Round 7 added `AgentsHandler.publishedNameFor` but `remove` only used its
answer when `allNames` also carried that spelling — and `scanTeamForPull`
reports an agent by its bare stem, never `<ns>/<stem>`. So the resolution was
inert on the real command path: `teamai remove agents vr` fell back to the bare
stem and deleted every `vr` in every namespace, exactly as before. The cross
check is gone; `publishedNameFor` has already proved the file is in the team
repo, which is stronger evidence than membership in a list the scans spell
differently per type.

The bare-stem tombstone that round went with it. Agents deploy FLATTENED, so
both the push scan and the post-pull cleanup read a bare tombstone globally:
removing `fe/vr` suppressed and deleted `be/vr` the moment that namespace
became active. Only the published name is tombstoned now. The author's own
flattened copy is still swept, but only where this machine's record says the
file just removed is where push put it — without that, the copy on disk may be
another namespace's deployment.

Covered end to end this time: the new case drives `teamai remove agents vr`
through the built CLI, which is the join the round-7 unit tests skipped.

* fix(push): raise a project agents-axis failure the scan would otherwise swallow

`--project <id>` resolved the agents destination before scanning and dropped
the failure on the floor. A project with no agents namespace then looked
identical to a run with no flag at all: the scan skipped the agent as "no
active source", the item never reached placement, and the command exited 0 with
"No new or modified resources" — on a flag it could not honour. The error is
carried forward and raised as soon as the scan contains an agent. It cannot
wait for the selection the way the skills axis does, because the item that
would prove the axis is needed is exactly the one the scan removes.

`placedResourcePath` matched the recorded filename by prefix, so a record
pointing at `rules/<ns>/foo.backup.md` was trusted whenever that file existed,
and scanning, the pre-push sync and removal would all follow it onto somebody
else's file. The filename must now be exactly the resource's own.

* fix(agents): reach the canonical source, and hold the record to what it proves

Four review findings, all on the agent side of placement.

The single-repo canonical source in `.teamai/agents/` is picked up directly,
never reverse-parsed, and that branch ignored the placement record: a root
`vr.yaml` placed at `agents/fe/vr.yaml` read as new on the next push, and the
collision guard then refused the very agent this machine published. Removal
missed the same directory, so the agent republished itself on the next push —
which a bare-stem tombstone cannot prevent without suppressing that stem in
every other namespace, since agents deploy flattened.

The project agents-axis error now counts only agents that actually need a
destination. A modified agent already in a namespace is written in place, so an
empty agents axis is none of its business; blocking it contradicted the rule
that only new shared-root resources are placed.

And the record is no longer taken as licence to overwrite. It admits a namespace
this directory never activates, which also means `pull` never refreshed a copy
and the pre-push sync does not cover agents — so if the canonical file moved on
since the last pull, push now says so and asks for a pull instead of writing a
stale rendering over whoever changed it.

* fix(agents): deliver recorded agents on pull, and let a named destination win

The staleness guard added last round was defeated by the pull it recommended:
`pull` advances lastPullRev without deploying an inactive namespace, so the
next push saw an unchanged canonical and wrote the stale rendering anyway. It
also never fired right after the first PR merged, when the file did not exist
at lastPullRev.

The guard is gone, and the cause with it. `pull` now delivers an agent whose
placement record names it, so the local copy tracks the team file and the
ordinary comparison is valid — the inactive case stops being special instead of
needing its own machinery. A stem an ACTIVE namespace already claims is left
alone, since agents deploy flattened and the active one is what is deployed
here; the scan follows the same order, treating the record as a fallback rather
than an extra candidate.

Pending-PR reuse matched on type and name alone, so an open PR for a different
resource of the same name captured a push that named another namespace and
force-pushed into that review. Neither silent answer is safe, so the flag the
user typed decides, the open PR is left untouched, and the collision is
reported. This supersedes the original #331/#654 rule that a PR's destination
always won; that rule still holds whenever no destination is named.

* fix(pull): stop revoking the agent pull had just delivered

Self-review of the branch, before the next review round.

Round 11 taught delivery about placement records but not revocation, and both
run in the same pull: `filterAgentsByNamespaces` wrote the agent and
`cleanupInactiveNamespaces` deleted it again, byte-equal to the render so the
data-safety gate passed it straight through. The record-based delivery was
inert and the file churned on every pull. Both halves now resolve through one
exported `selectAgentsForDirectory`, so they cannot disagree — the same
treatment `placedResourcePath` already gives the push scanner and the pre-push
sync.

Two smaller ones from the same pass. `--dry-run` grouped against the unfiltered
pending list, so it reported a destination the real push no longer uses, which
breaks the property that a dry run matches the run it describes. And the
partial-selection warning counted entries the run had already declined to
reuse, contradicting the warning given for them; both now share the filtered
list, which is computed once there is actually something to push.

* fix(pull): stop the stale sweep from deleting the author's own placed rule

`pullAllRules` sweeps a local rule whose name is absent from the desired set.
A rule published into a namespace keeps the author's copy at the rules ROOT
under its bare name, while the desired set holds `<ns>/<name>` — or nothing at
all when that namespace is not active here — so the sweep deleted their own
file, local edits included. The placement record marks it as theirs, and only
while the team file it points at still exists.

Pending-PR conflict detection trusted `PendingPushItem.namespace`, but the
agent scan records a namespaced destination without setting that field, so
those entries slipped past the check and a push naming another namespace could
force-push into the PR under review. The namespace is derived from the recorded
path when the field is absent, and the agent scan now sets it too — the field
was the only thing telling `pendingNamespaceFor` where to put the resource.

* fix(push): defer the agents-axis failure to selection, reload projects.yaml after the pull, and stop a named namespace reusing a shared-root PR

A project with no agents namespace failed before the listing whenever any new
agent was present locally, so a rules-only push under `--project` was blocked
on an agent the user was never given the chance to deselect. Only an agent the
scan itself skipped (`needsDestination`) fails early now — that one never
reaches the listing, so deferring its error means never raising it. A new agent
is listed, and step 4 raises the same error if it stays selected.

`manifest/projects.yaml` was read in `push`, before `pushCore` pulled the team
clone, so a project whose namespaces changed on the remote placed this run's
new rules and agents by the previous pull's mapping. It is read inside
`pushCore` now, after the pull, and in self mode from the fresh worktree.

Pending-PR conflict detection treated a recorded path with no namespace as
non-conflicting, so an explicit `--role`/`--project` reused a shared-root PR's
branch and rebuilt it with the namespaced path, moving a review the user did
not name from "everyone" to one namespace. The shared root is a destination
like any other: it conflicts with any namespace the flag names.

`pull` delivered a rule this machine placed at `<tool>/rules/<ns>/<name>`
beside the author's copy at the rules root, so a tool that loads rules
recursively applied the same rule twice, disagreeing as soon as the team file
moved on. The placement record names the root copy as this rule's local file,
so delivery updates it and removes the namespaced duplicate an earlier pull
wrote. A namespace another member placed is untouched.

* fix(push): stop on a stale clone under --project, retire a renamed canonical agent's recorded file, and prune placement records

A failed refresh of the team clone was only warned about, after which
`--project` resolved every destination from the previous pull's
`manifest/projects.yaml`. A namespace the remote had changed sent this run's
new rules and agents to the members of the old one. `--project` stops now, and
so does a new resource that would resolve from `manifest/roles.yaml`; a team
with no roles manifest resolves from nothing that can go stale and keeps its
behaviour. `--role` names the namespace itself and is unaffected.

In self mode the placement record redirected a root canonical agent to the
recorded file, extension included. `pushItem` writes by the source's
extension, so an author who rewrote `vr.md` as `vr.yaml` had the `.yaml`
written and the `.md` staged: the change never reached the branch and the
`.md` stayed. The destination now keeps the record's directory and the
source's extension, `pushItem` deletes the file it retires, `push` stages that
deletion and moves the record to the new path.

Placement records were written when a branch reached the remote and never
removed. Once the PR was closed unmerged, or the file deleted upstream, the
record pointed at nothing — until another member created the same path, at
which point it came true again and their unrelated resource read as this
author's. `push` and `pull` now drop a record whose target is neither on the
default branch nor awaiting review on a branch origin still has, before
anything reads the records. A record is kept when origin cannot be asked.

* fix(push): record a placement only once it has landed, withdraw it when a shared-root file takes the name, and drop every stale record

Placement records were written when the branch reached the remote, so a PR
closed without merging left one behind for as long as its branch stayed —
and no provider here can say whether a PR is open. The record now travels on
the pending PR entry (`PendingPushItem.placed`, with the blob push wrote) and
becomes a `placedRules`/`placedAgents` record only when that blob is in the
default branch's history for the path: the PR merged, however the platform
merged it. A path that merely exists is not enough, since another member may
have created it after the PR was closed. `push`, `pull` and `remove` settle
this before reading the records; the stale sweep spares a root copy whose
placement is still awaiting review.

`localNameFor` redirected a recorded rule onto the bare root path without
asking whether a shared-root rule of the same name was being delivered too,
so both landed on the one file in loop order, and the next push could follow
the record and carry the shared rule over the namespaced one. Delivery keeps
the namespaced path when the shared root holds that name, and the reconcile
pass withdraws the record with a warning: the root copy follows the shared
rule from then on.

Dropping stale records destructured from the ORIGINAL map each iteration, so
a later deletion put back what an earlier one had removed and only the last
stale record went. The kept entries are rebuilt in one pass.

* fix(push): consume a pending placement once it is recorded

Recording a landed placement left its `placed` mark on the pending entry. Had
the team then deleted the file — which drops the record — and another member
recreated the path, the next reconcile recorded it again: the path existed and
the blob push had written was still in the default branch's history, so both
checks passed, and the unrelated replacement read as this author's resource.
The mark and the blob are cleared when the record is written, so a placement
is recorded exactly once.

* fix(remove): keep placement records until the removal lands, and retire flattened copies

`remove` dropped the placement record as soon as its branch was pushed, so a
retry during review resolved `vr` to the bare stem and removed that agent from
every namespace. Records are now left to the reconcile pass, which drops one
once its file is gone from the default branch, or was deleted since the last
check (`placementsCheckedAt`) even if another member has recreated the path.

A namespaced removal tombstones only `<ns>/<name>`, which never matched the
flattened `<agents>/<name>` copy members hold, so that copy survived pull and
the next push republished it. `AgentsHandler.removedStems` reads the tombstone
as the flattened stem while no namespace still has that agent, for both the
pull cleanup and the push scan. Rules no longer write a bare tombstone, which
swept and suppressed other members' unrelated rules of the same name.

Agent and rule push scans skip tools the member excluded: remove leaves those
copies behind by design, so reading them republished the removed resource.

Landing is proven only by history after the full commit the push branch was
built on, and a placement whose path was deleted after it landed is spent
unrecorded. Single-repo pull no longer reconciles against the member's own
checkout; push and remove still do, in a fresh origin/<default> worktree.

* fix(pull): reconcile single-repo records through origin/<default>, and retire flattened copies per directory

Single-repo pull skipped the reconcile pass, so a merged placement stayed
unrecorded until the next push or remove. It now reads the default branch as
the ref origin/<default> (existence via `<ref>:./<path>`, history up to that
ref) instead of the member's own checkout, and changes nothing when the ref
cannot be resolved.

`placementsCheckedAt` survived its last record, and a record made in the same
run was checked against it, so re-placing a resource at a path deleted earlier
dropped the new record at once. The checkpoint is cleared with the last record,
and records made in this run are not held to it.

`removedStems` retired the flattened stem only when no namespace had it at all,
so an fe member kept a removed fe/vr while an unrelated be/vr existed. It now
asks what this directory is meant to hold, through the same selection pull
delivers with.

* fix(remove): stop on a stale clone, and judge PR conflicts by the destination the flag gives

`remove` ignored a failed refresh and reconciled the stale clone as if it were
the default branch. A placement merged since the last pull was then not
recorded, and the bare name fell back to the stem, removing that agent from
every namespace. `remove` now stops with exit 1 and removes nothing.

Under --role/--project, a pending PR counted as conflicting whenever its
recorded namespace differed from the flag's, although only skills and new
shared-root rules and agents are moved by it. A modified rule already in a
namespace kept its path yet went to a second PR on the same file. Only an
item the flag actually moves can conflict with it now.

* fix(push): prefer an agent's delivered source over the flag, baseline new placements, validate the record's namespace

With --role/--project the requested namespace always chose the team file a
local agent was compared with, so an untouched copy delivered from an active
namespace read as an edit of the requested namespace's agent and overwrote
it. Candidates now follow delivery: an active source (the shared root
included), then this machine's record, and only then the requested namespace.

A placement that landed after the last pull had no `lastPullRev` version, so
the pre-push sync skipped it and a teammate's edit before the author's next
pull was pushed over. Rules take the version the file was added with as their
base; a recorded agent, which has no pre-push sync, is held with a pull-first
message when the team file moved past that baseline.

`placedResourcePath` now requires a safe namespace segment: a backslash in it
is a separator on Windows and walked out of the resource root.

* fix(push): close the round-21 findings on placement, removal and agent sources

- A pending namespaced placement whose name a shared-root file now takes is
  left out of the push with a warning, instead of the open PR being rebuilt
  with the author's copy over the shared file.
- `remove` stops when the placement records cannot be reconciled and saved:
  a missing record sends the bare name to the stem, which spans namespaces.
- An agent skipped for want of a --project agents namespace no longer blocks
  the rest of the push; the error stands only when nothing else is left.
- A placement is marked only with a blob that can prove it landed, and one
  without is spent rather than recorded because its path exists.
- The single-repo `.teamai/rules` scan source is not a tool, so an
  `enabledAgents` list no longer hides it.
- A namespaced agent tombstone retires the flattened stem only where that
  agent could have been delivered; reconcile keeps the author's dropped
  record (`retiredPlacedAgents`) so their own copy still counts.
- Two active same-name agents stay ambiguous under a flag, and a flag naming
  a namespace that already holds the agent is a collision, as for rules.
- In single-repo mode a root copy equal to an older version of its placed
  file is held as stale rather than pushed over a teammate's edit.
- The recorded-agent hold runs only for a changed copy and says to set the
  edit aside first; a pending placement is routed to its PR, not skipped;
  a flag that does not move a shared-root edit says so; several candidate
  namespaces without a terminal fail with a --role hint; a placement that
  landed with other content is reported once.

* fix(push): stop on unsaved records and stale placements, name namespaced agents on remove

- `push` stops, pushing nothing, when the reconciled placement records
  cannot be saved: the sync and the scan read them back from disk, and a
  record that could not be withdrawn still redirects the author's copy.
- On a stale clone every unflagged placement stops, not only one resolved
  from an existing roles manifest: the manifest's absence and the skills
  namespaces detected from the tree are clone state too.
- `remove agents <ns>/<name>` names one namespaced agent; a bare name only
  one namespace has resolves to it, and a bare name found in several places
  is refused rather than removed from all of them.
- A record dropped because its file was deleted and recreated is retired
  like one whose file is simply gone, so the author's flattened copy of the
  removed agent is still recognised.

---------

Co-authored-by: Saul Moro <smoro@ai-lab.knowmadmood.com>
2026-09-23 14:58:21 +08:00
Saul Moro 94eb1484bc fix(init): refuse provider logins without a terminal instead of hanging (#711) (#713)
* fix(init): refuse provider logins without a terminal instead of hanging (#711)

`teamai init` with no session spawned `gh auth login --web` (or `gf auth
login`, `cnb login`) with inherited stdio and waited for a browser device
flow nobody could complete, about five minutes for GitHub, then exited
with the provider's error and no hint of the missing credential.

Cause: the non-TTY guard lived only in utils/prompt.ts. A login is a child
process that owns the terminal, so it never went through that guard.

- `isInteractive()` in utils/prompt.ts: stdin is a TTY and neither `CI`
  nor `TEAMAI_NONINTERACTIVE` is set. Every prompt and the four
  prompt-semantics `isTTY` checks use it; the six hook-payload checks are
  untouched.
- github, tgit and cnb logins throw before spawning when not interactive,
  naming the token variable, the way gitcode already did.
- index.ts exports GIT_TERMINAL_PROMPT=0 when not interactive, so a
  missing clone credential fails at once instead of prompting or opening
  a credential helper dialog. An explicit caller value wins.
- e2e test with a fake `gh` whose `auth login` sleeps: exit 1 in under a
  second naming GITHUB_TOKEN, also under CI=true.

Closes #711

* fix(init): close every git prompt and point TGit at the credential that works

Review follow-up on #713.

- utils/git-env.ts: GIT_TERMINAL_PROMPT=0 only closed git's own terminal
  question. The askpass chain (GUI dialog), ssh's passphrase / unknown-host
  question through /dev/tty, and Git Credential Manager's window each still
  parked an unattended clone until the 180s timeout. All four are now closed
  together (GIT_ASKPASS=echo, GIT_SSH_COMMAND='ssh -o BatchMode=yes',
  GCM_INTERACTIVE=never), each only where the caller set nothing.
- tgit: the guard suggested exporting TGIT_TOKEN, which cannot make an
  unattended run succeed — the PAT is REST-API-only and git.woa.com's git
  endpoint rejects it, so `gf auth whoami` still fails and the clone still has
  no credential. The message now names `gf auth login` (whose stored credential
  is the one that works) and says why the token is not it. Docs follow.
- local-agent: keep askViaTty's non-interactive decline synchronous. Awaiting
  the prompt module's import before declining shifted hook-path timing enough
  to break the once-per-session binding hint (local-agent.test.ts).

* fix(git-env): append batch mode to core.sshCommand instead of replacing it

Review follow-up on #713.

GIT_SSH_COMMAND overrides core.sshCommand rather than extending it, so setting
it blindly dropped a configured custom key, ssh binary or wrapper and left the
run unable to authenticate at all. The value is now composed: read
core.sshCommand and append `-o BatchMode=yes`, or use plain `ssh` when nothing
is configured. A command that already decides BatchMode is left alone, and the
config read is skipped entirely when the caller set GIT_SSH_COMMAND.

Test isolation, so the suite's own result can be trusted:

- shell-profile.test.ts: the three Windows cases never stubbed SHELL, and
  detectShellProfile reads it before the platform branch — a suite run from a
  zsh login shell resolved .zshrc and failed them without ever reaching the
  Windows branch. CI runners use bash, which is why only local runs saw it.
- local-agent.test.ts: the once-per-session binding-hint markers live in
  os.tmpdir() under one shared key, so a leftover marker decided whether the
  next test emitted a hint, and concurrent runs competed for the same paths.
  Each test now gets its own temp directory, which makes the markers per-test
  by construction.

Full suite: 3788 pass, 0 failures, five consecutive runs.

* fix(git-env): leave ssh to each repository instead of a process-wide override

Review follow-up on #713.

GIT_SSH_COMMAND is the only way to reach ssh's batch flag, and it overrides
`core.sshCommand` for *every* later git operation, not just the one the value
was derived from. Reading the launch directory's config and exporting it
process-wide therefore pushed that repo's key or wrapper onto the managed team
repo, and the plain default suppressed a `core.sshCommand` the managed repo had
configured for itself. Prompt suppression must not reach a repository's
transport, so the variable and the `git config` read are gone.

What remains is the three variables that name a prompt and nothing else, so one
value is right for every repository a run touches: GIT_TERMINAL_PROMPT=0,
GIT_ASKPASS=echo and GCM_INTERACTIVE=never.

An ssh remote that would still ask is now documented as the caller's to close,
per repository (`git config core.sshCommand 'ssh -o BatchMode=yes'`) or per run
(`GIT_SSH_COMMAND`). Measured first: with stdin closed, ssh's own tty read hits
EOF and fails in about a second, so the unattended paths this PR is about do
not depend on the flag.

Also reverts the shell-profile.test.ts and local-agent.test.ts isolation edits
from the previous round: neither traces to the unattended-login fix, so they
belong in their own PR.

* docs(providers): mark the provider logins interactive-only (#711)

docs/providers.md still described `teamai init` as running `gh auth login`,
`gf auth login` and `cnb login` unconditionally. Each now happens only in an
interactive terminal; an unattended run fails at once naming the credential to
prepare (a token for GitHub and CNB, a prior `gf auth login` for TGit, since a
TGIT_TOKEN PAT is REST-API-only and cannot clone).
2026-09-23 10:41:13 +08:00
Saul Moro cd3e0e6ef5 fix(hooks): stop the Stop hook nudge reaching the user twice (#720) 2026-09-22 22:47:35 +08:00
Saul Moro 2ed17e4fb2 feat: scope hooks, MCP servers and env variables by logical project (#700) 2026-09-22 22:44:56 +08:00
Saul Moro a3366b9bf6 chore(git): ignore local copilot agent sync files (#694)
Ignore Copilot sync targets under .github/ (.github/hooks/teamai.json, .github/skills/, .github/instructions/*.instructions.md, .github/agents/, .github/copilot-instructions.md, and .github/mcp.json) as a best-effort ignore pattern while preserving other GitHub files.
2026-09-21 17:05:52 +08:00
Saul MoroandSaul Moro 52525a9402 feat(doctor): extend the delivery check to rules, agents, MCP and env (#669)
* refactor(doctor): resolve delivery destinations through one handler seam

`doctor` asked "can this tool receive skills" through `skillsReachTool`, which
had to invent a skill name (`__teamai_probe__`) because `skillTargetForTool`
fused two questions: whether a tool receives skills at all, and where a given
skill lands. Only a comment said the invented name could not affect the first.

Split the gate from the path. `skillsDirForTool` answers the gate on its own —
OpenClaw's workspace, Hermes' home, Copilot's enabledAgents, else the tool root
— and `skillTargetForTool` is that directory plus the skill name, with Codex's
shared-directory redirect on top since only that one is per-skill.

Add `ResourceHandler.deliveryTargets`, the read-only seam #624 asks for: where
an item lands for each tool that receives it, `null` for a resource with no
per-tool file destination. `SkillsHandler` implements it, and its `pullItem`
now walks the same resolved targets, so the write path and the check cannot
answer differently. `buildDeliveryChecks` consumes the seam through the handler
registry instead of importing `SkillsHandler` directly; check names, failure
buckets and fix text are unchanged.

Also points the inactive-skill cleanup at the same gate. It probed the tool
root and then swept `<base>/<skills path>`, which for OpenClaw is a directory
delivery never writes to — the real workspace copy was never pruned.

* feat(doctor): check that rules reached each tool in its own format

A rule changes both its filename and its bytes per tool: `.md` verbatim for
Claude, `.mdc` with derived `globs`/`alwaysApply` for Cursor-compatible tools,
`.instructions.md` with `applyTo` for Copilot. Nothing exposed where one lands,
so `doctor` could not ask — the extension table lived inside `pullItem`.

`RulesHandler.deliveryTargets` answers it, and `pullItem` now walks the targets
it returns rather than rebuilding the gate chain, so the check and the write
path resolve the same paths. `resolveDesiredRules` joins `resolveDesiredSkills`
in pull.ts: the namespace convention and the tag channel are stated once, and
the check reads them rather than restating them.

The check reports two buckets per tool: a rule that never arrived, and one that
arrived without the frontmatter its tool reads — a `.mdc` without `alwaysApply`
is inert, which no write-time gate can see because the write succeeded. The fix
names the destination directory, since the filename is not the rule's name.

A legacy `.md` left beside a correct `.mdc` is deliberately not reported: it is
inert leftover that `pullAllRules` already sweeps, not a delivery failure.

* feat(doctor): check that agents reached each tool they target

Agents break the items × tools shape the skills check assumes: a spec carries
`targets:`, so the desired set is a relation, and each tool renders its own
format, so the filename comes from the render and not from the agent's name.
`AgentsHandler.deliveryTargets` is therefore the only thing that can say where
an agent lands, and `pullItem` now walks the same resolution — the YAML and
legacy paths merge into one loop instead of two gate chains.

`resolveDesiredAgents` joins its skills and rules siblings in pull.ts, so the
namespace filter and its stem-collision throw are stated once; `doctor` reports
that throw as a failing check rather than stack-tracing, as it already does for
skills.

Two bugs surfaced while unifying the paths:

- `renderedForTool` decided "legacy" from `item.legacy` alone while `pullItem`
  also accepted a non-`.yaml` source. An item built without the flag was
  therefore parsed as a spec by the cleanup and copied verbatim by the pull.
  `isLegacyAgent` now answers it in one place.
- The parse-failure warning was Chinese, which the repo forbids in production
  code, and went through `console.warn` rather than the logger.

Also adds a check for an agent that renders for no installed tool at all: the
file is in the team repo, `pull` names the reason once, and nothing afterwards
says it is still reaching nobody.

* feat(doctor): check that MCP servers reached each tool's own config

An MCP server is an entry inside a tool's native config, not a file of its own,
so this check takes the shape of the hook check rather than of the delivery
seam: it asks which servers the reconcile would want for a tool, then whether
that tool's config carries them.

The desired-set pass moves out of `reconcileMcpForConfig` into
`desiredMcpForTarget`, unchanged — the `tools:` and `roles:` filters, the
transport and policy gates, the `requires:` PATH check and the placeholder
resolution all stay in one place, and the reconcile now calls it. A second copy
of those filters is precisely how a server skipped once for an unresolved
variable gets reported as delivered forever after.

That skip reason is the point. A server dropped for `unresolved variable(s)`
prints one line during a pull and is never mentioned again, so the member sees
"MCP does not work" and goes looking at MCP. The check now names the server,
the variable and `env/env.yaml` — including that its top-level key must be
`variables:`, since a plain `KEY: value` mapping parses as no variables at all
and silently skips every injection (#662).

A server the member excluded on purpose is not reported; an unparseable tool
config is, because the write path abandons the injection there too.

* feat(doctor): check that env variables reach a shell, not just a marker

The env check asserted that `# [teamai:env:start]` appeared somewhere in the
profile. That is true of a block that cannot load and of a run that delivered
nothing, so both failures passed and surfaced three layers away as MCP servers
skipped for `unresolved variable(s)`, with nothing pointing back at env.

It now asks the three questions the marker stands in for:

- Does `env.yaml` declare anything? A file with content that parses to zero
  variables is the shorthand `KEY: value` form, which zod strips to an empty
  list — the pull then writes nothing and logs nothing (#662).
- Did every declared variable reach `env.sh`?
- Would the injected block load it? The block is built with the platform
  separator, so on Windows it carries backslashes; a POSIX shell reads an
  unquoted `\` as an escape, the `[ -f ... ]` test fails, `&&` short-circuits
  and `source` never runs, silently (#661). Whitespace in the path needs
  quotes for the same reason.

Neither underlying bug is fixed here — #661 and #662 own those. This is the
row missing from the issue's table: env had no check that looks at the payload,
which is why both of them reach `All checks passed!`.

The check is still emitted when there is nothing to deliver, passing, since
`doctor --json` consumers cannot tell an absent entry from a passing one.

* feat(doctor): let the caller pick the stage instead of flagging each check

The post-pull pass re-runs the registry under a 5s all-or-nothing budget that
covers building it as well as running it. Skills and docs cost a stat per item;
rules cost a read per rule per tool and agents parse every spec. Adding those
to the pass would spend the budget on the expensive checks and lose the cheap
ones — and going over means the member gets no check at all.

`buildChecks(ctx, stage)` takes 'pull' or 'doctor' and does not build the two
expensive registries for 'pull'. The stage is a property of the caller, not of
a check, so it is an argument rather than a third optional flag on `Check`
beside `source` and `reportedByPull` — which the issue flags as the point where
that object stops reading.

Skipping is at build time, not a filter over the result: the cost is in
building the registry, so filtering afterwards would save nothing.

* docs(doctor): describe the rules, agents, MCP and env delivery checks

* test(doctor): cover the delivery checks through the built CLI

* refactor(doctor): move the delivery checks out of the command file

Review findings, all three from the repo's own standards.

`doctor.ts` had grown to 959 lines, most of it domain logic: where a rule lands
for Cursor, which tools an agent's spec targets, whether a shell block would
load. CONTRIBUTING says commands in `src/*.ts` stay thin and the heavy lifting
lives elsewhere. The checks move to `doctor-delivery.ts`, and `doctor.ts` is
back to being the registry that runs them — smaller now than before this branch.

The three per-tool builders repeated one shape: walk items × targets, bucket
the failures by tool, remember the directory, format a check. `walkDelivery`
holds that walk and takes a `classify` callback for the part that genuinely
differs; `describeProblems` formats the buckets in the caller's label order, so
the same broken machine reads the same way twice rather than in the order its
failures happened.

`envDeliveryProblems` had its own copy of the `$SHELL` → `.zshrc`/`.bashrc`
choice, a second spelling of what `EnvHandler.detectShellProfile` already
decides — the exact failure this branch exists to prevent, one layer down: it
would check `.bashrc` while the pull wrote `.zshrc` and call a correct install
broken. That method is now public and the check calls it.

No check name, failure bucket or fix string changes.

* refactor(doctor): drop the unused null return from deliveryTargets

AGENTS.md's review rules reject unused flexibility, and this was some. The
seam returned `DeliveryTarget[] | null`, where `null` meant "this resource has
no per-tool file destination" and `[]` meant "no installed tool receives it
here". The single caller wrote `?? []` and treated them alike, so the
distinction only cost a branch nobody took.

The default is `[]` now, and the comment carries the meaning the type was
trying to.

* refactor(doctor): drop the imports the delivery move left behind

Moving the checks into `doctor-delivery.ts` left nine imports in `doctor.ts`
with no remaining user: `fs`, `expandHome`, `listFilesRecursive`,
`TEAMAI_ENV_END`, `getMcpSharing`, `usesCursorMdcRules`,
`usesCopilotInstructions`, `splitFrontmatter` and the `ResourceItem` type.

`tsc --noEmit` stays green either way because `noUnusedLocals` is off, so CI
could not have caught these. They make `doctor.ts` look like it still reaches
into frontmatter parsing and MCP sharing config, which is the impression the
move existed to remove.

* fix(doctor): report unreachable agents from the tools, not from the renders

`Every team agent reaches a tool` was gated on `byTool.size > 0`, using
successful deliveries as the proxy for "some tool was there to receive an
agent". It is the wrong proxy for exactly the case it exists to catch: when
every agent is malformed or targets tools that are not installed, no agent
renders anywhere, `byTool` is empty, no check is built at all, and `doctor`
reports success on a machine where nothing arrived.

The gate is now the installed tools themselves. `AgentsHandler.agentToolDirs`
answers that on its own — the tool-path, exclusion and install gates without
asking any agent to render — and `resolveRenders` and the inactive-agent
cleanup, which both carried their own copy of that loop, now go through it.

* fix(doctor): compare MCP entries with the team definition, not their names

The check asked whether the desired server name was a key in the tool's
config. Reconciliation never overwrites an entry teamai does not own, so the
one case the write path deliberately skips — a server of your own under a team
name — satisfied the check: the key is there, the team's server is not, and
every later pull skips it again without a word.

`installedMcpEntries` replaces `installedMcpServerNames` and returns the
entries in the rendered form `desiredMcpForTarget` produces, so the check
compares values. Structurally, via `isDeepStrictEqual`: key order in a JSON
config is not meaning, and a tool that rewrites its own file should not read
as a failure. Codex stores a TOML block rather than a JSON value, so
`codexBlockIn` extracts the block by the same regex `spliceCodexBlock` writes
with, trimmed to the single trailing newline `renderCodexBlock` emits.

A stale entry and a foreign one are reported alike, as `not the team's
definition` — both mean the tool is not running what the team declared — and
the fix says that a pull leaves an entry teamai does not own alone, so only
`--force` replaces it.

* fix(doctor): compare env.sh assignments with their values, not their keys

`export KEY=` as a substring is true of the value env.yaml declares and of the
one it replaced. A rotated credential that never reached `env.sh` — the pull
that would rewrite it skips a scope whose team repo has not changed — passed
the check while every shell and every MCP server kept exporting the old value,
which is the failure this check exists to name.

Each declared variable is now compared against the line `generateEnvFile`
would write for it, the injection's own rendering rather than a second copy of
its quoting, and a key present with a different value is reported as stale
rather than as missing. Neither value is printed: these are credentials, and
the key is the whole diagnosis.

The e2e fixture delivered an MCP entry and an `env.sh` that were not what
teamai writes; it now carries the rendered forms, and covers a foreign server
under a team name and a stale `env.sh` through the built CLI.

* docs(doctor): say what the delivery checks compare, not just that they check

The MCP and env paragraphs described a name lookup and a key lookup. Both now
compare values, and the MCP one reports a server of your own holding a team
name — which only `teamai pull --force` replaces — so the guide and the
changelog have to say so. Both language versions.

* feat(doctor): compare a delivered agent with its render, not its existence

The check asked only whether something readable sat at the destination, which
is the same class of gap the three review findings were: an agent rendered
from an older spec passes while the tool runs instructions the team replaced.
A plain pull syncs a scope only when its team repo changed, so the copy can
sit there indefinitely.

`DeliveryTarget` carries the bytes `pullItem` writes, which `resolveRenders`
already had in hand and threw away at the seam, and the agents check compares
them. It is the same equality the inactive-agent cleanup already uses to
decide a deployed copy is the team's. Absent `content` means the handler
renders nothing — a skill is a directory tree — and only existence is judged,
so skills and rules are unchanged.

`walkDelivery` passes the target to `classify` rather than its two fields.

The fixtures delivered the literal string `rendered`, which the new comparison
correctly rejects: the unit tests now deliver through the handler's own seam,
and the e2e fixture carries each tool's render byte for byte.

* fix(agents): leave a member's same-stem file alone beside a legacy .md

Routing the legacy `.md` path through `resolveRenders` also gave it the
stale-sibling sweep, which the old `pullLegacyMd` never ran. A team agent
named `helper` then deleted a `helper.toml`, `helper.json` or
`helper.agent.md` the member wrote, with no ownership or content check.

Only a rendered spec can leave a sibling behind: its extension follows the
tool's format and changes when `targets` does. A legacy `.md` is copied
verbatim to one extension for every tool, so anything else on the stem is
not ours.

* fix(doctor): compare a delivered rule with its render, not its key names

The check read the delivered file for the presence of `alwaysApply` or a
nonempty `applyTo`. A `.mdc` whose `globs` no longer match the team rule's
`paths:` passes that while Cursor applies it to the wrong files, and so
does a body that drifted from the team `.md`.

`RulesHandler.deliveryTargets` now carries the bytes `pullItem` writes, the
way the agents handler does, and the check compares against them. That
makes the render the single spelling of the mapping rather than a contract
`doctor` restates in terms of the keys it happens to know about.

* fix(doctor): keep the reason an mcp.yaml yielded no servers

`parseTeamMcpServers` answers `[]` to an absent file and to one that does
not parse alike. That is right for a pull, which can only skip the run, but
it left `doctor` unable to tell a team with no MCP from a team whose every
server reaches no tool: the desired set was empty, no per-tool check was
emitted, and `doctor --json` reported ok: true.

`readMcpYaml` returns the parse failure with its reason and the check
reports it. `parseMcpYaml` keeps its old shape on top of it, so the pull
path is unchanged.

* fix(doctor): tell a parse failure from a deliberately empty env.yaml

`parseEnvYaml` answers `[]` to four different files: absent, empty,
`variables: []`, and the shorthand `KEY: value` mapping whose unknown
top-level key zod drops (#662). The check equated zero variables with the
shorthand form, so an intentional `variables: []` was reported as
malformed.

`readEnvYaml` returns the reason instead of the count, so the shorthand
form and invalid YAML are both named while an empty configuration fails
nothing.

* docs(doctor): say what the rules, MCP and env checks compare after the review

The guides and the changelog entry describe what each check compares, and
three of them now compare something else: a delivered rule against its
render rather than its frontmatter keys, an unparsable `mcp.yaml` as its
own failing check, and an explicit `variables: []` as an empty
configuration rather than a malformed file.

The e2e suite covers all four cases through the built CLI.

* fix(doctor): check the two rule destinations that are not a file per tool

`deliveryTargets` covers what `pullItem` writes under `toolPath.rules`.
`pullAllRules` delivers two more things it cannot see, and both fail
silently:

OpenCode does not auto-scan a rules directory. Every `.md` can be there
byte for byte and be inert, because `opencode.json` no longer lists the
glob the pull owns — and `Rules delivered to opencode` passes throughout.

Hermes has no rules directory at all: its rules are the contents of a
managed block in SOUL.md. A deleted or stale block is a tool reading the
wrong rules with nothing on disk to show for it.

Both take the shape of the hook and MCP checks — one destination, not one
per tool. `opencodeInstructionsTarget` and `hermesRulesText` are the single
spelling each, so the check reads the answer the pull writes rather than
deriving a second one.

* fix(doctor): match a multiline env value instead of calling it stale

A YAML block scalar is a legal env value, and `generateEnvFile`
single-quotes it into an export spanning several physical lines. The check
split env.sh on newlines and compared each line with a whole generated
export, so such a value could never match: a correct pull was reported as
a stale value on every run.

`parseEnvFile` is the generator's inverse — it reads the assignments back,
including the `'\''` encoding of an embedded quote — and the check compares
values rather than lines.

* docs(doctor): describe the two rule activation checks and the env inverse

Two checks are new and one comparison changed, so the guides and the
changelog entry describing them change with it. The e2e suite covers both
through the built CLI: OpenCode rules delivered byte for byte while the
glob is gone, and a multiline env value that the old line scan called
stale.

---------

Co-authored-by: Saul Moro <saul.moro@darstelecom.es>
2026-09-20 23:29:01 +08:00
Saul Moro 701e472b0b feat(doctor): run the checks after a pull, and check what actually landed (#625)
* feat(pull): report failing checks at the end of an interactive pull

Every line a pull prints reports what it did; none reported what is on
disk. That gap is the shape of #574, #525, #342 and friends: "Synced N
skills" and the tool receives nothing.

An explicit `teamai pull` now re-runs the doctor registry and prints only
the checks that failed, with the fix each one already carries. The
SessionStart hook path (`pull({ silent: true })`) and `--dry-run` run no
checks at all, so session startup is unchanged.

Checks now declare `source: 'local' | 'provider'`. The post-pull pass runs
the local ones only: the pull just used the provider successfully, so
re-probing `gh auth status` would add a subprocess to every sync and prove
nothing new. `teamai doctor` still runs the full registry.

For #598.

* feat(doctor): fail when an enabled tool is not installed

buildHookChecks skipped any tool whose settings directory was missing —
the same silent skip #574 reports in pull, reproduced inside doctor. With
`claude` installed and `codex` not, the report was all green while codex
received nothing.

A tool listed in `enabledAgents` is the user's own claim that they use it,
so it now yields a failing `<tool> is installed` check with a fix that
points at `teamai uninstall --agent <tool>`. Without `enabledAgents` the
team's tool list is aspirational and an absent tool stays silent, so no
existing install grows a new red line.

For #598.

* refactor(pull): extract resolveDesiredSkills from pullForScope

pullForScope computed the desired skill set — role namespaces union the
subscribed tags, minus the exclusions — and dropped it when the run ended.
The delivery check needs the same set, and re-deriving it there would put
that policy in a second place that drifts on its own.

The block moves to an exported, read-only resolveDesiredSkills(), called
from where it stood. roleContext stays an explicit argument: pullForScope
already holds one, and null means "no roles configured", not "not looked
up yet". No behaviour change.

For #598.

* feat(doctor): check that the desired skills actually landed on disk

Every other check verifies plumbing — provider CLI, clone, config, hooks,
env. None verified the payload, which is what #574, #525, #342 and #372
are actually about: the run reports success and the agent finds nothing.

`Skills delivered to <tool>` compares the desired set (role namespaces
union subscribed tags, minus exclusions) against what is on disk for each
installed tool, and names the skills that are missing. It catches what a
write-time gate cannot: per-tool skips, and drift after a correct pull —
a directory deleted by hand, a tool reinstalled, a role changed.

Destination resolution moves into skillTargetForTool(), so pull writes and
doctor checks the same paths, including Codex's shared .agents/skills
directory. A tool that is not installed is asked for nothing; enabledAgents
covers that case with its own check.

For #598.

* feat(doctor): report a skill that landed but stays invisible

A copy can arrive intact and still never be discovered: SKILL.md deleted,
frontmatter that does not parse, or a `name` that does not match its
directory (#372's class). The write succeeded, so no write-time gate has
anything to report.

The delivery check now separates the two causes — "not delivered" from
"delivered but unreadable" — and the fix says which one `teamai pull` can
repair and which one needs the team repo fixed.

For #598.

* fix(doctor): check every enabled tool, not only the ones with hooks

Hanging "is this tool installed" off the hook registry made it invisible
for exactly the tools most likely to be declared and absent: OpenCode and
CodeBuddy ship skills and no hook configuration, so `buildHookChecks`
returned before the question was ever asked. Found driving the real CLI:
`enabledAgents: [claude, opencode]` with no OpenCode root printed nothing.

The check moves to its own builder over ctx.toolPaths, and probes a
resource path rather than the settings path — resources land under
resolveToolBaseDir (the project root in project scope), which is the root
a pull would have to write into.

For #598.

* fix(doctor): survive a team repo it cannot resolve a desired set from

Two findings from the review pass.

A repo with the same skill in two active namespaces makes
scanRoleAwareSkills throw. pullForScope catches it and the post-pull pass
catches it, but `teamai doctor` called buildChecks unguarded: the command
whose job is explaining bad state stack-traced on it. It now reports a
failing "Skills to deliver can be resolved" check carrying the collision.

`copilot is installed` could never fail — isToolInstalledForConfig counts
Copilot as installed as soon as enabledAgents names it — so that dead check
is gone. Copilot's delivery check still reports what did not arrive.

Also: pull and doctor now share formatCheckResult instead of two copies of
the same glyphs, the timeout message interpolates its constant, and the
DesiredSkills block no longer sits between skillSafeToRemove's doc comment
and its function.

For #598.

* docs: document the post-pull checks and the delivery check

Both guides, both languages, same positions: the manual-pull block, the
doctor section, the exclusion and tag-subscription paragraphs, and the
packages-section one-liner that enumerated what doctor checks.

No README change: the `teamai doctor` row still describes it, and touching
it would cost five synchronized translations.

For #598.

* feat(doctor): check the team docs bundle landed

Docs are the one payload with a single destination instead of one per
tool, so the check is a tree comparison rather than a per-tool loop:
every non-dot file under the team repo's `docs/` against
`sharing.docs.localDir`, with the same filter the copy uses.

The destination resolution moves out of DocsHandler.pullItem into
resolveDocsDestination(), so pull writes and doctor checks the same
directory — including the project-scope rule that a `~/` prefix means the
project root, not HOME.

Found while validating it: a doc deleted by hand is not restored by the
next pull, because the rev fast-path skips the scope. The check is what
makes that visible.

For #598.

* fix(doctor): stop telling people to run the pull they just ran

The delivery fixes said "Run `teamai pull`" — printed at the end of a
`teamai pull`, and wrong besides: a scope whose team repo has not moved is
skipped by the revision fast-path, so a plain pull cannot restore a
resource deleted after a correct sync, which is the main case these checks
exist to catch.

Both fixes now say `teamai pull --force` and why. The underlying gap —
that a plain pull does not heal drift — is filed as its own issue.

For #598.

* fix(pull): do not repeat, at the end of a pull, what the pull already said

#621 landed a "Contributed learnings are published" check on the same
registry this pass now runs. A pull with a stuck queue therefore said the
same thing twice, and contradicted itself doing it: pullForScope warns
"run `teamai doctor` for what to check", then the post-pull block answers
with the check's own fix, "Run `teamai pull` to publish them" — the pull
that had just run.

The warning is the better of the two and has to stay: it carries the push
error, which the check cannot learn without attempting a push of its own,
and `doctor` is a read-only diagnostic. Rewording the fix is no good
either, because in `teamai doctor` — where the queue publish runs before
the revision fast-path, so a plain pull really is the retry — that advice
is correct.

So `Check` gains `reportedByPull`, and the post-pull pass skips a check
whose topic this run reported. Evidence, not a declaration: pullForScope
returns before the publish step when the team repo fails to refresh, and
swallows a publish throw into a debug line. On both paths the pull says
nothing about the queue, so suppressing the check unconditionally would
leave a stuck queue reported by nobody.

The topic is a union rather than a boolean for the same reason: the pull
proves what it reported by naming it, so a second tagged check cannot be
silenced by the first one's evidence.

For #598.

* fix(doctor): cap the delivery fix's name list, as the docs one already does

`nameList` was written for the docs check and used only there, while the
delivery check — the one most likely to have a long list, since a fresh
machine is missing every skill at once — joined its names unbounded. A
member with forty desired skills got all forty pasted into one fix line.

Both now go through the helper, which moves above its first caller.

Also drops a stray blank line that a rebase left between
`skillSafeToRemove`'s docstring and the function, detaching the two.

For #598.

* fix(skills): stop reporting a Codex conflict nobody can act on

`resolveSkillDestination` warns when a skill exists in both `.agents/skills`
and `.codex/skills`, unless it can prove the two are identical. That proof
needs the team copy, so the check is written as `sourcePath && ...` — and
without a sourcePath the guard short-circuits into the warning instead of
past it.

The read-only callers are the ones that omit it. `teamai remove skills`
already did on main; this branch added `buildDeliveryChecks`, so the warning
now fires once per skill on every `teamai doctor` and at the end of every
pull, for copies the write path silently reconciles.

Omitting sourcePath now returns the shared destination before the
reconciliation branch, which is what the function's own docstring already
promised. The write path is untouched: with a source, an unprovable pair
still warns.

For #598.

* fix(doctor): own the post-pull evidence, bound the whole pass, report both ways

Three things a review of the post-pull pass turned up.

The set of what a run already said was a module-level `const` cleared at the
top of `pull()`, and the topic it held was a union with one member: two pieces
of machinery where one value does. `pull()` now owns the set and passes it
down. It is a required parameter of `pullForScope` rather than a field on its
optional `policy`, because a call site that forgot it would stop recording
silently, which is the failure the mechanism exists to prevent.
`PullReportedTopic` is gone and `Check.reportedByPull` is a plain string.

The 5s budget wrapped `runChecks` only, while the I/O is in `buildChecks`:
the delivery checks stat every desired skill for every tool as the registry is
built. Both are inside it now. And the pass no longer goes quiet when it gives
up — silence after spending the whole budget is the same "reported success,
nothing happened" shape these checks exist to catch, so it says one line and
points at `teamai doctor`. The reason stays on the debug channel.

`<tool> is installed` only pushed a check when it already failed, so an
installed tool had no entry at all. `doctor --json` is consumed by hooks and
CI, where a missing entry cannot be told apart from one that passed, and no
other check in the registry behaves that way. It now reports both ways.

Docs in both languages and the CHANGELOG follow, including a note that the
checks at the end of a pull cover the scope resolved from the current
directory.

For #598.

* fix(doctor): judge a tool where the sync writes, and keep off a busy clone

Three findings from the Codex review.

A pull that found a scope's lock held by another process drops that scope from
every stage that reads the shared clone, because the other process may have it
on a transient branch. The post-pull checks resolve their own context from that
same clone and ran anyway, so a diagnostic could report a failure about someone
else's work in progress. They now stay out entirely when any scope was
contended; `teamai doctor` runs them once the other process is done.

`<tool> is installed` probed the tool root while skill delivery asks
`skillTargetForTool`, which sends OpenClaw to its workspace directory, Hermes
to its home, and Copilot through `enabledAgents`. A `~/.openclaw` with no
workspace therefore passed the check while delivery skipped the tool and its
delivery check vanished — "reported success, received nothing" inside the
command written to catch it. The probe is now `skillsReachTool`, which asks
that same resolver; a tool that configures no skills path keeps the generic
one, having no such resolver to ask.

`Team docs delivered` called `pathExists`, which follows symlinks and says yes
to a directory, so a name occupied by something other than the document passed
while the document was no more readable than a missing one. It now requires a
file. Reading each one would cost more than the job needs on a bundle of
hundreds of documents, so this stats rather than reads, and the guides no
longer claim the docs check does everything the skills check does.

For #598.
2026-09-18 20:52:36 +08:00
Saul Moro 1f79b2d45e feat(contribute): make the pending queue the write path, and say when it is stuck (#621)
`pending-learnings/` was a durable queue that only ever caught a failed push,
and nothing mentioned it. A member without push rights queued notes forever
while being told each time that the next pull would retry, and in single-repo
mode a rejected push lost the note outright.

Contributing is now: write to the queue, index it, publish from the queue.

- A queue entry is dropped only once its content is confirmed on origin, in
  every mode. Single-repo mode goes through the same queue, which is what stops
  it losing notes.
- One function knows the destination, so contribute runs no git command itself.
- The queue is indexed ahead of the published roots, so a contribution is
  recallable the moment it is written, online or not, and a queued edit of a
  published learning is the copy recall serves.
- `teamai pull` reports what it published and warns, with the reason, when
  anything is still queued.
- `teamai doctor` carries a check for the queue with an actionable fix, on the
  existing registry, so `doctor --json` gets it for free. Both stay silent when
  the queue is empty.

Three things a review of the queue's edges turned up, fixed here:

- A queue entry nobody can read was skipped on every run, so the warning said
  "1 learning is not published" forever with no reason. It now names the file.
- The hidden-entry filter split paths on the platform separator, while the
  directory walk always joins with `/`, so on Windows a hidden entry inside a
  namespace was published.
- The publisher's docstring claimed it stops at the first failure; every entry
  rides in one commit.

Closes #615
2026-09-18 11:35:20 +08:00
Saul Moro d4c57a0961 feat(learnings): write learnings to the teamai-learnings branch (#616)
* refactor(branch): extract the orphan-branch worktree engine

`reports-branch.ts` managed the `teamai-reports` orphan branch: cold-start
creation across two git versions, stale-worktree repair, a non-blocking lock,
commit, and push with fetch+rebase retry. Learnings need the same machinery on
their own branch (#485), and must not share the branch, the worktree or the
lock.

The engine moves to `branch-worktree.ts`, parameterised by a spec of three
names plus an init commit message. Reports become one instance of it and keep
every exported name, so no caller changes. `reports-branch.ts` keeps what is
not a side branch: `EmptyRepoError` and `withKnowledgeWorktree`.

- A publish returns `PublishResult` instead of a boolean. A caller holding the
  only durable copy of its data must be able to tell "landed on origin" from
  "another writer holds the lock", "nothing to commit" and "rejected". Reports
  map it back to today's boolean, so their behaviour is unchanged.
- A publish confirms the ref actually moved before reporting success, and
  separates "nothing to commit" from "nothing to deliver": an earlier attempt
  may have committed the content and failed to push it. A successful push
  updates the remote-tracking ref locally, so the check is a local rev-list.
  Anything unreadable counts as not landed, which costs one extra push instead
  of losing data.
- `getWorktreeDir` and `getBusinessRoot` come out of the reports-specific
  helpers; `getReportsDir` is now one call to the first.
- `usesReportsBranch` was a one-line forward; its eleven call sites use
  `usesBranchWorktree`, which says what it means now that two branches use it.
- A side branch's `.gitignore` lists every worktree directory, so no worktree
  can nest-track another.
- `ensureReportsDir` had no callers and duplicated `ensureReportsWorktree`.

Refs #485

* feat(learnings): write learnings to teamai-learnings, read them from every root

Contributing no longer touches the default branch, so a member whose team
protects `main` can contribute. Everything the team wrote before the switch
stays readable exactly where it is: nothing is copied, deleted or migrated.

**One accessor.** Eleven call sites built the learnings path by hand, from a
different base each time, so a forgotten one did not fail — it read an empty
directory and recall quietly returned less. `learningsRoots(localConfig)` now
answers with a write root and an ordered read list, and a test fails on a new
hand-built path. Two call sites have no config to resolve (the `viz --repo`
flag and CI) and opt out in place, with the reason on the line.

**Every root is read.** `buildIndex` takes the whole list. For one relative
path the first root wins, deduplicated while collecting: that path is also the
id votes are counted by, so two entries would double-count votes and then have
one silently dropped by recall's dedup. The clone's `learnings/` is always the
last root, which is what keeps the pre-switch corpus searchable.

**One publish path.** `teamai contribute` writes into the `teamai-learnings`
worktree and pushes, for independent clones and single-repo installs alike.
Single-repo mode stops opening a pull request per contribution, so the
disposable knowledge worktree and the machine-wide cache copy that stood in for
it are gone — and with them the cross-project leak of writing one project's
learnings into a cache every project shares. What cannot be published stays in
the durable queue outside the clone and is retried by the next `teamai pull`.

**Maintenance.** Pruning, promotion and confidence write-backs used to mutate a
checkout nothing pushes: the result reached no teammate, and a realign could
undo it. Promotion was worse — the `promoted_to` mark it wrote was discarded,
so the same learning was promoted again on the next run and paid for another
model call. They now write to the write root, publish what they changed, and
refuse to prune an inherited learning instead of deleting a file that comes
back.

Fixes that fall out of reading the same code:

- A recall that had to rebuild the index passed no project namespaces, so every
  project-private learning vanished from the rebuilt index while a pull-built
  one had them.
- Recall printed the absolute path an entry had when it was indexed; a worktree
  that is removed and rebuilt leaves that dangling, and the agent reading the
  output got a dead pointer.
- `pull`'s mirror deletes whatever its source does not have, so it now takes
  every published root except the mirror itself — pointed at one root it would
  have stopped propagating upstream deletions (#458).
- The learnings count in `pull` deduplicates across roots, the way the index
  does.
- Contributing self-heals the single-repo `.gitignore` first, the way `pull`
  and `push` already do, so the new worktree never shows up in the user's own
  `git status`.
- Recall says why a scope was skipped instead of swallowing the reason: an
  invalid projects manifest reported "No learnings available" and nothing else.

Real-git coverage against a bare origin whose `update` hook refuses the default
branch: the learning lands on the side branch with `main` untouched, a refused
branch leaves the note recoverable, a commit whose push failed is delivered on
the next attempt, two members publishing at once both land, one member reads
the other's learning after a refresh, and the inherited corpus is still there.

Refs #485

* docs: document the branch layout and the minimum Git permissions (#486)

A team that turns on branch protection needs to know what still works and what
access its members actually need. The docs said learnings live on the default
branch, and the providers guide said the push target is hardcoded to `master`.

- The data-layout table covers knowledge, learnings, reports and machine-local
  data, and says for each how it is written and whether it needs write access to
  the default branch. EN and ZH match row for row.
- A minimum-permissions section in the usage guide and in the providers guide:
  what a member needs, what they do not, and that `provider: git` still cannot
  open a pull request for you while `contribute` never needs one.
- The directory trees show `learnings-wt/` and `pending-learnings/`.
- The admin checklist says `init` commits an empty `learnings/` and that
  contributions do not go there.
- `docs/providers.md` no longer says the push target is hardcoded: it is
  resolved by `getDefaultBranch()`. The same stale claim in the docstring of
  `pushRepoDirectly` goes with it.
- The share-learnings skill names the branch instead of "directly to master".
- All five READMEs name the branch in the `teamai contribute` row.
- Both design documents describe the new split.

Closes #486

* fix(contribute): keep a contribution findable when nothing can be published

Splitting this change in two left a gap the full version did not have. Before
learnings moved to their own branch the note was written into the clone, so it
was indexed no matter what git did. Now a contribution is written into the
branch worktree — and when that worktree cannot be created at all, the durable
copy was kept but never indexed, so the member could not recall what they had
just written.

- The queue is a learnings root for the index, in `contribute` and in `pull`.
- The durable copy is written before the index is rebuilt, not after, or there
  would be nothing to index.

The queue stays the fallback here; making it the write path is #615.

Refs #485

* fix(learnings): do not call a successful push a failure, and cap how long it can take

A review pass against edge cases found three ways this change could report or
behave worse than what it replaced.

**A successful push read as a failure.** The publish confirmed delivery by
checking that HEAD was no longer ahead of `origin/<branch>`. That ref is only
updated through the remote's FETCH refspec, so in a clone made with
`--single-branch` — what CI checkouts and many business repos are — pushing a
side branch leaves no tracking ref behind and the check said "not landed".
The result was five pushes of the same content and a reported failure although
the data was on origin, and for reports a `save-session` that printed a push
timeout for a write that had worked. A push that resolves is a push the remote
accepted; git exits non-zero when it refuses one. The tracking ref is now only
consulted to answer "does this worktree still owe origin a commit?", where an
unreadable answer means push and find out.

**No cap on how long publishing can take.** `contribute` used to give its push
ten seconds. That cap was lost, so an unreachable origin or a credential prompt
could block the CLI through five push and rebase rounds. Both publish paths
carry it again; timing out is safe because the durable copy stays.

**A warning aimed at a backend with no branch.** Maintenance on an HTTP team
repo has nothing to publish, and said so as a failure. It stays quiet.

Also from the same pass:

- `teamai pull` publishes the queue for every repo kind. It sat inside the
  team-repo refresh, which returns early for single-repo and HTTP, so the
  promise contribute makes — "will retry on the next pull" — was only true for
  independent clones.
- The auto-migration skips `learnings-wt/` along with the other worktrees,
  driven by the shared list so the next worktree is covered without anyone
  remembering this file. `pending-learnings/` deliberately still travels: it is
  work the member has already done.
- `recall`'s own index rebuild reads the queue, so a contribution that could not
  be published does not vanish when anything invalidates the index.
- The learnings mirror carries every visible file in an active namespace again,
  not only Markdown.

Refs #485
2026-09-17 20:53:11 +08:00
Saul Moro 3d4353f45e feat(remove): add --force to skip the confirmation prompt (#594)
`askConfirmation` returns false when stdin is not a TTY, and `teamai
remove` had no flag to skip it, so every scripted run printed
"Cancelled" and exited 0. The command could not be used from a script,
and it had no end-to-end test, which is how both defects in #576
survived.

`--force` is spelled and described the same way as `teamai uninstall
--force`, and the prompt keeps the same guard shape so the two stay
refactorable together. The command's action was also dropping its
`cmdOpts` argument, so the option is wired through the way `uninstall`
does it.

The new end-to-end test drives the built CLI against a bare repository
and asserts the whole contract. The deployed copy goes immediately, and
the deletion plus its tombstone are published as a branch for review.
`checkoutMaster` returns the clone to the default branch, so the clone's
own working tree keeps the file until that branch merges.

Fixes #591
2026-09-17 12:08:00 +08:00
Saul Moro 0982976baf feat(doctor): export the check registry and add --json (#599)
The checks that catch "reported success, nothing on disk" lived inside
doctor() as a local array, so nothing else could run them, and the only
machine-readable result was the exit code #569 added.

Extract resolveDoctorContext() and buildChecks(), which render nothing,
and add --json: one object on stdout, every log line on stderr, exit code
unchanged. Human output is byte-for-byte what it was.
2026-09-17 11:55:06 +08:00
Saul Moro c674ffe9b8 fix(remove): honour enabledAgents when removing rules and skills (#592)
`removeItem` in `RulesHandler` and `SkillsHandler` deleted the resource
from every tool in `toolPaths`, without checking `isAgentExcluded`. A
member who whitelists `enabledAgents: ["claude"]` still lost the Codex
copy. The tombstone pass in `teamai pull` checks that whitelist, so the
two sides of the same removal disagreed.

PR #589 added the gate to the agents handler while closing #576, and
left these two out so the tombstone regression stayed reviewable. This
applies the same gate to both.

Gating the two removal loops is not enough on its own. `RulesHandler
.removeItem` ends by calling `pullAllRules`, whose stale-file sweep
iterates every tool and deletes any local rule missing from the team
set, including the one just removed. `pullItem` in the same file already
skips excluded tools, so the sweep now does too. Without that, the gate
in the removal loop is bypassed whenever the team still holds another
rule, which is the normal case. The sweep also runs during `teamai
pull`, so an excluded tool keeps its rule files there as well.

For skills the check sits above the OpenClaw branch, so an excluded
tool's workspace copy and Codex's shared `.agents/skills` destination
are both covered.

`MCPHandler.removeItem` needs no gate. It only rewrites the team repo's
`mcp.yaml` and never touches a tool directory.

Fixes #590
2026-09-16 22:33:10 +08:00
Saul Moro d2d1f78d85 fix(agents): clean up tombstoned agents under every render extension (#589)
`teamai remove agents <name>` deleted the `.md`, `.toml` and `.json`
renders on the machine that ran it, but the tombstone pass in `teamai
pull` tried only `.md`. The Codex `.toml` and Kiro `.json` copies of a
removed agent stayed on every other machine.

The cause was duplicated extension knowledge. `removeItem` and the
tombstone pass each carried their own list, and the two drifted.
`rule-format.ts` already documents the same split for rules, a per-tool
function for writers and one shared list for scanners and deleters.
Agents now get that split through `AGENT_FILE_EXTENSIONS`, and the three
other sites that hardcoded the same triple read it too.

The same code had two further gaps:

- `removeItem` ignored `isAgentExcluded`, so `teamai remove` deleted
  agents from tools the member had excluded through `enabledAgents`.
- `pull` returns early when the team repo rev is unchanged, before the
  tombstone pass. Every machine this bug affects is in that state,
  because it already pulled the tombstone with the older CLI, so the
  upgrade never reached it. The cleanup now also runs on that path,
  next to the built-in deploys that are there for the same reason.

Fixes #576
2026-09-16 20:41:10 +08:00
Saul Moro cccbe1d5aa feat(agents): scope team agents by role and project namespace (#577)
* feat(roles): add agents resource namespaces to roles and projects manifests

Optional `agents:` key on roles.yaml and projects.yaml resources, merged
into the active namespace set. Absent means no agents namespaces, which
keeps today's behaviour for existing manifests. Part of #563.

* feat(agents): scan one level of agents/<namespace>/ in the team repo

Root files stay shared. Subdirectory files carry a namespace like
rules/<namespace>/ so pull can filter them by role. Part of #563.

* feat(pull): filter team agents by active namespaces and reject stem collisions

Root-level agents ship to everyone; agents/<ns>/ ships only when <ns> is an
active agents namespace, mirroring filterRulesByKnowledgeNamespaces. Two
kept agents with one stem would overwrite each other on disk, so pull
fails loudly like scanRoleAwareSkills does for skills. Part of #563.

* feat(agents): remove deployed agents of namespaces that stop being active

A role or project change now revokes the previous namespaces' agents on
every installed tool. A deployed file is deleted only when it is byte-equal
to what pull renders from the team source; a local edit is kept and
reported, the same data-safety rule inactive skills follow. Part of #563.

* feat(agents): locate team agents by stem across namespaces in push, remove and uninstall

A modified agent whose source lives in agents/<ns>/ is written back there
instead of creating a root duplicate; remove and uninstall find namespaced
sources the same way. New agents still land at the root. Part of #563.

* docs: describe role-scoped agents in both READMEs, usage guides and the changelog

* refactor(agents): share the agents directory walk and drop dead nullish guards

Review follow-ups: one listTeamAgentDirs used by scan, stem lookup,
uninstall and the self-mode pickup (which now sees agents/<ns>/ too);
removeItem iterates every match once; the roles/projects resolvers rely on
the zod default for agents instead of runtime guards; design doc example
gains the optional agents key.

* fix(agents): resolve active sources and clean up per tool
2026-09-16 15:38:26 +08:00
Saul Moro f7c7bdd58c feat(mcp,hooks): scope team MCP servers and hooks by role (#578)
* feat(mcp,hooks): parse an optional roles list on mcp.yaml servers and hooks.yaml hooks

Carried through to McpServerDef and HookDef; not yet applied. Part of #563.

* feat(roles): add activeRoleIds and matchesRoles for per-entry role filters

Shared by the MCP and hooks reconcilers. Omitted roles match everyone, an
empty list matches nobody (like tools), no configured role means no
filter, the same fallback skills and rules already use. Part of #563.

* feat(mcp): apply the roles filter when reconciling team MCP servers

A server with roles: ships only to members whose primaryRole or
additionalRoles it lists; a role change removes it on the next reconcile
through the existing desired-set rebuild. Unknown role ids get one warning
per id so a typo does not hide a server in silence. Part of #563.

* feat(hooks): apply the roles filter when resolving team hooks

resolveTeamHooks drops hooks whose roles: does not list one of the member's
roles, before the security gates so the transparency print shows only what
will run. A role change removes the previous role's hooks through the
existing managed-slice rebuild. Unknown role ids warn once. Part of #563.

* feat(mcp,hooks): show the roles restriction in mcp list and hooks list

* docs: describe the roles field for MCP servers and hooks in both guides, READMEs and the changelog

* refactor(roles): review follow-ups for the roles filter

Warn once per process for an unknown role id so a member with user and
project scopes reads it once per pull; matchesRoles accepts an undefined
active set; list commands print nobody for an empty roles list, matching
the docs; docs example uses the devops role the guide already defines.
2026-09-16 11:05:50 +08:00
Saul Moro 8efeaf73d2 fix(dashboard): normalize Unicode correction keywords (#575)
* fix(dashboard): normalize Unicode correction keywords

* docs: limit guide changes to Unicode matching
2026-09-16 10:39:32 +08:00
Saul Moro d13780c7e5 fix(dashboard): match correction keywords as whole words, add team keywords (#567)
* fix(dashboard): match correction keywords as whole words, add team keywords

`isCorrectionPrompt` used raw substring matching, so the built-in `undo` and
`redo` fired on ordinary Spanish and Portuguese words ("segundo", "mundo",
"redondo"). One false correction scores 20, which is the share-learnings nudge
threshold on its own.

Keywords in a space-separated script now match as whole words (Unicode-aware,
so accented letters count as letters). Keywords containing Han, Hiragana,
Katakana or Hangul keep substring matching.

Teams can add their own words via `sharing.intervention.correctionKeywords` in
teamai.yaml. The prompt_submit hook resolves them and stores a `correction`
flag on the event, because the machine-level events file mixes sessions from
every team. Events without the flag fall back to the built-in list.

For #564

* fix(dashboard): resolve team keywords from the hook cwd, treat _ as a word char

Cursor runs hooks from ~/.cursor and sends the project in workspace_roots, so
resolving the team via autoDetectInit() (process.cwd()) silently dropped the
team's correctionKeywords there. Resolve the project from resolveHookCwd(stdin)
and fall back to the user-scope config.

Underscore joins the word boundary so identifiers such as "test_undo" do not
count as `undo`. Drop the regex cache and fold both keyword lists into one
`.some`. CHANGELOG separates the whole-word fix from the team-keywords feature
and notes the new `correction` event field; both usage guides say the built-in
list still covers only zh/en/ja.

* fix(dashboard): keep autoDetectInit for team keywords in the prompt hook

hook-dispatch-cli already chdir's to the hook payload's cwd before running
handlers, so autoDetectInit() resolves the right project on Cursor too. The
explicit detectProjectConfig(resolveHookCwd(stdin)) path added in the previous
commit was redundant; drop it and its tests.
2026-09-15 21:49:13 +08:00