A busy event loop stretched the usage rewrite wait past its 5s budget because the wait counted sleeps. Bound it by wall time. The viz fixture also cloned a bare repo whose HEAD stayed on master under git 2.39, so worktree add failed with "invalid reference: HEAD".
Co-authored-by: Cursor <cursoragent@cursor.com>
oxlint --deny-warnings flags /Error$/ as a warning, so every main CI run has failed since Swift support landed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(review): scope e2e requirement, drop provider×agent matrix, add P3
The auto-review gate over-flagged: it required a full provider × agent
e2e matrix as [P1 blocking] on every PR, raised theoretical/low-confidence
risks at blocking severity, and repeated already-resolved findings.
- e2e record is [P1 blocking] only for behavior-changing PRs; docs-only /
tests-only diffs need none. One representative real-CLI run is enough;
missing extra provider/agent coverage the author flagged as untestable
here or deferred to CI is at most [P3 nit], never blocking.
- Add a third severity [P3 nit] (minor/optional polish, theoretical edge
case, coverage deferred to CI).
- Tell the reviewer not to over-review: report only high-confidence
findings, prefer few high-signal over exhaustive, never restate a
resolved finding.
Kept in sync across AGENTS.md (## Code Review Rules, the trusted base
criteria) and the workflow prompt in codex-review-on-assign.yml.
* chore(review): relax PR 前测试 matrix, gate P1 by evidence not count
Address SaulMoro's changes-requested on #781:
- The provider × agent matrix requirement still lived in `## PR 前测试`
(AGENTS.md/CLAUDE.md), which the Codex reviewer loads whole — so the
matrix could return as a P1 despite `## Code Review Rules` relaxing it.
Relax `## PR 前测试` to match: one representative real-CLI run suffices,
docs/tests-only needs no e2e, extra provider/agent coverage defers to CI.
- Replace "prefer a few high-signal findings" (which caps real bugs and
pushes them to later passes) with evidence-gated severity: every
[P1 blocking] must cite a concrete failure scenario or the exact rule it
breaks, else it drops to P2/P3. Count is uncapped; real bugs surface in
one pass. Synced into the workflow prompt.
---------
Co-authored-by: review <review@local>
ClawHub's SkillSpector scan flagged three doc-level issues in the teamai
skill guidance:
- PE3 (Credential Access): the GitLab setup showed a literal
`export GITLAB_TOKEN=glpat-...`, which lands in shell history and process
listings. Switch to a no-echo prompt, recommend a short-lived api-scope
token, and unset it after init.
- SQP-2 / P4 (silent install + behavior manipulation): the TGit login guidance
told the agent to run install/login itself and 'never tell the user'. Require
disclosing that it installs a binary and stores a credential, and getting the
user's OK first (or showing the commands if they prefer).
- SDI-4 (trigger ambiguity): publishing a skill from a plain-language request
had no confirmation gate. Require showing what will be shared and getting an
explicit go-ahead before teamai push / contribute.
Doc-only; edits land in skill-data/ (served by `teamai skill get`).
The gf CLI installer downloaded the platform tarball over plaintext HTTP
and piped it straight into tar with no integrity check, then executed the
extracted binary. ClawHub's security scan flagged this (finding T03).
Fetch over HTTPS to a unique temp file, verify its SHA-256 before extracting,
then clean up. The mirror 302-redirects to a content-addressed backend whose
URL path is the artifact's sha256 (the backend does not echo a checksum
header), so the expected digest is read from the redirect target, falling
back to x-checksum-sha256 for a direct serve. Fail closed when no digest is
advertised. Update the TGit provider reference doc to match.
Verified end-to-end on darwin-arm64: download, digest match, extract.
Put the product overview on the landing page so visitors can see Git-based Execution / Context / Improvement and per-agent coverage without opening a secondary doc.
Closes#731. Feed current PR body plus earlier review comments into
each pass, stop hand-listing config mocks, and run credential-free e2e
on fork PRs without exposing fixture tokens.
Co-authored-by: review <review@local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Step 9 of setup-admin.md told users to run
`/teamai 我想把学到的经验分享给团队` to share what they learned. This
contradicts the teamai SKILL.md, which states twice that sharing a
session's learnings is automatic and must NOT be routed through
`/teamai` — it is handled by the separate teamai-share-learnings skill.
Replace that bullet with the correct guidance:
- `/teamai` share entry is for publishing a reusable skill
(contribute-member.md), not loose learnings.
- Session learnings are shared automatically via teamai-share-learnings.
* fix(test): eliminate CI flaky tests
- lock-atomic: accept 1–2 winners under high-load concurrent reclaim
instead of exactly 1 (fs rename window widens under load)
- dashboard-collector: refresh `now` in beforeEach so timestamps are
never stale when rebuildSessions checks the 30s expiry window
- codebase-reconcile, import-dir: increase timeout from 15s to 60s
for tests that hit disk-heavy operations on slow CI
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(test): address review — keep strict lock assertion, reduce concurrency
Preserve exact-one mutual-exclusion assertion per reviewer feedback.
Reduce concurrency from 32→8 to lower CI load pressure, and add
retry:3 so transient fs scheduling races don't block the pipeline.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(test): revert lock-atomic changes, keep original assertion
Revert lock-atomic to its original form (32 concurrency, strict
exact-one assertion, no retry). The sentinel serialization is
correct under Node.js single-threaded event loop; the rare 2-winner
case is an OS/fs-level anomaly under extreme CI load, not a product
bug. This PR focuses on the deterministic fixes for the other 3
test files.
chmod 0o000 has no effect for root (CI runs as root), causing both
pending learnings to be published and the cleanup chmodSync to ENOENT.
Skip the test on root and guard the finally block with existsSync.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Consolidate the Tencent TGit (工蜂) flow — reachability probe, gf
install/login, and init repo-creation behavior — into a single
references/provider-tgit.md, which both setup-admin.md and join-member.md
now point to instead of repeating the probe twice and re-describing the
gf login across files.
Also dedup the "don't trust the 'Hooks injected into all AI tool
settings' message" caveat down to the canonical table in
troubleshooting.md, and merge the two overlapping auto-share blockquotes
in SKILL.md into one.
No command or flag changes — cheat sheet stays verbatim against src.
#701: The webhook handler forwarded the entire hook stdin as `data`, so
wildcard subscriptions leaked tool args (mock API keys) and full
tool_response. Forward only a per-event field whitelist (skillName /
sessionId) and route the outbound body through the shared redact module
as defense-in-depth. Skill-name resolution goes through the shared
resolveSkillUse helper, which validates with isValidSkillName, so an
arbitrary tool-arg string cannot escape through it, and a normal
(non-SKILL.md) Read never produces skill data.
#702: The handler read `stdin.event`, which hosts never send, so every
event was forwarded as `unknown` and skill-use / session-start /
session-stop subscriptions never matched. Map `hook_event_name` to the
canonical event names (case-insensitive, so both Claude's PascalCase and
Cursor/CodeBuddy's camelCase resolve) and drop unmapped events instead of
emitting `unknown`. Register the webhook handler on session-end too, so
Copilot's SessionEnd emits a session-stop notification. Cursor represents
skill use as a `Read` of a SKILL.md file dispatched under the `Skill`
matcher (git f0ab4eb switched Cursor tracking from Read to the Skill
matcher); the webhook now reaches parity with the usage tracker by sharing
resolveSkillUse, which handles both the Skill and Read+SKILL.md shapes.
Wire the promised `push` / `pull` command events from their CLI actions,
gated on a real completion signal so dry-run / cancel / no-change /
handled-failure runs never fire a misleading notification — pushGroup
reports a distinct outcome ('pushed' | 'nochange' | 'pr-failed' |
'failed'). A reuse branch recorded with prUrl:null (an earlier PR creation
failed) now retries PR creation instead of silently skipping it — including
when the tree is unchanged (hasChanges false), where it (re)creates the PR
for the already-pushed branch without re-pushing.
#703: getWebhookSharing dropped `secret` and force-overrode
timeout/retries, so a configured signing key never reached the request
and receivers with signature verification rejected the unsigned webhook.
Preserve all schema-accepted fields and sign the exact bytes sent.
docs: Document sharing.webhooks (config shape, events + timing, whitelisted
payload, signature) in docs/usage-guide.md and docs/usage-guide.zh-CN.md.
Co-authored-by: review <review@local>
Keep README as title, Why TeamAI, quick start, docs links, contributors, and contributing. Move architecture and capability details into product-overview, and the command table into the usage guide.
#704: An independent teamai-learnings update never moves main's revision, so
the "Already synced" fast-return in pull skipped the learnings-branch refresh
and index rebuild — a teammate's contribution only surfaced after `pull --force`.
Hoist the learnings-sync + index-rebuild into a helper and run it on the
fast-return path too. It is read-only (pushIfCreated:false) and never publishes,
so it stays inside the caller's partition sync lock and does not reintroduce the
flush-outside-lock bug.
#705: contribute indexed pending-learnings, then published + deleted the pending
files without refreshing the index, so recall handed back a File path under the
now-emptied pending dir. Rebuild the index once more after a successful publish
so every published entry resolves to its durable worktree copy.
#706: A --single-branch clone's fetch refspec covers only the default branch, so
`fetch origin <branch>` moved only FETCH_HEAD and left origin/<branch> absent,
then `worktree add --track` failed and the knowledge worktree was never created.
Fetch the tracking ref with an explicit refspec, and check the worktree out with
`--no-track` (nothing relies on git upstream config; every sync references
origin/<branch> directly). Shared helper — reports branch verified not regressed.
Regression tests: real-CLI e2e for #704/#705 (learnings-sync-704-705) and a
real-git single-branch checkout test for #706 (git-kind-learnings). Both confirmed
to fail on the pre-fix code and pass after.
Co-authored-by: review <review@local>
* docs(teamai-skill): drive GitHub login from the assistant, not the user
The teamai skill's GitHub login step was a bare `gh auth login`, unlike the
TGit/CNB steps that spell out "you run the login; the user only approves in the
browser." That asymmetry made the assistant hand the raw command to the user
("run this with the ! prefix"), which non-technical users can't do.
Rewrite the GitHub step in join-member.md and setup-admin.md to match TGit/CNB:
the assistant runs the login (or lets `teamai init` auto-start `gh auth login
--web`, which it already does), relays the device code + URL, and the user's
only action is approving in the browser. Add the headless GITHUB_TOKEN escape
hatch for parity.
* docs(teamai-skill): make 'log in first' the GitHub step, not init auto-trigger
Drop the 'let teamai init auto-start gh auth login' path as the primary route.
The GitHub step now mirrors TGit: the assistant runs `gh auth login --web`
first, as an explicit step, and the user only approves in the browser.
The Codex review posts findings tagged bare 'P1'/'P2', but the review
rules never defined what those codes mean, so first-time readers cannot
tell a blocking issue from a suggestion. Add a Code Review Rule requiring
each finding's severity to be spelled out inline (keeping the P marker):
[P1 blocking] must-fix before merge, [P2 non-blocking] suggestion, in the
PR author's language.
Co-authored-by: review <review@local>
Wrap Install / Team admin / Team members in a <details> block (collapsed by
default) so Quick Start leads with the /teamai flow. The usage-guide link
stays outside the fold since it's useful to everyone. Applied to
en/zh-CN/ja/ko/th; content unchanged, only wrapped.
Co-authored-by: review <review@local>
* feat(dashboard): unify workspace views, themes and localization
* fix(dashboard): address review, decouple cache metric, fix session attribution
Review fixes (@laolaoPlayer):
- #1 self-mode workspace read the wrong KB: route dashboard workspace
discovery through the exported readConfigFrom so the self-mode repo
rebind is applied; log the real /api/context error instead of swallowing.
- #3 workspace membership frozen at startup: recompute on a short TTL and
route unmatched sessions to a dedicated "unassigned" bucket instead of
silently inflating User scope.
- #4 add workspaceEvents unit tests and a ?workspace= scoping e2e (project
scope + linked worktree + unassigned).
- #5 delete the drifted, unreferenced demos/dashboard/ prototype.
- #6 harden report i18n: tag every report title with data-i18n, drop the
brittle regex special-cases, add a guard test that every data-i18n label
has a translation.
Cost/metric accuracy:
- feat(pricing): modelAliases config maps gateway model aliases (e.g.
ep-qxst1hw4) to known Claude models so cost estimation works behind a
gateway; falls back to the built-in table when unset.
- fix(trends): decouple cache-read share from pricing — derive it from the
session's own transcript tokens so it shows even when the model can't be
priced.
Session-data fixes:
- fix(collector): drop UserPromptSubmit events that are purely injected
content (task-notifications, system-reminders, interrupt markers) so they
no longer inflate the prompt count or appear as prompts.
- fix(collector): dedupe cross-tool duplicate events at the readEvents
boundary — a host (e.g. Cursor) that also loads claude's hooks double-fires
every event; collapse the pair, keep the specific host tool. Fixes Cursor
sessions being labelled claude and turn counts doubling.
UI: drop the "All local workspaces" option, the sidebar accent dot and the
bottom "TeamAI Dashboard" text; remove the low-signal "Usage & sessions"
panel row and the "Active duration" / "Session success" trend cards; rename
"团队执行" to "团队执行环境" (zh only).
Move each /teamai prompt (and the bootstrap line) into its own fenced code
block so GitHub renders a copy button — blockquotes and inline code have
none. Labels become bold headings above each block. Applied to
en/zh-CN/ja/ko/th; prompt text unchanged.
Co-authored-by: review <review@local>
Replace the top of Quick Start with a one-line bootstrap prompt (load the
teamai skill from its repo URL, then set up a team) followed by the four
/teamai flows: set up from scratch, join, share a skill, open the dashboard.
Applied to en/zh-CN/ja/ko/th; the /teamai lines, URLs, and commands stay
verbatim while the surrounding prose is localized per file.
openai/codex-action@v1 now resolves to v1.12+, which rewrote the
codex-exec runner to wait on the whole process tree's stdio. The
backgrounded Responses API proxy keeps those descriptors open, so the
step never sees stdio close and hangs until timeout — the review is
produced (final-message is printed) but the job idles for hours and the
Post-comment step never runs, so the review is thrown away.
Pin to v1.11, the last release before the regression (identical
final-message output contract), and add timeout-minutes: 15 to the
review job so any future hang is capped instead of idling for the
default 6h.
Refs: openai/codex-action#151, openai/codex#9269
* feat(skills): add teamai onboarding skill (#572)
Publish a standard SKILL.md-format `teamai` skill that interactively
guides users (including those unfamiliar with Git) through setting up,
joining, managing, and contributing to a team AI repo.
- Manual-trigger only (/teamai); progressive-disclosure menu, no action
on bare invoke
- Scenario references: admin setup, member onboarding, daily management,
contribute, uninstall, plus a shared troubleshooting guide
- Replies in the user's language; agent-specific hook caveats
(Codex trust-gate, Cursor, CodeBuddy/WorkBuddy, ChatGPT App)
- Privacy note: no third-party data reporting; only the team repo
Supersedes #517.
* feat(skills): address review feedback on teamai skill (#572)
- Point 'contribute learnings' to the dedicated teamai-share-learnings
skill; refocus contribute-member.md on publishing a reusable skill
- Spell out daily management as publish/update skills, rules, MCP, env;
add MCP, projects, and dashboard sections to manage-admin.md
- Add a dashboard menu entry / cheat-sheet line (teamai dashboard)
- Don't restrict install to one agent: set up all installed AI tools by
default and report which agents were configured (new global rule 9)
- Add projects to the cheat sheet (manage several projects from one repo)
* docs(skills): address teamai skill review — TGit provider, member access, tighter description (#572)
- SKILL.md: trim description to <100 tokens; keep the "invoke only on
explicit /teamai, never auto-trigger" constraint.
- setup-admin.md: add Tencent TGit (工蜂) as a first-class, top-listed
platform when a request to git.woa.com returns the `x-env: tgit` header.
Covered in 2a/2b probes, the 2c create-repo table, and a new `gf auth
login` step (teamai auto-installs the gf CLI; no GITLAB_URL needed).
- setup-admin.md: new Step 7 — admin must grant each member read/write
access on the platform before handoff, or their init/pull/push fails
(teamai has no permission model of its own). Renumber later steps.
- join-member.md: add the matching git.woa.com → `gf auth login` entry.
* docs(skills): fix member contribute hints — auto-share is automatic, any member can publish a skill (#572)
The member-facing "what's next" hints were wrong on two counts:
- Sharing session learnings is NOT a manual `/teamai` menu choice — it is
auto-triggered by the Stop hook (contributeCheckHandler) at session end and
gated by the admin's team-sharing toggle (sharing.contributeHint.enabled,
on by default). Reworded the menu, routing note, join-member wrap-up, and
contribute-member to say so, and stop offering a `/teamai` line for it.
- Any member (not just admins) can publish a skill, by asking in plain
language ("share this xxx skill with my team"). Added that as the member
menu row / routing entry and stated it in contribute-member.
Also documented the admin on/off toggle + resolution order in manage-admin
(verified against isContributeHintEnabled; contribute-hint-toggle tests pass).
* docs(skills): TGit needs no manual login — init auto-installs gf and prompts (#572)
The 工蜂 login instructions were wrong: they told the user to run
`gf auth login` in the agent. In fact the tgit provider does it all inside
`teamai init` — ensureInstalled() auto-downloads the gf CLI and authenticate()
launches the interactive login mid-init when the user isn't authorized yet.
- setup-admin Step 3 / 2c note / Step 5 caveat: drop the manual `gf auth
login` command; state that init installs gf and handles the login on its
own (user only approves in browser / iOA when prompted).
- join-member Step 3: same fix for the member side.
- Keep TGIT_TOKEN only as the headless/CI escape hatch.
* docs(skills): document optional gf pre-install using teamai's own method (#572)
Add an optional path for the agent to install the gf CLI before init, using
the exact same source, install dir, and verification teamai uses internally
(gf-cli.ts) — not a hand-rolled URL:
- source: http://mirrors.tencent.com/repository/generic/gongfeng-cli/.../gf-<os>-<arch>.tar.gz
- dir: ${TEAMAI_HOME:-~/.teamai}/gf ; verify: test -x <dir>/gf/bin/gf
- darwin|linux × x64|arm64 only
Login stays automatic (init launches it); this only lets you pre-fetch gf.
join-member points to the same commands rather than duplicating the script.
Verified end-to-end on darwin/arm64 (download + extract + test -x all pass).
* docs(skills): agent installs gf + logs in; prefer auto-creating the TGit repo (#572)
Per review: stop describing gf install/login as something teamai/init does
automatically. The agent drives it.
- setup-admin Step 3 (TGit): the agent installs gf (teamai's own download +
test -x verify) AND runs `gf … auth login` (verified: binary + auth login
subcommand exist); user only approves the browser/iOA prompt. Drop the
"init auto-installs/auto-logs-in" wording (2a note, 2c note, Step 5 caveat).
- Repo creation: prefer letting `teamai init` create the TGit repo via the API
(init offers the create prompt and calls provider.createRepo); only fall back
to git.woa.com/projects/new when the group/namespace is missing or the user
lacks create permission.
- join-member: member side also has the agent install gf + log in, reusing the
setup-admin commands.
* docs(skills): agent runs gf install AND login itself, user only approves (#572)
Per review: never tell the user to run a gf command. The agent runs both the
install and `gf … auth login`; the user's only action is approving the login in
the browser / iOA.
- setup-admin Step 3 (TGit): reword to "YOU run gf install and login"; describe
gf auth login's real interactive flow (iOA / browser device code / token —
verified via `gf auth login --help`); confirm with `gf auth whoami`.
- join-member: same — agent runs both, user only approves the URL.
---------
Co-authored-by: review <review@local>
The auto re-review added in #631 never worked for fork PRs. Our gate job
correctly authorizes a re-review when the PR carries a maintainer assignee,
but codex-action then runs its OWN write-access check against the triggering
actor — which on a fork PR's `synchronize` is the PR author (read access).
So every fork auto re-review failed with:
Actor '<author>' is not permitted to run this action ... Detected 'read'.
Pass `allow-users: "*"` to disable codex-action's actor check. This does
NOT widen who can trigger a review: our `gate` job is the real, stricter
authorization door (an un-authorized PR never reaches this job at all), and
codex-action's check is redundant with it while being wrong for the fork
push case. All secret-protecting layers are unchanged: gate authorization,
trusted base checkout, PR code read as diff data only (never built/installed/
run), persist-credentials: false, and Codex staying :read-only.
Previously a maintainer had to re-assign the PR after every push to get an
updated review. Now assigning once authorizes the PR, and later pushes
re-review automatically.
- Add synchronize + ready_for_review to the trigger.
- The gate no longer keys off who triggered the run (a push's actor is the
fork author, who would always fail the check). Instead:
* assigned -> the assigner must be a maintainer (unchanged intent)
* synchronize / -> the PR must already carry a maintainer assignee,
ready_for_review i.e. it was authorized by an earlier assign
- Skip re-review on pushes to a draft PR; wait for ready_for_review.
Security is unchanged: an un-authorized fork PR (no maintainer assignee)
never runs on its own pushes, the workflow still comes from the base branch,
PR code is still read-only diff data, and Codex stays read-only.
Gate logic verified locally with a 10-case truth table covering authorized/
unauthorized assigns, re-review with/without a maintainer assignee, and the
draft skip.
The pull_request trigger never receives secrets on fork PRs, so the Codex
review silently no-op'd for external contributions (the majority of PRs).
Switch to pull_request_target so fork PRs can access the API key, and add
the safeguards that trigger requires:
- A maintainer gate job: only a user with write/admin/maintain permission
can trigger the review; unauthorized assigns skip it. This stops
strangers from burning the API budget or exercising the token.
- Check out the trusted BASE commit, never the PR head. The PR changes are
exposed to Codex only through git diff — read as passive data, never
built, installed, or executed (npm install alone would run a fork's
pre/postinstall hooks).
- Review rules come from the base AGENTS.md, so a PR cannot rewrite its own
criteria.
- persist-credentials: false keeps the repo token off disk.
Codex stays read-only throughout.
Add a GitHub Actions workflow that runs an automated Codex review when a
pull request is assigned. Codex reads the PR diff in read-only mode,
follows the new AGENTS.md "Code Review Rules" section, and posts its
findings as a PR comment in the PR author's language.
The review flags bugs, rule violations, and PRs whose description lacks a
test plan or an end-to-end record. PR title/body/diff are treated as
untrusted data to guard against prompt injection. The Responses API
endpoint and review model are optional, configurable via secret/variable
and fall back to Codex defaults when unset.
Co-authored-by: review <review@local>
The ClawPro project-binding prompt fired for every host that runs the
teamai hook — Claude, Cursor, Codex included — even though ClawPro
project binding only backs CodeBuddy/WorkBuddy. Those other users got a
"[ClawPro项目 绑定提示]" choice list injected into their session with no
way to act on it, which is the poor UX being reported.
Gate the whole prompt (both the SessionStart TTY prompt via
ensureWorkspaceBinding and the UserPromptSubmit hint via emitBindingHint)
on the current tool being a buddy agent, at the single `reportAndSyncLocalAgent`
entry point. Non-buddy tools now short-circuit before resolving the
workspace, so they fork no git process and never fetch /projects/mine.
Reuses `modelAgentKind` for the check so tool-name variants like
`codebuddy-internal` still match (a raw Set would miss them).
Tests: existing bind-hint cases retargeted to a buddy agent so they keep
their discriminating power; added a parametrized case asserting claude
and cursor emit no hint and skip the project fetch. Full suite green
(3307). Verified end-to-end against a mock backend with the built CLI:
codebuddy/workbuddy/codebuddy-internal inject the hint; claude/cursor/codex
produce no output.
src/doctor.ts imported LocalConfig twice — once in the top type-only
import (line 5) and once in the grouped import block from './types.js'.
TypeScript rejected this with TS2300 (Duplicate identifier 'LocalConfig'),
so `tsc --noEmit` failed and the main CI has been red since #599.
The two imports landed cleanly as a merge (no textual conflict) but
collided at the type level, so each PR's own branch build was green.
Drop the LocalConfig binding from line 5 and keep it in the grouped
block alongside TeamaiConfig, matching the surrounding style.
* ci: add informational code-erosion (slop metrics) workflow
Report SlopCodeBench verbosity/erosion metrics on every PR using the
official scb-check tool, pinned to 0.2.0 (the first release with
TypeScript support; 0.1.3 is Python-only).
The workflow is informational and never blocks a merge: scb-check's exit
code is swallowed, and the numbers are posted as a deduplicated PR comment
with the run's job summary as a fallback (so fork PRs, whose token is
read-only, still surface the report). Tests are excluded via scb-check.toml
so metrics reflect the product surface.
On TypeScript the ast-grep verbosity rule component is Python-only and
contributes 0, so verbosity reflects clone + wrapper detection only;
erosion is fully faithful. This caveat is documented in the bilingual
docs/ci-code-erosion.{md,zh-CN.md}.
* ci(code-erosion): add independent TS verbosity rule layer
scb-check only runs its ast-grep rules on Python files, so on this
TypeScript repo its verbosity rule component is always 0. This adds a
standalone ast-grep pass with a small, hand-ported rule set to fill that
gap, reported as a separate "Rule hits (TS verbosity layer)" section in
the same non-blocking PR comment.
Only purely structural rules are ported. Rules that hinge on truthiness or
type semantics (len==0, ==True, redundant template strings) were tried and
deliberately dropped: they are false positives in TypeScript, where
arr.length>0 is idiomatic and x!==true is not equivalent to x===false
(TS has undefined). Ported rules verified against src/ for false positives:
unnecessary-else-after-return, empty-catch-block, redundant-ternary-same,
if-return-boolean-literal, return-ternary-boolean-literal,
duplicated-if-condition, self-assignment.
Rules use severity: hint and the scan step has `|| true`, so the layer
never blocks CI. Bilingual docs updated with the honest scope: this is an
extra signal, not a reproduction of the paper's verbosity number.
* ci(code-erosion): slim down the PR comment, defer detail to docs
The comment carried long inline explanations (verbosity footnote, rule-layer
paragraph). Move the prose to docs/ci-code-erosion.md and keep the comment to
numbers plus a one-line pointer. Also replace the ambiguous "(informational)"
tag with plain "never blocks the merge".
* ci(code-erosion): drop the two repeated doc links in the comment
The top line already points to docs/ci-code-erosion.md; the verbosity
footnote and rule-layer note repeated the same link. Keep one pointer.
* fix(partition): adopt pre-#546 legacy-named partitions instead of stranding them
#546 widened the partition slug prefix from the anchor's basename to its
whole path but kept the sha256 suffix — without migrating existing
installs. On an upgraded CLI, a partition still named <basename>-<hash>
(every install created before #546) stops matching projectSlug(anchor):
detectProjectConfig finds no config (the project looks uninitialized),
planMigration would re-copy a retired workspace into a second, empty
partition, and status --all flags the perfectly good data as "corrupt —
dir name does not match anchor".
Both formats share the same hash, so an anchor's legacy name is
computable exactly — no directory scanning. resolvePartitionDir becomes
the seam for "the partition that actually holds this project's data"
(detection, init, migration): when the canonical current-format
directory is absent and a legacy-named one exists, it ATOMICALLY RENAMES
the legacy partition into place — a same-parent metadata move, no data
copied, and an interruption leaves either name intact. An authoritative
current-format partition is never clobbered by a leftover legacy one; a
rename that is genuinely impossible (read-only home) keeps serving the
legacy directory so no data is stranded. status --all stays read-only
and reports a legacy-named partition as "active (legacy name; renamed
automatically on next command)" instead of corrupt.
Verified end-to-end with the real CLI in an isolated HOME: a legacy
.teamai migrates into the readable whole-path slug; status --all
reverse-resolves it as [active]; a partition renamed to the pre-#546
format is reported active-legacy (no rename, no corrupt), then adopted
by the next command (detection) with data intact, and pull runs through
the adopted partition.
* fix(partition): rebase repo.localPath when adopting a legacy partition
A pre-#546 partition stores repo.localPath as an ABSOLUTE path to its
team-repo clone (<legacyPartition>/team-repo). resolvePartitionDir renamed
the directory but left config.yaml pointing at the now-gone old path, so
every later `pull` read the team config from a dead directory and silently
skipped the sync ("Team config (teamai.yaml) not found. Skipping.", exit 0)
— the project could never sync again after an upgrade.
Adoption now rebases repo.localPath onto the new partition (mirroring
migrate.ts's rebaseConfigPaths). The rewrite is idempotent and self-healing:
a modern install whose localPath already sits in the canonical dir is left
untouched, an external clone outside the partition is left untouched, and an
adoption interrupted between the rename and the config rewrite is finished by
the next command.
Reproduced with the real CLI (isolated HOME, real local git team repo): a
legacy-named partition + two `pull --force` runs printed "Team config not
found. Skipping." (exit 0) before this fix; after it both runs print
"Synced 1 skills" and localPath points at the live clone.
Regression tests: localPath rebased off the legacy dir / external localPath
left alone / modern-install no-op. Verified they fail when the rebase is
disabled.
* fix(partition): write the adopted config.yaml atomically to prevent truncation
The localPath rebase overwrote config.yaml with a plain (non-atomic) write.
By that point the legacy partition has already been renamed away, so
config.yaml is the partition's ONLY copy — a write that fails partway
(ENOSPC, EFBIG, crash mid-write) truncates it with no source to recover from,
and the CLI still prints "Your original data is unchanged, re-run to retry"
while the data is in fact corrupt and cannot self-recover.
Add writeFileAtomic (same-dir temp + rename, preserving mode) alongside the
existing writeJsonAtomic, and use it for the adopted config.yaml. rename(2)
is atomic, so a failed write removes the temp file and leaves the original
config.yaml byte-for-byte intact; the next command retries the idempotent
rebase and converges.
Reproduced with the real CLI (isolated HOME, real local git team repo, NO fs
mock) by injecting a write failure with RLIMIT_FSIZE=128:
before: pull --force reports EFBIG, config.yaml truncated 453 -> 128 bytes,
legacy partition already moved, retry after lifting the limit exits
0 but never syncs and cannot recover
after: config.yaml preserved at 453 bytes, retry after lifting the limit
prints "Synced 1 skills" and localPath points at the live clone
Regression test: the localPath rewrite failing mid-write leaves config.yaml
intact with no leftover temp file. Verified it fails when the write is made
non-atomic. Full suite: 3109 passed, 0 regressions; tsc clean.
Project-scope machine data lives in a partition under
~/.teamai/projects/<slug>/, but the slug prefix was only the anchor's
basename (e.g. teamai-cli-<hash>), so the directory name did not reveal
which project it belonged to — unlike Claude Code's ~/.claude/projects/,
whose names encode the full path.
Build the prefix from the WHOLE anchor path (leading separator dropped,
path separators -> '-'), e.g. /Users/x/Project/app ->
Users-x-Project-app-<hash>. The 16-hex sha256 suffix is unchanged, so
uniqueness (escape-ambiguous paths, shared basenames) and the
slug===dir corrupt check are untouched; the prefix is length-bounded so
a very long path can never overflow NAME_MAX. The authoritative reverse
lookup remains the per-partition 'anchor' file.
Verified end-to-end with the real CLI in an isolated HOME: a partition
is laid down under the readable slug, 'status --all' reverse-resolves it
to the project path and reports [active], a mismatched dir name is
flagged [corrupt], and detection from the project cwd loads the
partition config via the new slug.
PR #520 changed cnbExec to resolve the CLI path (resolveCliPath) and
launch it via cross-spawn instead of a bare spawnSync('cnb', ...). PR
#536 landed cnb-create.test.ts mocking node:child_process. Both were
green on their own branches, but once merged the test mocks the wrong
layer: resolveCliPath returns null in CI, cnbExec short-circuits with
'cnb CLI not found', and the canned responses never reach the code —
so cnbOrganizationExists returns false and 7 cases fail on main.
Mock cross-spawn's default.sync + utils/cli-path.resolveCliPath (the
pattern #520 already uses in cnb-login-host.test.ts). Assertions are
unchanged; only the mock injection point moves.
The cnb login OAuth token carries neither group-manage:rw (create org) nor
group-resource:rw (create repo), so create-repo failed with a raw 403/404 that
gave users no next step.
- Detect a missing organization up front (cnb get-group, a read-only scope the
login token has) before prompting to create the repo.
- Surface OrganizationNotFoundError / RepoCreatePermissionError from the CNB
provider, each carrying the platform's web create URL.
- init now prints the create page (cnb.cool/new/groups or /new/repos) and exits,
instead of a confusing scope error.
Adds GitProvider.organizationExists?/getOrganizationCreateUrl? (CNB only) and
tests; docs updated (providers.md, usage-guide.*).
Cover the two gaps in real-command e2e:
- data-layout-migration.test.ts: seed a legacy <repo>/.teamai (real
team-repo clone + plaintext env + localPath into legacy), run the
compiled CLI 'pull --force', and assert the preAction hook migrated
data into ~/.teamai/projects/<slug>/, kept the git clone intact,
rebased repo.localPath onto the partition, retired the source to
.teamai.bak, and pulled through the migrated clone.
- multi-project.test.ts: 'projects set' + 'pull --force' deploys only
the active project's skills and prunes an inactivated one; 'push
--project' routes a skill into the project's skills/<ns>/ on the
remote, matching what pull reads.
Hermetic: sandbox HOME + local bare remote (provider: git), no network,
tokens, or PR creation. Land in src/__tests__/e2e/ so vitest.e2e.config
picks them up.
The cnb CLI infers its platform URL from the first git remote of the
current directory when --host is not given. Running `cnb login` inside a
repo whose remote points at a non-CNB host (e.g. an internal git server)
therefore sends the OAuth2 device-auth request there and fails with 401,
while the same command succeeds in a directory with no such remote.
Pass --host ${CNB_HOST} explicitly. CNB_HOST (TEAMAI_CNB_HOST, default
cnb.cool) is already the single source of truth for every other CNB
operation (clone / create-repo / PR), so anchoring login to it keeps the
provider's auth consistent and self-hosted-friendly.
`teamai init` cloned CNB repos via `git -c credential.helper='!cnb git-credential'
clone`, but the `-c` flag only applies to that single invocation — it never
reached the cloned repo's `.git/config`. So `remote.origin.url` stayed
credential-free, and the subsequent push (member registration) plus `teamai pull`
fell back to an interactive Username/Password prompt despite the user being
logged into the `cnb` CLI.
GitHub/TGit solve this by embedding the token in the clone URL (persisted into
remote.origin.url); CNB's CI path (CNB_TOKEN) does the same. Only the
interactive-login path was missing persistence. Persist the helper into the
repo's local config after a successful interactive clone so every later git
operation authenticates transparently.
Verified end-to-end against a real cnb.cool account (logged-in, no CNB_TOKEN):
`teamai init https://cnb.cool/test1122444/test` now pushes member registration
without prompting, and `teamai pull` reports up to date instead of
"couldn't find remote ref".
GitHub's automatic base selection takes the newest tag reachable from the
tagged commit. Release commits are created detached and never land on a
branch, so no release tag is an ancestor of the next one and every recent
release compared back to v0.22.0 — the last release that happened to be
merged into main. The 0.23.x stables are worse: they were cut from an
internal mirror sync, so their merge-base with main sits at v0.17.4 and
their notes span three months of unrelated PRs.
Resolve the base explicitly and feed it to the generate-notes API. A
prerelease compares against the tag before it; a stable release compares
against the previous stable tag. Both skip tags whose parent is absent
from this release's history, so a mirror-sync tag can no longer poison
the range.
Co-authored-by: Cursor <cursoragent@cursor.com>
A pre-P3 partition (written before anchor-on-save) has no anchor file, so
statusAll fell back to the config's businessRepoRoot/projectRoot. That is a
persisted *workspace* path that can point at a linked worktree — the partition
itself is keyed by the shared project anchor (main checkout) and stays active.
When such a worktree was removed but the main checkout remained, status --all
labeled the still-active partition 'ORPHAN — safe to delete' and printed an
rm -rf hint, so following the advice would delete config/state still in use.
Base the active/ORPHAN/corrupt verdict solely on the anchor; use the config path
for display only. A partition with no anchor is now 'unknown', never ORPHAN — we
never recommend deleting data we cannot confirm is dead.
Adds regression tests (anchor-gone → ORPHAN, no-anchor+removed-worktree →
unknown, anchor-present → active) and syncs the design + bilingual usage docs.
Two parts, closing out issue #374.
A set of top-level path constants were computed once at module import
(`export const TEAMAI_HOME = path.join(getUserHome(), '.teamai')` and its
derivatives). Frozen at import, they made a test's later `HOME` swap a no-op — so
`HOME`-based isolation silently failed and tests worked around it with
`vi.resetModules()` / `vi.mock('../types.js')`. Converted them to call-time
getters (`getTeamaiHomeDir`, `getUserConfigPath`, `getUserStatePath`,
`getTokenPath`, `getUpdateLockPath`, `getSessionLogsDir`, `getUserLearningsDir`,
`getUserSearchIndexPath`, `getUserVotesDir`), matching the existing
`getUserHome()` / `getDataHome()` pattern, and updated all ~45 call sites
(incl. the module-private consts in import-mr.ts / import-local.ts / gf-cli.ts).
Removed seven dead consts that already had runtime getters and no live consumers
(TEAMAI_SOURCES_DIR, TEAMAI_USAGE_PATH, TEAMAI_KNOWN_SKILLS_PATH,
TEAMAI_PUSHIGNORE_PATH, CONTRIBUTE_SESSIONS_DIR, DASHBOARD_EVENTS_DIR/PATH).
Functionization is NOT project-scoping: every getter still returns
`~/.teamai/...` (class A2, landing unchanged); the project-scoped equivalents
already route through getDataHome(). The dashboard stays an A2 singleton keyed by
event cwd/sessionId — "two projects' events don't mix" holds because
getEventsPath() reads HOME at call time.
Cleaned up the now-unnecessary test workarounds (home.test resetModules dance,
the vi.mock('../types.js') const overrides in init/gf-cli/update).
Previously only migration wrote a partition's `anchor` reverse-lookup file, so
freshly-init'd partitions had none. saveLocalConfigForScope now writes it whenever
a config lands in a partition (shared writeAnchorFile helper, also used by
migrate). `teamai status --all` (a new flag on the existing status command)
enumerates every partition under ~/.teamai/projects, recovers each project path
(anchor file, falling back to config businessRepoRoot/projectRoot), and flags it
active / ORPHAN (project gone → safe to rm -rf) / unknown / corrupt. teamai never
auto-collects orphans (no gc command — out of scope), so this is how a user finds
partitions to delete by hand.
Tests: p3-functionize.test.ts (getters follow a runtime HOME change; dashboard
isolates by HOME so two projects don't mix; a partition config gets an anchor,
user scope does not). Full suite 2917 passed. Verified end-to-end: getters honor
HOME, and `status --all` lists an active + an orphan partition with the cleanup
hint. Docs: design doc P3 section + bilingual usage guides (status --all).
Follow-up to the shell-env guard: it treated ANY non-empty ANTHROPIC_* in
process.env as user-owned. But Claude injects settings.json.env into the hook
subprocess, so a value TeamAI itself wrote last sync reappears in process.env.
The guard then read its own managed gateway as a user conflict, skipped every
follow-up sync, and cleared manifest.claudeEnv — after which the gateway could
never be updated or removed. Token rotation / gateway retirement left the user
stranded on dead config, all acked success. (Reported as a P1 regression.)
- Recognize a managed value: process.env[key] whose hash matches previousHashes
or the current settings.json env[key] is our own re-injection, not a user's
independent shell config, so it is not a conflict.
- Split the two skip reasons: a genuine shell conflict skips the write but keeps
manifest.claudeEnv (so reconcile resumes once the shell config is gone); only a
user edit to our managed settings.json entry clears the managed record.
- Add regression tests driving the 4-step update/remove flow with the managed env
re-injected into process.env; verified they fail without this fix.
The Claude model-config guard only checked ~/.claude/settings.json's env,
so a user who runs Claude via shell `export ANTHROPIC_*` (keeping no gateway
in settings.json) read env[key] === undefined and had the enterprise gateway
written into settings.json. settings.json outranks the shell env, so their
working gateway was silently overridden and Claude broke.
- Extend the conflict check to inspect process.env as well as settings.json.
TeamAI never exports to the shell, so any non-empty ANTHROPIC_* shell value
is user-owned and is never reconciled away.
- Widen the conflict key set to auth the gateway swap would break
(ANTHROPIC_CUSTOM_HEADERS) and the user's model choice
(ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU}_MODEL).
- On skip, log the conflicting keys to reporter/errors.jsonl for diagnosis.
- Isolate the test env from the runner's own ANTHROPIC_* shell vars.
[P1 regression] A self install synced the wrong branch's knowledge across
worktrees. After P2 the self config lives in the SHARED partition (keyed on the
main checkout's projectAnchor), so every worktree reads the same config. Its
persisted repo.localPath / businessRepoRoot name the checkout that first migrated
(main). detection only re-anchored projectRoot, leaving localPath pointing at
main — and getKnowledgeDir === repo.localPath. So a `teamai pull` in a feature
worktree read MAIN's knowledge and injected it into the feature tree, reporting
success. This makes the current branch's AI use the wrong rules/skills.
Repro (real git worktree): main skill = MAIN_BRANCH_ONLY, feature worktree skill
= FEATURE_BRANCH_ONLY; pull --force on both → feature ended up with
MAIN_BRANCH_ONLY (baseline 999f57c: FEATURE_BRANCH_ONLY).
Fix: readConfigFrom re-anchors a self config's repo.localPath to
<workspaceRoot>/.teamai and businessRepoRoot to <workspaceRoot> (self's
invariant), so knowledge + business-root track the CURRENT worktree while
dataHome stays the shared partition (machine data is shared across worktrees on
purpose). Non-self localPath is the team-repo clone path — shared deliberately —
and is left untouched.
Tests (self-worktree-rebind.test.ts): a feature worktree resolves knowledge to
its own .teamai (not main); the main checkout still resolves to itself; a
git-mode config's clone localPath is NOT rewritten. Full suite 2793 passed.
- members.ts: mergeMemberConfig (append+dedupe project membership)
- init.ts: member registration merges active projects into members/<user>.yaml
(both non-self and self/reports-branch paths), no longer no-op on re-init
- projects-cmd.ts: projects list / set / members commands
- index.ts: register projects subcommand + push --project option
- push.ts: --project resolves to the project's manifest skills namespace
(manifest-resolved, not raw id) and reuses the --role landing logic
- tests: mergeMemberConfig union/dedupe/role cases
Two P1 issues found in review of the self-mode relocation, both reproduced and
fixed with regression tests.
[#1] Directory entry dropped on relocation → data loss. migrateSelfA1's
"destination exists → remove(src)" branch treated a DIRECTORY entry (workspaces/)
the same as a file: it deleted the whole legacy workspaces/ without merging.
Reachable: a prior run relocated workspaces/ then crashed before config.yaml;
with the partition config still absent, dataHome stayed legacy, so a reconcile
between crash and retry wrote a NEW worktree under <repo>/.teamai/workspaces/<id>;
on retry the dest existed → remove(src) dropped that worktree's managed-mcp +
resource cache, unrecoverable. Fix: a directory entry whose dest already exists is
MERGED (relocate only the children the partition lacks, each via temp + atomic
rename, never overwriting an authoritative partition child), then the drained
source is removed. File entries keep the "partition wins, drop source" behavior.
[#2] self pull/push shared no lock with migration → write-back / torn write.
pull.ts lockScope returned true for any non-git kind (no lock); push.ts's self
branch returned before its acquireLock. So an in-flight self pull, whose dataHome
resolved to legacy before migration ran, could saveStateForScope /
search-index.json back into <repo>/.teamai AFTER the migration deleted them —
re-polluting the workspace and stranding that pull's update in the repo while the
partition kept the old value. Fix: self now contends on <getDataHome>/.sync-lock —
the exact path migrateSelfA1 takes (pre-migration that is <repo>/.teamai/.sync-lock)
— so pull skips (idempotent) and push errors on contention, same as git-mode.
http still takes no lock (no clone, no relocation).
Tests: workspaces/ merge carries a legacy-only child over while leaving an
authoritative partition child untouched; a held sync-lock makes self pull skip
(without overwriting machine data) and self push exit non-zero. Full suite 2790
passed. Verified end-to-end with the built CLI: with the sync-lock held, self pull
skips and leaves state.json unchanged, self push errors — the write-back window is
closed.
The How It Works diagram already covers the flow; skill paths and project-scope directory rules belong in the usage guide.
Co-authored-by: Cursor <cursoragent@cursor.com>
Keep the README short: drop self-hosted GitLab setup from Quick Start, and label Team Context / Team Improvement as unfinished.
Co-authored-by: Cursor <cursoragent@cursor.com>
Before P2, self (single-repo) mode kept its class-A1 machine data (config,
state, env backup, search index, managed-mcp, the per-worktree resource cache)
inside the business repo at <repo>/.teamai/, next to the class-B team knowledge
committed to main. A hand-maintained gitignore blacklist kept git status clean —
fragile (the per-worktree workspaces/ tree and user-scope managed-mcp.json were
never listed, so a self repo running MCP reconcile leaked them into the tree).
P2 physically relocates A1 to the partition ~/.teamai/projects/<slug>/, leaving
.teamai/ with only class-B knowledge. Same lever as non-self installs: attach a
partition dataHome to the self LocalConfig, and every getDataHome()-based write
follows.
Invariant kept: getKnowledgeDir / repo.localPath stay <repo>/.teamai (the class-B
anchor, committed to main), so the ~230 path.join(localPath, …) sites don't
change; reports-wt/knowledge-wt stay in the repo (git worktrees anchor there).
- init (initSelfRepo): attaches the partition as dataHome; writes config/state
there. The old "retire the stale partition" step is gone — self now USES the
partition, so an overwrite of the same partition config supersedes it.
- bootstrap: the already-initialized check and config write target the partition
(legacy fallback keeps a pre-P2 install recognized).
- detection seam: on a fresh clone the partition config does not exist yet, so
partition-first misses; the legacy branch runs the self-heal bootstrap (which
now writes into the partition) then reads it back FROM the partition
(selfHealAndReadPartition). A pre-P2 install is read via the legacy branch
(double-read compat) until migration relocates it.
- migration (mode: 'self'): self CANNOT use the git-mode whole-dir copy→rename
(that carries the knowledge off and renames .teamai to .bak, breaking
"knowledge on main"). It selectively relocates the A1 whitelist entry-by-entry
into the partition, leaving class-B knowledge and the worktrees untouched and
never renaming .teamai/. self repo.localPath is NOT rebased.
- gitignore hardened: workspaces/ + managed-mcp.json + .sync-lock added (the
pre-P2 leak sources); existing entries kept for pre-migration double-read.
Adversarial self-review found and fixed three issues (regression-tested):
- [C1] self relocation was non-atomic at the planner level: config.yaml was the
first entry AND planMigration keyed on legacy config.yaml existing, so a crash
after moving config.yaml went blind and stranded the plaintext env.local/env.sh
in the repo forever. Fix: config.yaml is relocated LAST (sentinel), and
planMigration gates on ANY A1 entry (reading kind from the partition config
when legacy config.yaml is already gone).
- [H1] a self pull/push does NOT take the partition .sync-lock (lockScope skips
non-git kinds), so a concurrent hook-pull could read a half-copied env.local.
Fix: relocate each entry via a temp sibling + atomic rename, so a reader sees
the complete file or nothing.
- [M2] a pathological <repo>/.teamai carrying self knowledge but a kind:git
config would take the git-mode retire path and rename .teamai to .bak, wiping
committed knowledge. Fix: retireLegacy refuses to rename a dir holding a
teamai.yaml mode:self (fail closed).
Tests: 28 migrate cases incl. self relocation (A1 moved / class-B kept /
.teamai not renamed / localPath not rebased), idempotent + interrupted-run
finish, the C1 crash state, the M2 guard, and dry-run. Full suite 2788 passed.
Verified end-to-end with the built CLI: a pre-P2 self install's pull relocates A1
to the partition, git status goes clean, the pre-P2 workspaces/managed-mcp leak
is gone, class-B knowledge stays on main, and the C1 crash state is recovered
(the stranded plaintext secret is relocated, not left in the repo).
Docs: design doc P2 section + bilingual usage guides updated (self data layout,
upgrade note).
Write apply_model_config to ~/.workbuddy/models.json and workspace
.codebuddy/models.json using the documented object wrapper, while
preserving legacy top-level arrays and user-owned entries.
--story=1020422209136345350
Co-authored-by: Cursor <cursoragent@cursor.com>
`teamai source add` previously succeeded on a repo with no teamai.yaml
with only a vague "It can still be used" hint, while pull silently skipped
it (log.debug). Members ended up subscribed to a source that syncs nothing,
with no signal at add time.
Warn explicitly at add time, distinguishing the two zero-skill cases:
- no teamai.yaml at all
- teamai.yaml present but no publicSkills declared
Extract the wording into a pure `sourceSyncWarnings()` helper and unit-test
all three branches. Closes#416 (UX point 4 only; the auto-upstream-tracking
points are intentionally out of scope — auto-following third-party upstreams
would bypass team review and open a supply-chain injection vector).
Refs #416
Two P1 issues found in review of the retireLegacy step, both reproduced locally
and fixed with regression tests.
[P1] Existing .teamai.bak was destroyed. retireLegacy did an unconditional
remove(backup) before the rename, so a pre-existing `.teamai.bak/` (from a prior
migration or the user's own) was deleted — a plain `teamai pull` on upgrade could
cause unrecoverable data loss. Fix: never remove an existing backup; pick the
first FREE name instead (`.teamai.bak`, else `.teamai.bak.1`, …).
[P1] The renamed backup could leak credentials into git. An old install's
`.teamai/` was often protected ONLY by a repo-root `.gitignore` rule matching
`.teamai/`, which does NOT match `.teamai.bak/`. After the rename a `git add -A`
staged the plaintext env/token (confirmed: `git show :.teamai.bak/env` returned
the secret). Fix: drop a self-contained `.gitignore` (`*`) INTO the dir BEFORE
renaming, so the backup ignores its own contents regardless of its final name or
the repo's ignore rules — no window in which the credentials sit unignored.
Tests: a pre-existing .teamai.bak is preserved while the migration backup goes to
.teamai.bak.1; after migration `git add -A` stages nothing under .teamai.bak and
`git show :.teamai.bak/{env,token}` both fail, while the files remain on disk for
rollback. Verified end-to-end against the built CLI across three repo-root ignore
spellings (`.teamai/`, `/.teamai/`, `.teamai`).
Adversarial self-review of the migration found three issues; all fixed with
regression tests.
[H1] A crash between the partition rename (step 3) and the source retire (step 5)
left the partition authoritative but the legacy dir — including its plaintext env
— lingering in the workspace forever: the next run's planMigration hit the
"partition exists" check and returned null, so the legacy dir was never retired,
breaking the zero-residue guarantee. Fix: planMigration now returns a
'retire-only' plan when the partition exists but a legacy dir still lingers;
runMigration finishes the job by retiring the leftover to .teamai.bak WITHOUT
re-copying onto the authoritative partition. The same path handles the
under-lock TOCTOU case (a sibling built the partition while we waited).
[L5] verifyStaging now smoke-checks the staged team-repo clone with
`git rev-parse HEAD` (on staging, before the rename), so a partial/corrupt copy
aborts with the source untouched instead of promoting a broken clone the next
pull would choke on. Existence of .git alone is no longer taken as proof.
[M3] The async preAction hook under program.parse() would surface a migration
failure as a raw unhandled-rejection stack. maybeMigrate is now wrapped: on
failure teamai prints a clean error ("your original .teamai is unchanged, re-run
to retry") and exits non-zero, rather than crashing into the command on partial
state.
Tests: retire-only finishes an interrupted run without overwriting the
partition; a corrupt staged clone aborts leaving source + no partition/.bak/
.staging; http-mode install (no team-repo) migrates. Verified end-to-end with
the built CLI: the interrupted-state pull retires the lingering legacy dir and
moves the plaintext env out of the workspace while keeping the partition intact.
An install created before partitioning keeps its machine data in the business
repo at <workspaceRoot>/.teamai/. P1-2 routed NEW installs to the partition and
read old installs via a legacy fallback; this moves a real legacy .teamai INTO
the partition on the next write command so the workspace ends up zero-residue.
src/migrate.ts:
- planMigration(): gate WITHOUT detectProjectConfig (which short-circuits on an
existing partition and runs the self-heal bootstrap — both would mask the raw
legacy state). Migrate iff git repo + legacy config.yaml exists + no partition
config yet + scope=project + kind!=self. self mode is a hard no-op (its
.teamai is team knowledge committed to main).
- runMigration(): copy -> verify -> atomic rename so an interruption never
leaves data half-in-both-places. Holds <legacyDir>/.sync-lock (the exact path
an un-migrated pull/push contends on) across the whole window; contention
skips (idempotent). Copies with raw fse.copy NOT copyDir — copyDir filters out
.git and would corrupt the team-repo clone. Skips disposable worktrees
(reports-wt/knowledge-wt) and lock files. Writes an anchor reverse-lookup file
(the slug is a one-way sha256). Retires the source to .teamai.bak (never
auto-deleted — the manual rollback path).
index.ts: wire maybeMigrate into the global preAction hook (now async), narrowed
to init/pull/push and excluding TEAMAI_HOOK_SUBCOMMANDS (hook-dispatch must
never move 12MB). --dry-run previews without writing.
Downgrade after migration is not supported — flagged in the design doc and the
bilingual usage guides, which also correct the now-stale "<cwd>/.teamai" layout
to the partition.
Tests: src/__tests__/migrate.test.ts builds a REAL legacy layout (real anchors +
a real nested git clone) and covers every gate, .git survival, atomicity,
idempotence, interrupt recovery, and dry-run. Verified end-to-end with the built
CLI: dry-run no-op, real pull/push migrate, git clone stays usable, workspace
zero-residue, self mode / status / hook-dispatch never migrate.