Files
teamai-cli/vitest.e2e.config.ts
T
Jiahe Gengandjaelgeng 7fd9f436a4 Phase 0+Phase 1+P4.4: 实现了subagent推送,检索subagent,初始化知识库导入以及MR自动learnings&codebase更新功能(团队级别codebase文档功能有待完善) (#28)
* docs:知识库飞轮系统设计roadmap文档

* feat: phase 1 - agents support, multi-index search, recall subagent & todowrite hint

- Add agents resource type (pull/push/remove support)
- Add builtin-agents with teamai-recall agent
- Add multi-source search index (rules + skills + agents)
- Add todowrite hint injection for recall subagent
- Add phase1 e2e tests and unit tests for new features
- Update README (zh-CN + en) with agents documentation

* docs: update roadmap文档

* feat(search): P1.4 domain inference + search weighting

Introduce KnowledgeDomain (technical/ops/support/neutral) inferred from
frontmatter > tags > path > type fallback. Apply per-domain score
multipliers at search time (technical ×1.0, neutral ×0.85, ops ×0.5,
support ×0.3) plus a skills/rules type bonus (×1.1).

Bump SEARCH_INDEX_VERSION 2→3; legacy v2 indexes auto-rebuild on next pull.

--other=P1.4 domain inference and search weighting

* test(validation): add Phase 1 E2E suite and acceptance report

Move phase1-e2e.test.ts to validation/ alongside the acceptance report,
update import paths and vitest.e2e.config.ts to pick up the new location.

Fix pre-existing E2E isolation bug: loadStateForScope mock used
mockResolvedValue (shared object reference), causing test-1 to mutate
state.lastPullRev and trigger the rev-based early-exit in test-2/3.
Switch to mockImplementation(() => ({ lastPull: null })) so each call
gets a fresh object. All 5 E2E tests now pass.

--other=Phase 1 validation

* docs: update roadmap — add P4.6 learning promotion mechanism

Learning entries that accumulate confidence ≥ 0.90 (5+ upvotes, 2+
contributors, 14+ days old) are prompted for promotion to docs/skills/rules
based on content type, regardless of origin or domain. Add P4.6 step to
Phase 4, dependency graph, implementation detail, work-estimate table,
and architecture overview diagram.

--other=roadmap update

* docs(validation): update phase1 report to reflect E2E bug fix

All 5 E2E tests now pass. Update P1.2 status to ✅, remove E2E
failure notes and known issues, set pass rate to 100%.

--other=phase1 report update

* docs(validation): add runtime evidence appendix to phase1 report

Append Appendix A1-A4 with captured outputs from demo-phase1.test.ts:
agents sync paths, CLAUDE.md full content after injection,
search-index.json entry listing (4 types covered), and recall("api")
full STDOUT including envelope markers and domain-weighted scores.

Knowledge content in A3/A4 is anonymised (titles, authors, file paths
replaced with generic placeholders).

Also commit demo-phase1.test.ts as the reproducible evidence runner.

--other=phase1 report runtime evidence

* feat(search): query-aware domain weights + IDF scoring (v4 index)

改动 A — 查询感知 domain 权重:
将静态一维 DOMAIN_WEIGHT 改为二维查询-文档权重矩阵,新增
inferQueryDomain() 从查询 token 推断查询域。搜索 k8s/deploy 等
ops 相关问题时,ops 条目不再被打五折,technical 查询行为与原来一致。

改动 B — IDF 降权:
buildIndex() 末尾计算 df(文档频率)map 并写入索引;search() 中
为每个 token 匹配乘以 log((N+1)/(df+1))+1 的 IDF 权重,高频通用词
(api、deploy、error)自动降权,低频专有词(deepgemm、mooncake)
权重保持不变。

索引版本 3 → 4;isLegacyIndex() 新增 !index.df 判断触发重建。
旧 v3 索引在下次 teamai pull 时自动重建,search() 对无 df 字段的
旧索引降级为 idf=1.0,不报错。

--other=search quality improvements

* feat(import): add teamai import command — Phase 0 cold-start + P4.4 MR pipeline

## 新增命令:teamai import

支持五种知识来源:
- --dir <path>:扫描本地目录,AI 分类为 rule/doc/learning
- --from-claude:迁移 ~/.claude/rules 等 AI 工具规则目录
- --workspace:基于当前 git 仓库生成 codebase.md
- --from-mr <url>:从已合并 MR 提炼 learning + codebase 更新建议(P4.4)
- --from-iwiki <id/url>:从 iWiki Space 批量导入文档

## 新增核心模块

- src/utils/ai-client.ts:claude -p 子进程封装(并发 ≤ 3,60s 超时)
- src/utils/dedup.ts:Jaccard 相似度重复检测(14 天窗口,≥ 60% 标记 superseded)
- src/utils/iwiki-client.ts:iWiki MCP HTTP 客户端(JSON-RPC 2.0,零外部依赖)
- src/import-local.ts:本地文件扫描/AI 分类/交互确认/推送
- src/import-mr.ts:MR 三层解析/双路 AI 提炼/dedup/推送
- src/import-iwiki.ts:iWiki 导入(复用 import-local.ts 基础设施)
- src/codebase.ts:codebase.md 生成/增量更新

## 扩展现有接口

- providers/types.ts:GitProvider 新增可选 fetchMergeRequest() 方法
- providers/github/mr-fetch.ts:gh pr view 实现
- providers/tgit/mr-fetch.ts:gf mr 实现
- types.ts:新增 MRData/ClassifiedItem/LearningDraft/CodebaseSuggestion/ImportSession

## 测试 & 文档

- ai-client.test.ts:5 tests(spawn mock + 并发控制)
- dedup.test.ts:11 tests(关键词提取 + Jaccard + 文件扫描)
- validation/phase0-p44-acceptance-report-public.md:Phase 0 + P4.4 验收报告

--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线

* fix(import): support claude-internal CLI + gh REST API fallback + real PR demo

## ai-client.ts
- detectClaudeCli():按 claude → claude-internal 优先级探测,结果缓存,
  进程内只探测一次,两者均不可用时给出清晰错误提示
- DEFAULT_TIMEOUT_MS:60s → 180s,适应大 diff 下 AI 提炼耗时

## providers/github/mr-fetch.ts
- gh CLI 不可用时自动降级到 GitHub REST API(内置 https,零依赖)
- 支持公开仓库无 token 访问;有 GITHUB_TOKEN 环境变量时自动携带

## import-mr.ts
- codebase 建议 JSON 解析:先提取 {…} 块再解析,
  兼容 AI 在 JSON 前附加说明文字的输出格式

## index.ts
- import 子命令新增 --all 选项,跳过交互确认

## validation
- phase0-p44-acceptance-report-public.md A4:替换为基于真实 PR #2
  的端到端操作记录(真实终端输出 + AI 真实生成的 learning.md
  和 codebase-suggestions.json 完整原文)

--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线

* fix(codebase): require file-path prefix in module descriptions for agent guidance

codebase.ts 和 import-mr.ts 的提示词中"主要模块"格式要求均已更新:

- 之前:**模块名** — 功能说明(AI 可能只写中文名,无路径索引)
- 之后:**文件或目录路径** — 功能说明(明确要求带路径,便于 agent 定位)

codebase.ts: 全量生成 prompt 示例改为 **src/utils/git.ts** — 功能说明
import-mr.ts: codebase 建议 prompt 新增正确/错误示例,强制路径前缀

同步更新验收报告 A4:
- codebase-suggestions.json 从 8 条无路径条目 → 11 条带路径条目(真实重跑输出)
- 更新后 codebase.md 主要模块列表格式与更新前一致

--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线

* feat(codebase): upgrade workspace + MR prompts to A1-level documentation quality

codebase.ts — gatherRepoContext 扩展:
- 增加 package.json(依赖和 scripts)
- 增加入口文件(src/index.ts)命令注册全文
- 增加类型定义文件(src/types.ts)关键接口
- 文件树深度 maxdepth 3→4,过滤 dist/ worktrees/
- 截断上限:FILE_TREE 3000→5000 字符,DOC 1000→2000 字符
- 新增 META_MAX_CHARS 常量(2500 字符)

codebase.ts — 全量生成 prompt 重写:
- 提供完整 8+ 章节格式骨架(项目概述/技术栈/目录结构/数据配置/
  核心数据流/关键接口/配置系统/性能可靠性/测试覆盖/备注)
- 目录结构要求带分组框树形图(┌─ 功能分组 ──┐ 风格)
- 技术栈要求表格含版本信息
- 项目概述要求带 emoji 核心能力 bullet list
- 核心数据流要求带缩进 → 的流程图格式

import-mr.ts — codebase 建议 prompt 升级:
- 新增 existingCodebaseMd 参数,注入现有文档全文作为格式样本
- AI 参考现有文档的分组和粒度生成风格一致的增量条目

import.ts — --from-mr 分支:
- 调用前读取 repoPath/docs/codebase.md 传入 existingCodebaseMd
- 确保 MR 增量更新与初始生成风格一致
- 复用已导入的顶层 fs 模块,删除内联 dynamic import

--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线

* docs(validation): replace A1+A4 codebase docs with real teamai-cli generated output

A1 附录:替换为 teamai import --workspace 对 upstream/main 真实生成的
codebase.md(含分组框目录树、表格技术栈、emoji 核心能力、流程图数据流)

A4 附录:
- Step 1:更新 codebase-before.md 为同一真实生成版本
- Step 2/3:终端输出更新为 3 条建议(主要模块 7 条/关键路径 4 条/架构决策)
- Step 3:learning.md 更新为最新真实 AI 输出
- Step 4:更新前后对比基于真实文档
- Step 5:飞轮闭环统计数字更新(3 条建议)

--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线

* feat(mr-hint): add SessionStart hook to hint AI about unimported merged MRs

P4.4 优化:在每次 Session 开始时检测当前 git 仓库的 origin remote,
查询近 7 天内已合入但尚未通过 teamai import 处理的 MR,并通过
additionalContext 提示 AI 在任务完成后建议用户运行 teamai import --from-mr。

核心实现:
- src/mr-hint.ts:新增 mrHint() 入口,支持 TGit REST API 和 GitHub gh CLI
  双路查询;per-repo 磁盘缓存(30 天 TTL)避免重复提示相同 MR
- src/hooks.ts:注册 SessionStart hook,更新 TEAMAI_COMMAND_MARKERS
  和 TEAMAI_HOOK_SUBCOMMANDS,同步 buildCursorHooks
- src/index.ts:注册 mr-hint 子命令
- 同步修复 hooks.test.ts、usage-tracking.test.ts、doctor.test.ts
  中 todowrite-hint 加入后遗留的计数断言

--other=P4.4 MR 合入统一处理流水线(SessionStart 触发提示)

* feat(mr-hint): add GitHub REST API fallback when gh CLI unavailable

gh CLI 不在环境中时自动回退到 GitHub REST API(/repos/.../pulls),
逻辑与 providers/github/mr-fetch.ts 保持一致;
支持公开仓库无 token,有 GITHUB_TOKEN 时自动携带以提升限速上限。

--other=P4.4 MR 合入统一处理流水线(mr-hint GitHub REST fallback)

* docs(validation): update public acceptance report for P4.4 mr-hint trigger mechanism

新增 P4.4 触发机制优化章节,更新附录 A1(基于含 mr-hint 模块的代码库
真实生成),追加附录 A4.2(SessionStart hook 自动感知 merged PR 的真实
运行场景,含 GitHub REST API fallback 演示与幂等性验证)。

--other=P4.4 MR 合入统一处理流水线验收报告更新

* feat(ai-client): support login shell PATH + extend CLI candidates

- 用 bash -lc 包裹探测和调用,解决 ~/.nvm 路径下 CLI 不在 PATH 的问题
- 探测顺序扩展为 claude / claude-internal / codex / codex-internal /
  codebuddy / workbuddy / openclaw
- 新增 shellEscape() 避免 prompt 中单引号破坏 shell 命令
- 超时从 180s 降至 120s
- 测试补充 execFileSync mock,修复预存 5 个失败用例

feat(import-mr): interactive codebase review loop + apply suggestions

- 新增 reviewCodebaseSuggestions():AI 实时修订循环,用户输入意见
  → AI 修订 → 再展示,直到 y 确认或 n 跳过
- --output 模式:apply 后写 codebase-after.md(完整 Markdown)
- repoPath 模式:apply 后写回 docs/codebase.md
- 修复 codebase.ts apply prompt,确保输出完整文档而非摘要

feat(import): add --existing-codebase option

allow 用户显式指定 before codebase.md 路径,不依赖团队仓库;
优先级高于从 repoPath/docs/codebase.md 自动读取

--other=P4.4 MR 合入统一处理流水线优化

* docs(validation): update A4 with real before/after codebase demo

用真实运行产物替换 A4 中的 codebase before/after 内容:
- Step 1 before:由 teamai import --workspace 在 PR #2 合入前的代码库生成
- Step 3 suggestions:teamai import --from-mr PR #2 的真实 AI 输出
- Step 4(新增):apply suggestions 后的 codebase-after.md,含 diff 展示
三阶段格式统一,均为同一版本 prompt 的真实运行产物

--other=P4.4 验收报告 A4 codebase before/after 更新

* fix(providers): add TGit REST API fallback + multi-shell CLI detection

- mr-fetch.ts: 新增 fetchTGitMRViaApi(),gf CLI 不可用时自动 fallback
  到 git.woa.com REST API(使用 ~/.netrc OAuth token);
  diff 获取失败时降级为空字符串而非中断流程
- ai-client.ts: detectClaudeCli() 对每个候选依次尝试
  bash -lc → zsh -lc → which,覆盖 fish/CI 容器等非标准 shell 环境

--other=fix-risk-items

* docs: 验收文档typo更正

* fix(ai-client): multi-CLI compat + shell injection hardening

- ai-client.ts: detectClaudeCli 解析 CLI 绝对路径并校验存在性,
  spawn 改为直调 absPath + 参数数组,删除 shellEscape 去 shell;
  新增 buildCliArgs 区分 codex/codex-internal 用 'exec' 子命令、
  其余 CLI 用 '-p',修复 codex 系 CLI 调用失败问题
- providers/tgit/mr-fetch.ts: execSync → execFileSync 数组参数,
  消除 mrIid/repoArg 命令注入风险
- mr-hint.ts: TEAMAI_MR_HINT_CWD 增加 path.resolve + statSync
  校验,非法路径静默跳过
- ai-client.test.ts: mock 适配新探测语义(command -v 返回路径
  + existsSync=true)

--other=phase0-p44-cli-compat-and-security

* feat(codebase): align with llm-wiki — frontmatter / index / lint / multi-source

参照 docs/llm-wiki.md 的持久化知识库理念,对 codebase 文档生成做四项优化:

- frontmatter:generateCodebaseMd 输出顶部注入 YAML frontmatter
  (title / lastUpdated / source / generator / schemaVersion),
  支持去重旧 frontmatter,便于跨会话溯源
- 索引体系:新增 generateCodebaseIndex 导出,从 codebase.md 抽取
  二级章节 + 一句摘要 + 关键词,输出 codebase-index.md,加速 LLM 跨
  会话定位
- 健康检查:新增 lintCodebaseMd 导出,AI 检测矛盾/过时/孤儿/缺失
  四类问题,返回 LintReport(含 severity 分级),不修改文档
- 多源聚合:generateCodebaseMd 入参新增 learningsSuggestions 与
  learningsDir,gatherLearningsContext 内部函数读取 learnings/*.md
  frontmatter tags 做高频统计,融合 P4.4 MR 建议进 prompt
- prompt 模板新增"架构决策与权衡""已知限制与演进方向"两章节
- import.ts workspace 流程串入索引生成 + lint 报告打印
- types.ts 新增 LintIssue / LintReport 接口
- 新增 codebase.test.ts 11 个单元测试,全部通过

--other=phase4-codebase-llm-wiki-alignment

* feat(search): codebase-index.md high-weight + skip codebase.md

- 新增常量 CODEBASE_INDEX_FILENAME / CODEBASE_FULL_FILENAME /
  CODEBASE_INDEX_WEIGHT_BOOST(×1.5)
- entryFromMdFile:同目录存在 codebase-index.md 时自动跳过
  codebase.md,避免全量文档与索引文件重复命中
- search():codebase-index.md 命中时额外乘以 1.5 权重 boost,
  recall 时章节摘要优先返回
- 兼容 subagent 与 fallback recall 两条路径,boost 在本地索引
  阶段生效,无额外 AI 调用开销
- 新增 3 个测试用例(跳过逻辑 / 权重 boost / fallback 路径),
  search-index.test.ts 共 26 tests 全通过

--other=phase4-codebase-index-search-boost

* docs(validation): refresh A1/A4 with llm-wiki-optimized codebase output

用新版 CLI(llm-wiki 优化后)重新执行 A1/A4,更新公开版验收报告产物:
- codebase-before.md:含 YAML frontmatter、架构决策与权衡、
  已知限制与演进方向两新章节
- codebase-index.md:新增章节索引文件(11 行索引表)
- codebase-after.md:PR #2 应用建议后的更新版
- learning.md / codebase-suggestions.json:最新 AI 提炼产物
- 记录实际使用模型:claude-internal v1.1.9(DeepSeek-V3.1-Terminus)
- 保持 tgit → [internal] 等脱敏风格

--other=phase0-p44-acceptance-report-refresh

* docs(validation): replace codebase-after full content with before/after diff

A4 Step 4 中 codebase-after.md 展示方式由全文改为 unified diff,
更直观反映 MR 建议应用后的变更:新增"主要模块"章节(+8 行),
包含 import-local/import-mr/import-iwiki/codebase 四个关键模块说明

--other=phase0-p44-acceptance-report-refresh

* docs(roadmap): add Phase 6 — Phase 5 hardening

Phase 5 shipped the team-level codebase aggregation pipeline; in
shipping it we deliberately deferred several reliability concerns to
keep each step deliverable. Phase 6 captures those deferrals as a
focused hardening pass — no new capability surface, just turning the
Phase 5 deliverables into something safe to run in production
indefinitely.

Six sub-steps, three independent (P6.0 / P6.1 / P6.5) and three
chained (P6.2 → P6.3 → P6.4):

- P6.0  Real TGit listOrgRepos (replace stub)
- P6.1  Cache lifecycle (LRU + size cap + GC command)
- P6.2  Section-level diff with HTML-anchor in-place updates
- P6.3  pending-review CLI (review / apply / reject)
- P6.4  Domain-drift auto-apply workflow
- P6.5  Global codebase doc lint (cross-file consistency)

Appendix C dependency table updated with all six rows.

Phase 5 leftovers explicitly out of scope here (二级业务域 / 跨仓重复
检测 / search-index 联动 / agent 检索效果量化) are listed under the
new "遗留至 Phase 7" block.

--other=phase6-roadmap

* feat(import): Phase 5 — team-level codebase aggregation

Lift teamai-cli's codebase knowledge base from a single-repo, local view
into a team-wide, multi-repo aggregation that can be initialized in one
command, kept in sync incrementally, and audited end-to-end. The work
ships in five sub-steps but is a single coherent feature; merging as one
commit per the project's MR rules.

What's new
==========

P5.0  Business-domain dictionary (src/domains/*)
    Zod schema + YAML store + AI batch clustering + single-repo
    recommendation + interactive review CLI + jsonl audit log. Library
    only; no CLI wiring at this step.

P5.1  Single remote repo import
    `teamai import --from-repo <url>` shallow-clones into
    ~/.teamai/cache/repos/<provider>/<owner>/<repo>, reuses the existing
    generateCodebaseMd scanner, and writes a per-repo summary at
    docs/team-codebase/repos/<slug>.md. Three-tier auth (HTTPS+token /
    HTTPS-anonymous / SSH); tokens are always redacted in error output.

P5.2  Batch import + domain aggregation
    `--from-repo-list <yaml>` drives a whitelist with per-entry
    { url, domain, auth, priority } plus org entries (deferred). Failures
    on one repo don't block siblings. Pure-template aggregator emits
    docs/team-codebase/domains/domain-<name>.md and a top-level
    docs/team-codebase/index.md from the per-repo files. Default output
    root moved to docs/team-codebase to avoid colliding with the
    existing teamai-cli self-codebase at docs/codebase.md.

P5.3  Incremental sync + domain drift
    `--incremental` skips the full clone when the cache is hit and
    LAST_SYNC is present, falling back to a fresh shallowClone if fetch
    fails. After scan, a fresh recommendDomain pass is compared against
    the existing assignment; divergent recommendations (different
    domain + confidence > 0.5 + delta > 0.4) land in
    domains.history.jsonl as a drift event without auto-reassigning.
    AI failures never block the main flow. CI scheduling examples
    shipped under examples/ci/ for GitHub Actions and Coding CI.

P5.4  Org bootstrap + iWiki dual output
    `--from-org <org> --bootstrap` lists repos via gh api (paged with
    /orgs/<o>/repos -> /users/<o>/repos fallback), AI-clusters them,
    and walks the user through reviewDomains to produce both
    domains.yaml and repo-whitelist.yaml before chaining into
    importFromRepoList for the first full sync. TGit listOrgRepos is a
    stub (Phase 6).
    `--from-iwiki --iwiki-dual` extracts business APIs / external
    knowledge / glossary into docs/team-codebase/external-knowledge.md
    guarded by HTML comment anchors so future syncs replace bodies in
    place. `--require-review` defers section writes to
    .teamai/pending-review.jsonl. A small source-conflict helper flags
    multi-source updates within a 24h window.

Filesystem layout introduced
============================

  docs/team-codebase/
    index.md                  # business-domain map + repo index
    domains/domain-*.md       # per-domain aggregate
    repos/<slug>.md           # per-repo detail (from --from-repo)
    external-knowledge.md     # iwiki-extracted sections
  .teamai/
    domains.yaml              # business-domain dictionary
    domains.draft.yaml        # AI cluster draft
    domains.history.jsonl     # decision audit
    repo-whitelist.yaml       # repo allowlist
    source-marks.jsonl        # multi-source conflict tracking
    pending-review.jsonl      # deferred high-risk changes
  ~/.teamai/cache/repos/      # shallow-clone cache + LAST_SYNC

Surface area
============

CLI flags added to `teamai import`:
  --from-repo / --from-repo-list / --from-org / --bootstrap
  --depth / --ssh / --domain / --concurrency / --skip-aggregate
  --incremental / --max-repos / --exclude-archived
  --include-pattern / --exclude-pattern / --skip-import
  --iwiki-dual / --require-review

Tests
=====

~95 new unit tests across 14 new test files. Full suite passes
1165/1170 (the one remaining failure in types.test.ts pre-dates this
change). tsc clean (only the pre-existing recall.test.ts error is left).

Out of scope (tracked in Phase 6)
=================================

- TGit listOrgRepos real implementation (currently a stub)
- Cache LRU + 5GB cap + GC command
- Section-level diff with in-place anchor updates
- pending-review CLI to consume the deferred changes
- Domain-drift auto-apply workflow
- Global codebase doc lint

--other=phase5-team-codebase

* feat(agents): multi-CLI subagent sync + security hardening

- introduce YAML intermediate spec for team agents with renderers
  for claude / claude-internal / codebuddy / codex / codex-internal / cursor
- add agents path for codex / codex-internal / cursor in toolPaths
- pull renders per target tool format (.md / .toml); push reverses
  native files back to YAML, warns and skips on conflicts
- legacy agents/*.md still synced to claude-family for back-compat

Security fixes:
- clone.ts: drop token-in-URL, use http.extraHeader + sanitizeGitUrl
- ai-client.ts: execFileSync with shell:false, timeout, CLI whitelist
- path-safety.ts: assertSafePath + assertSafeResourceName
- import-local.ts: path traversal guard on --dir / output
- push.ts --skill / status.ts --agent: assertSafeResourceName
- env-commands.ts: env list masked by default, add --reveal flag

CSIG fixes:
- parseAgentYaml returns ParseResult instead of throwing
- fileContentEqual catch logs the error instead of swallowing

* feat: Phase 6 — Phase 5 hardening pass

Phase 5 shipped the team-codebase aggregation pipeline; in shipping
it we deliberately deferred a handful of reliability concerns to keep
each step deliverable. Phase 6 closes those gaps -- no new capability
surface, just turning the Phase 5 outputs into something safe to run
in production indefinitely. Six sub-steps, three independent and
three chained, merged here as one commit per project MR rules.

What's in the box
=================

P6.5  Global codebase doc lint  (src/codebase-{lint,cmd}.ts)
    A deterministic, AI-free cross-file lint over docs/team-codebase
    and the .teamai/ controls. 12 categories spanning anchor
    integrity, repo/whitelist consistency, sync staleness, frontmatter
    completeness, multi-source conflict, etc. `--fix` is intentionally
    narrow: only mechanical low-risk actions (orphan-md → archived,
    schemaVersion backfill, index counts refresh). High issues that
    aren't fixable end up in the skipped list. Exit 1 when any high
    remains so CI can gate merges. `--json` for downstream tooling.

P6.0  TGit listOrgRepos real implementation
    Replace the Phase 5.4 stub. Uses TGit's GitLab-style OpenAPI:
    GET https://git.woa.com/api/v3/groups/<encoded-path>/projects
    with token from the existing gfGetOAuthToken() helper. Multi-level
    group paths are URL-encoded whole. Pagination loops to maxRepos
    (200 default). Field mapping matches the GitHub side; primaryLanguage
    is left blank because the list endpoint doesn't return it. Errors
    are clean: 404 → "group not found or no access", other HTTP →
    "TGit API HTTP <code>: <body>", missing token → explicit hint
    about ~/.netrc / TAI_PAT_TOKEN. Token never reaches a log line or
    Error message.

P6.1  Cache lifecycle (LRU + size cap + GC command)
    The Phase 5 shallow-clone cache only grew. P6.1 adds an explicit
    metadata file at ~/.teamai/cache/repos/.cache-index.json; every
    successful clone/fetch in importFromRepo now refreshes its row
    via touchCacheEntry() (wrapped in try/catch + log.debug so cache
    bookkeeping never blocks the import). GC algorithm: stale-evict
    > 30d, then if over cap (5GB default, override via
    TEAMAI_CACHE_MAX_BYTES or --max-bytes) evict by ascending
    last_used until totalBytes ≤ cap*0.8. Eviction order is safe:
    fs.remove first, only then splice from index; failures land in
    skipped[]. getCacheStatus auto-heals index entries whose physical
    directory is gone. CLI: `teamai cache --status | --gc [--dry-run]
    [--max-bytes N] [--stale-days N] [--json]`.

P6.2  Section-level diff + in-place anchor updates
    The Phase 5 --incremental flag skipped clone but still rewrote
    docs/team-codebase/repos/<slug>.md whole, producing churn even
    when the source repo had no real change. P6.2 turns those
    summaries into anchored sections so unchanged content survives
    byte-equal across sync runs. Every `## title` block is wrapped in
        <!-- managed-by: import --from-repo, section: <slug>,
             source: ..., syncedAt: ... -->
        ## <title>
        <body>
        <!-- /managed-by: <slug> -->
    Section slugs derive mechanically from the title; duplicate slugs
    in one file get -2 / -3 suffixes so split / parse stay aligned.
    generateCodebaseMd is intentionally NOT touched -- the AI still
    emits one whole markdown blob, and the --workspace path that
    maintains the teamai-cli self codebase (docs/codebase.md) is
    untouched. Anchors only apply to per-repo team-codebase outputs;
    importFromRepo runs the AI output through mergeWithAnchors() per
    slug:
      - same body hash         → kept (old syncedAt + source preserved)
      - body changed           → rewritten with fresh body + new meta
      - present in fresh only  → added (appended)
      - present in old only    → removed (dropped)
    The frontmatter rule that actually delivers byte-equal: if all
    sections are kept, the prelude is also taken from old, otherwise
    from fresh. Without this tie-break the fresh `lastUpdated: <ISO>`
    in frontmatter would mtime-bump the file every run.

P6.3  pending-review CLI
    Phase 5.4 wrote .teamai/pending-review.jsonl when --require-review
    fired but provided no consumer. P6.3 adds the consumer:
        teamai review                # list (sorted by risk desc)
        teamai review <id>           # show details
        teamai review <id> --apply   # apply + drop + audit
        teamai review <id> --reject [--reason ...]
        teamai review --all-apply [--max-risk medium|low]
    Schema upgraded to {id, ts, kind, target, payload, source, risk}
    with backward-compat: loadPendingReview() normalises old rows
    on read, computing id from sha1(file|section|ts).slice(0,12) and
    inferring risk from a small hardcoded set of high-risk sections.
    iwiki-dual.ts now writes through appendPendingReview() so new rows
    always land in canonical shape. --apply only handles
    kind=codebase-section -- it calls patchManagedSection (P6.2),
    writes a fresh syncedAt, drops the row, appends an audit event.
    Other kinds gracefully degrade to "not auto-applicable". Atomic
    write through .tmp+rename so partial writes can't corrupt the
    jsonl.

P6.4  Domain-drift auto-apply workflow
    P5.3 detected drift and wrote history.jsonl, but gave the user
    no way to act on it. P6.4 turns drift into an actionable backlog:
    detectDomainDrift now dual-writes -- besides history.jsonl, every
    new event lands in pending-review.jsonl as kind=domain-drift,
    deduped 24h per url so a re-drifting repo stays one open item
    instead of growing. CLI:
        teamai domains drift                  # list
        teamai domains drift <repoUrl> --apply
        teamai domains drift <repoUrl> --lock
        teamai domains drift --apply-all [--threshold 0.8]
    Apply does the actual reassignment (splice old → push new,
    auto-create the new domain after a TTY confirmation; non-TTY
    refuses), updates confidence/signal from the recommendation,
    audits via appendHistory(reassign), drops the row, then calls
    regenerateAggregate so domain-*.md and index.md catch up. Lock
    sets RepoEntry.locked=true and clears stale drift items for the
    url. apply-all walks confidence-desc, applies above threshold,
    failures don't abort the batch.

CLI surface added
=================

  teamai cache --status | --gc
  teamai codebase --lint [--fix]
  teamai review [id] [--apply | --reject | --all-apply]
  teamai domains drift [url] [--apply | --lock | --apply-all]

Tests
=====

~150 new unit tests across 14 new test files. Full suite passes
1344 / 0 failing on this branch (Phase 5's pre-existing
recall.test.ts / types.test.ts failures were fixed in main between
Phase 5 and now and stay green here too). tsc clean.

Out of scope (tracked as Phase 7 in roadmap_jael.md)
====================================================

- Two-level domain hierarchy (e.g. AI/inference, platform/CI)
- Active cross-repo duplicate detection
- codebase.md ↔ search-index/recall integration
- agent retrieval effectiveness metrics

--other=phase6-team-codebase-hardening

* chore: stop tracking local drafts (roadmap, validation, .codebuddy)

These files exist in the working tree to support local iteration on
the team-codebase pipeline (personal roadmap notes, internal phase
acceptance reports, and the .codebuddy plan tree), but they should
not land in the upstream open-source repository. Add them to
.gitignore and untrack them via `git rm --cached` so they:

- stay on disk for local use
- stop showing up in `git status` for daily work
- disappear from the diff against upstream when sending PRs

This is a tracking change only -- no code or test behaviour is
affected.

* docs(readme): trim subagent phase notes and add Phase 5/6 commands

Three adjustments based on mentor feedback:

1. Remove the "Recall via subagent (Phase 1)" subsection from both
   README.md and README.zh-CN.md. The phase-numbered design note
   was useful during development but reads as roadmap detail on
   the public README. The downstream paragraph that explains what
   teamai recall actually returns -- the [<type>] tags plus the
   four-category index table -- is kept; that one is product
   behaviour, not phase trivia.

2. Tighten the public-facing tool list to the openly distributed
   editors (Claude Code, Codex, Cursor, CodeBuddy IDE, OpenClaw,
   WorkBuddy) and the matching ~/.claude/skills, ~/.codex/skills,
   ~/.cursor/skills, ~/.codebuddy/skills paths. No code changes:
   the underlying tool registry, sync paths, usage tracker, agent
   format dispatch, and AI-client probing are all left untouched,
   so existing setups keep working as before -- only the public
   README is shorter.

3. Add the recent commands that landed in PR #6 / #8 to the table:
     teamai import --from-repo / --from-repo-list / --from-org /
                   --from-iwiki [--iwiki-dual]
     teamai cache --status | --gc
     teamai codebase --lint [--fix]
     teamai review [id] [--apply | --reject | --all-apply]
     teamai domains drift [url] [--apply | --lock | --apply-all]

Documentation only; no code or test changes.

* fix(p5-p6): address audit findings — 2 blockers, 1 major, 5 medium

Independent review of the Phase 5 / Phase 6 codebase surfaced eight
issues; this commit fixes all of them. No new dependencies.

Blockers
========

1. iwiki anchor prefix mismatch broke `teamai review --apply`.
   iwiki-dual writes `<!-- managed-by: import --from-iwiki, ... -->`
   but the parser in section-patcher locked the prefix to
   `--from-repo`, so any pending-review item produced via
   --iwiki-dual --require-review threw `section not found` on apply.
   Fix: relax the parsing regex on both sides (parseSections and
   patchManagedSection) to accept `--from-(?:repo|iwiki)`. Writers
   stay as-is so the source-of-truth is still recoverable from the
   anchor metadata.

2. `--from-org` silently dropped private repos.
   `&type=public` was hardcoded into the GitHub list-org-repos URL,
   so on enterprise / internal orgs (mostly private) bootstrap
   produced a near-empty draft without error. Removed the query
   parameter -- relying on the caller's auth visibility (gh CLI or
   GITHUB_TOKEN) is the right default and matches GitHub's `type=all`.

Major
=====

3. tryEndpointPrefix returned success when the first page was empty.
   Combined with the bug above, an internal org with all-private
   repos returned [] from `/orgs/<x>` and never tried `/users/<x>`.
   Fix: when items.length === 0 && page === 1, return false so the
   outer code falls back to the user endpoint. Applied to both the
   gh CLI branch and the fetch branch.

Medium
======

4. ReDoS hardening on section-patcher anchor regexes.
   `[^>]*?` could be coaxed into exponential backtracking by hostile
   input. Replaced with `[^>\n]{0,256}?` -- bounded character class
   plus length cap. Applied to all four open/close anchor regexes
   (parseSections + patchManagedSection, both directions).

5. 10 MB hard cap on YAML / JSON config reads.
   loadDomains, loadCacheIndex, loadRepoList, and loadPendingReview
   now stat() before readFile and reject anything over 10 MB. Stops
   a malformed or hostile config file from blowing up memory.

6. Final path-safety check before per-repo writeFile.
   importFromRepo now calls assertSafePath() (an existing helper from
   PR #7) on the resolved repos/<slug>.md path. Defence-in-depth on
   top of the existing slug sanitisation; refuses to write outside
   the configured reposDir even if a future code path generates a
   weird slug.

7. SSRF guard + 50 MB response cap on outbound HTTP.
   gh-org and gf-org's fetch path now sets `redirect: 'manual'` and
   throws on any 3xx, and reads the body as a stream that cancels
   the reader and throws once total bytes exceed 50 MB. The gh CLI
   branch is unaffected -- gh handles redirects itself.

8. Backup before mergeWithAnchors fallback.
   When the existing repo file has corrupt / unclosed anchors,
   parseSections used to silently throw and importFromRepo fell
   back to a full-rewrite, losing every prior syncedAt timestamp.
   It now writes the old file to <repoMdPath>.bak (single overwrite,
   no accumulation) before doing the fallback wrap, so the prior
   state is recoverable.

Tests
=====

- iwiki-review-apply.test.ts (new): end-to-end -- iwiki-dual writes
  to pending-review.jsonl with --require-review, then `teamai review
  <id> --apply` is asserted to actually mutate external-knowledge.md
  (string contains the new body, anchor still says `--from-iwiki`).
  This was the regression that the parsing-regex fix unblocks.
- gh-org.test.ts (new): three cases -- private repos visible without
  type=public; first-page empty on /orgs/ falls back to /users/;
  /orgs/ 404 also falls back to /users/.
- section-patcher.test.ts: added cases for splitToSections /
  parseSections / patchManagedSection on iwiki-flavoured anchors.
- domains-store / cache-index / repo-list / review-store: each
  gained a test that writes an actual 11 MB file to a tmpdir and
  asserts the loader rejects it (not mocked).
- import-repo-merge.test.ts: corrupted-anchor case asserts the
  .bak file appears with the original content.

96 test files / 1356 tests pass / 0 failures (12 added by this
commit). tsc has only the pre-existing recall.test.ts error (carried
over from main, unrelated). Line length still ≤ 120 across all
touched files.

--other=fix-p5-p6-audit

* fix(test): add missing `type` field to recall.test.ts SearchIndexEntry mock

upstream CI's `Type check` step (npx tsc --noEmit) blocked PR #28 with:

  src/__tests__/recall.test.ts(62,7): error TS2741: Property 'type'
    is missing in type '{ filename, title, author, date, tags,
    tokens, votes }' but required in type 'SearchIndexEntry'.

SearchIndexEntry was extended in Phase 1 with a required `type:
KnowledgeType` field for the multi-bucket index, but the mock
factory in the recall vote test didn't follow. Local vitest doesn't
type-check the source so the bug never surfaced; upstream CI does
run tsc strictly and the typecheck step gates everything else
(unit tests, build, e2e), which is why PR #28 failed at the very
first step.

The test only exercises recallVote's counter path -- the `type`
field is never read -- so 'learnings' is just a representative
default; behaviour is unchanged.

Verified locally:
  - npx tsc --noEmit          → 0 errors (was 1)
  - npx vitest run recall     → 9/9 passing (unchanged)
  - npx vitest run            → 1360/1360 passing (unchanged)
  - npm run build             → success

* fix(test): resolve cross-platform path assertion failures in review-cmd.test.ts

--story=fix-github-actions-test-failures

Replace absolute path assertions with expect.stringContaining() to handle different tmp directory paths on macOS (/private/var) vs Linux (/var)

---------

Co-authored-by: jaelgeng <jaelgeng@tencent.com>
2026-06-12 17:28:43 +08:00

21 lines
641 B
TypeScript

import { defineConfig } from 'vitest/config';
export default defineConfig({
test: {
include: [
'src/__tests__/e2e/**/*.test.ts',
'src/__tests__/*-e2e.test.ts',
'validation/*.test.ts',
],
testTimeout: 60_000,
hookTimeout: 30_000,
// E2E tests spawn child processes and touch the real filesystem.
// Run test files sequentially to avoid race conditions (parallel
// file-level execution causes intermittent "Cannot find module
// dist/index.js" on GitHub Actions CI runners).
fileParallelism: false,
// Retry once: flaky tests recover, real bugs stay failed.
retry: 1,
},
});