mirror of
https://github.com/Tencent/teamai-cli.git
synced 2026-10-02 03:14:40 +08:00
* docs:知识库飞轮系统设计roadmap文档
* feat: phase 1 - agents support, multi-index search, recall subagent & todowrite hint
- Add agents resource type (pull/push/remove support)
- Add builtin-agents with teamai-recall agent
- Add multi-source search index (rules + skills + agents)
- Add todowrite hint injection for recall subagent
- Add phase1 e2e tests and unit tests for new features
- Update README (zh-CN + en) with agents documentation
* docs: update roadmap文档
* feat(search): P1.4 domain inference + search weighting
Introduce KnowledgeDomain (technical/ops/support/neutral) inferred from
frontmatter > tags > path > type fallback. Apply per-domain score
multipliers at search time (technical ×1.0, neutral ×0.85, ops ×0.5,
support ×0.3) plus a skills/rules type bonus (×1.1).
Bump SEARCH_INDEX_VERSION 2→3; legacy v2 indexes auto-rebuild on next pull.
--other=P1.4 domain inference and search weighting
* test(validation): add Phase 1 E2E suite and acceptance report
Move phase1-e2e.test.ts to validation/ alongside the acceptance report,
update import paths and vitest.e2e.config.ts to pick up the new location.
Fix pre-existing E2E isolation bug: loadStateForScope mock used
mockResolvedValue (shared object reference), causing test-1 to mutate
state.lastPullRev and trigger the rev-based early-exit in test-2/3.
Switch to mockImplementation(() => ({ lastPull: null })) so each call
gets a fresh object. All 5 E2E tests now pass.
--other=Phase 1 validation
* docs: update roadmap — add P4.6 learning promotion mechanism
Learning entries that accumulate confidence ≥ 0.90 (5+ upvotes, 2+
contributors, 14+ days old) are prompted for promotion to docs/skills/rules
based on content type, regardless of origin or domain. Add P4.6 step to
Phase 4, dependency graph, implementation detail, work-estimate table,
and architecture overview diagram.
--other=roadmap update
* docs(validation): update phase1 report to reflect E2E bug fix
All 5 E2E tests now pass. Update P1.2 status to ✅, remove E2E
failure notes and known issues, set pass rate to 100%.
--other=phase1 report update
* docs(validation): add runtime evidence appendix to phase1 report
Append Appendix A1-A4 with captured outputs from demo-phase1.test.ts:
agents sync paths, CLAUDE.md full content after injection,
search-index.json entry listing (4 types covered), and recall("api")
full STDOUT including envelope markers and domain-weighted scores.
Knowledge content in A3/A4 is anonymised (titles, authors, file paths
replaced with generic placeholders).
Also commit demo-phase1.test.ts as the reproducible evidence runner.
--other=phase1 report runtime evidence
* feat(search): query-aware domain weights + IDF scoring (v4 index)
改动 A — 查询感知 domain 权重:
将静态一维 DOMAIN_WEIGHT 改为二维查询-文档权重矩阵,新增
inferQueryDomain() 从查询 token 推断查询域。搜索 k8s/deploy 等
ops 相关问题时,ops 条目不再被打五折,technical 查询行为与原来一致。
改动 B — IDF 降权:
buildIndex() 末尾计算 df(文档频率)map 并写入索引;search() 中
为每个 token 匹配乘以 log((N+1)/(df+1))+1 的 IDF 权重,高频通用词
(api、deploy、error)自动降权,低频专有词(deepgemm、mooncake)
权重保持不变。
索引版本 3 → 4;isLegacyIndex() 新增 !index.df 判断触发重建。
旧 v3 索引在下次 teamai pull 时自动重建,search() 对无 df 字段的
旧索引降级为 idf=1.0,不报错。
--other=search quality improvements
* feat(import): add teamai import command — Phase 0 cold-start + P4.4 MR pipeline
## 新增命令:teamai import
支持五种知识来源:
- --dir <path>:扫描本地目录,AI 分类为 rule/doc/learning
- --from-claude:迁移 ~/.claude/rules 等 AI 工具规则目录
- --workspace:基于当前 git 仓库生成 codebase.md
- --from-mr <url>:从已合并 MR 提炼 learning + codebase 更新建议(P4.4)
- --from-iwiki <id/url>:从 iWiki Space 批量导入文档
## 新增核心模块
- src/utils/ai-client.ts:claude -p 子进程封装(并发 ≤ 3,60s 超时)
- src/utils/dedup.ts:Jaccard 相似度重复检测(14 天窗口,≥ 60% 标记 superseded)
- src/utils/iwiki-client.ts:iWiki MCP HTTP 客户端(JSON-RPC 2.0,零外部依赖)
- src/import-local.ts:本地文件扫描/AI 分类/交互确认/推送
- src/import-mr.ts:MR 三层解析/双路 AI 提炼/dedup/推送
- src/import-iwiki.ts:iWiki 导入(复用 import-local.ts 基础设施)
- src/codebase.ts:codebase.md 生成/增量更新
## 扩展现有接口
- providers/types.ts:GitProvider 新增可选 fetchMergeRequest() 方法
- providers/github/mr-fetch.ts:gh pr view 实现
- providers/tgit/mr-fetch.ts:gf mr 实现
- types.ts:新增 MRData/ClassifiedItem/LearningDraft/CodebaseSuggestion/ImportSession
## 测试 & 文档
- ai-client.test.ts:5 tests(spawn mock + 并发控制)
- dedup.test.ts:11 tests(关键词提取 + Jaccard + 文件扫描)
- validation/phase0-p44-acceptance-report-public.md:Phase 0 + P4.4 验收报告
--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线
* fix(import): support claude-internal CLI + gh REST API fallback + real PR demo
## ai-client.ts
- detectClaudeCli():按 claude → claude-internal 优先级探测,结果缓存,
进程内只探测一次,两者均不可用时给出清晰错误提示
- DEFAULT_TIMEOUT_MS:60s → 180s,适应大 diff 下 AI 提炼耗时
## providers/github/mr-fetch.ts
- gh CLI 不可用时自动降级到 GitHub REST API(内置 https,零依赖)
- 支持公开仓库无 token 访问;有 GITHUB_TOKEN 环境变量时自动携带
## import-mr.ts
- codebase 建议 JSON 解析:先提取 {…} 块再解析,
兼容 AI 在 JSON 前附加说明文字的输出格式
## index.ts
- import 子命令新增 --all 选项,跳过交互确认
## validation
- phase0-p44-acceptance-report-public.md A4:替换为基于真实 PR #2
的端到端操作记录(真实终端输出 + AI 真实生成的 learning.md
和 codebase-suggestions.json 完整原文)
--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线
* fix(codebase): require file-path prefix in module descriptions for agent guidance
codebase.ts 和 import-mr.ts 的提示词中"主要模块"格式要求均已更新:
- 之前:**模块名** — 功能说明(AI 可能只写中文名,无路径索引)
- 之后:**文件或目录路径** — 功能说明(明确要求带路径,便于 agent 定位)
codebase.ts: 全量生成 prompt 示例改为 **src/utils/git.ts** — 功能说明
import-mr.ts: codebase 建议 prompt 新增正确/错误示例,强制路径前缀
同步更新验收报告 A4:
- codebase-suggestions.json 从 8 条无路径条目 → 11 条带路径条目(真实重跑输出)
- 更新后 codebase.md 主要模块列表格式与更新前一致
--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线
* feat(codebase): upgrade workspace + MR prompts to A1-level documentation quality
codebase.ts — gatherRepoContext 扩展:
- 增加 package.json(依赖和 scripts)
- 增加入口文件(src/index.ts)命令注册全文
- 增加类型定义文件(src/types.ts)关键接口
- 文件树深度 maxdepth 3→4,过滤 dist/ worktrees/
- 截断上限:FILE_TREE 3000→5000 字符,DOC 1000→2000 字符
- 新增 META_MAX_CHARS 常量(2500 字符)
codebase.ts — 全量生成 prompt 重写:
- 提供完整 8+ 章节格式骨架(项目概述/技术栈/目录结构/数据配置/
核心数据流/关键接口/配置系统/性能可靠性/测试覆盖/备注)
- 目录结构要求带分组框树形图(┌─ 功能分组 ──┐ 风格)
- 技术栈要求表格含版本信息
- 项目概述要求带 emoji 核心能力 bullet list
- 核心数据流要求带缩进 → 的流程图格式
import-mr.ts — codebase 建议 prompt 升级:
- 新增 existingCodebaseMd 参数,注入现有文档全文作为格式样本
- AI 参考现有文档的分组和粒度生成风格一致的增量条目
import.ts — --from-mr 分支:
- 调用前读取 repoPath/docs/codebase.md 传入 existingCodebaseMd
- 确保 MR 增量更新与初始生成风格一致
- 复用已导入的顶层 fs 模块,删除内联 dynamic import
--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线
* docs(validation): replace A1+A4 codebase docs with real teamai-cli generated output
A1 附录:替换为 teamai import --workspace 对 upstream/main 真实生成的
codebase.md(含分组框目录树、表格技术栈、emoji 核心能力、流程图数据流)
A4 附录:
- Step 1:更新 codebase-before.md 为同一真实生成版本
- Step 2/3:终端输出更新为 3 条建议(主要模块 7 条/关键路径 4 条/架构决策)
- Step 3:learning.md 更新为最新真实 AI 输出
- Step 4:更新前后对比基于真实文档
- Step 5:飞轮闭环统计数字更新(3 条建议)
--story=132854480 【产品需求】teamai-cli Phase 0 冷启动 + P4.4 MR 提炼流水线
* feat(mr-hint): add SessionStart hook to hint AI about unimported merged MRs
P4.4 优化:在每次 Session 开始时检测当前 git 仓库的 origin remote,
查询近 7 天内已合入但尚未通过 teamai import 处理的 MR,并通过
additionalContext 提示 AI 在任务完成后建议用户运行 teamai import --from-mr。
核心实现:
- src/mr-hint.ts:新增 mrHint() 入口,支持 TGit REST API 和 GitHub gh CLI
双路查询;per-repo 磁盘缓存(30 天 TTL)避免重复提示相同 MR
- src/hooks.ts:注册 SessionStart hook,更新 TEAMAI_COMMAND_MARKERS
和 TEAMAI_HOOK_SUBCOMMANDS,同步 buildCursorHooks
- src/index.ts:注册 mr-hint 子命令
- 同步修复 hooks.test.ts、usage-tracking.test.ts、doctor.test.ts
中 todowrite-hint 加入后遗留的计数断言
--other=P4.4 MR 合入统一处理流水线(SessionStart 触发提示)
* feat(mr-hint): add GitHub REST API fallback when gh CLI unavailable
gh CLI 不在环境中时自动回退到 GitHub REST API(/repos/.../pulls),
逻辑与 providers/github/mr-fetch.ts 保持一致;
支持公开仓库无 token,有 GITHUB_TOKEN 时自动携带以提升限速上限。
--other=P4.4 MR 合入统一处理流水线(mr-hint GitHub REST fallback)
* docs(validation): update public acceptance report for P4.4 mr-hint trigger mechanism
新增 P4.4 触发机制优化章节,更新附录 A1(基于含 mr-hint 模块的代码库
真实生成),追加附录 A4.2(SessionStart hook 自动感知 merged PR 的真实
运行场景,含 GitHub REST API fallback 演示与幂等性验证)。
--other=P4.4 MR 合入统一处理流水线验收报告更新
* feat(ai-client): support login shell PATH + extend CLI candidates
- 用 bash -lc 包裹探测和调用,解决 ~/.nvm 路径下 CLI 不在 PATH 的问题
- 探测顺序扩展为 claude / claude-internal / codex / codex-internal /
codebuddy / workbuddy / openclaw
- 新增 shellEscape() 避免 prompt 中单引号破坏 shell 命令
- 超时从 180s 降至 120s
- 测试补充 execFileSync mock,修复预存 5 个失败用例
feat(import-mr): interactive codebase review loop + apply suggestions
- 新增 reviewCodebaseSuggestions():AI 实时修订循环,用户输入意见
→ AI 修订 → 再展示,直到 y 确认或 n 跳过
- --output 模式:apply 后写 codebase-after.md(完整 Markdown)
- repoPath 模式:apply 后写回 docs/codebase.md
- 修复 codebase.ts apply prompt,确保输出完整文档而非摘要
feat(import): add --existing-codebase option
allow 用户显式指定 before codebase.md 路径,不依赖团队仓库;
优先级高于从 repoPath/docs/codebase.md 自动读取
--other=P4.4 MR 合入统一处理流水线优化
* docs(validation): update A4 with real before/after codebase demo
用真实运行产物替换 A4 中的 codebase before/after 内容:
- Step 1 before:由 teamai import --workspace 在 PR #2 合入前的代码库生成
- Step 3 suggestions:teamai import --from-mr PR #2 的真实 AI 输出
- Step 4(新增):apply suggestions 后的 codebase-after.md,含 diff 展示
三阶段格式统一,均为同一版本 prompt 的真实运行产物
--other=P4.4 验收报告 A4 codebase before/after 更新
* fix(providers): add TGit REST API fallback + multi-shell CLI detection
- mr-fetch.ts: 新增 fetchTGitMRViaApi(),gf CLI 不可用时自动 fallback
到 git.woa.com REST API(使用 ~/.netrc OAuth token);
diff 获取失败时降级为空字符串而非中断流程
- ai-client.ts: detectClaudeCli() 对每个候选依次尝试
bash -lc → zsh -lc → which,覆盖 fish/CI 容器等非标准 shell 环境
--other=fix-risk-items
* docs: 验收文档typo更正
* fix(ai-client): multi-CLI compat + shell injection hardening
- ai-client.ts: detectClaudeCli 解析 CLI 绝对路径并校验存在性,
spawn 改为直调 absPath + 参数数组,删除 shellEscape 去 shell;
新增 buildCliArgs 区分 codex/codex-internal 用 'exec' 子命令、
其余 CLI 用 '-p',修复 codex 系 CLI 调用失败问题
- providers/tgit/mr-fetch.ts: execSync → execFileSync 数组参数,
消除 mrIid/repoArg 命令注入风险
- mr-hint.ts: TEAMAI_MR_HINT_CWD 增加 path.resolve + statSync
校验,非法路径静默跳过
- ai-client.test.ts: mock 适配新探测语义(command -v 返回路径
+ existsSync=true)
--other=phase0-p44-cli-compat-and-security
* feat(codebase): align with llm-wiki — frontmatter / index / lint / multi-source
参照 docs/llm-wiki.md 的持久化知识库理念,对 codebase 文档生成做四项优化:
- frontmatter:generateCodebaseMd 输出顶部注入 YAML frontmatter
(title / lastUpdated / source / generator / schemaVersion),
支持去重旧 frontmatter,便于跨会话溯源
- 索引体系:新增 generateCodebaseIndex 导出,从 codebase.md 抽取
二级章节 + 一句摘要 + 关键词,输出 codebase-index.md,加速 LLM 跨
会话定位
- 健康检查:新增 lintCodebaseMd 导出,AI 检测矛盾/过时/孤儿/缺失
四类问题,返回 LintReport(含 severity 分级),不修改文档
- 多源聚合:generateCodebaseMd 入参新增 learningsSuggestions 与
learningsDir,gatherLearningsContext 内部函数读取 learnings/*.md
frontmatter tags 做高频统计,融合 P4.4 MR 建议进 prompt
- prompt 模板新增"架构决策与权衡""已知限制与演进方向"两章节
- import.ts workspace 流程串入索引生成 + lint 报告打印
- types.ts 新增 LintIssue / LintReport 接口
- 新增 codebase.test.ts 11 个单元测试,全部通过
--other=phase4-codebase-llm-wiki-alignment
* feat(search): codebase-index.md high-weight + skip codebase.md
- 新增常量 CODEBASE_INDEX_FILENAME / CODEBASE_FULL_FILENAME /
CODEBASE_INDEX_WEIGHT_BOOST(×1.5)
- entryFromMdFile:同目录存在 codebase-index.md 时自动跳过
codebase.md,避免全量文档与索引文件重复命中
- search():codebase-index.md 命中时额外乘以 1.5 权重 boost,
recall 时章节摘要优先返回
- 兼容 subagent 与 fallback recall 两条路径,boost 在本地索引
阶段生效,无额外 AI 调用开销
- 新增 3 个测试用例(跳过逻辑 / 权重 boost / fallback 路径),
search-index.test.ts 共 26 tests 全通过
--other=phase4-codebase-index-search-boost
* docs(validation): refresh A1/A4 with llm-wiki-optimized codebase output
用新版 CLI(llm-wiki 优化后)重新执行 A1/A4,更新公开版验收报告产物:
- codebase-before.md:含 YAML frontmatter、架构决策与权衡、
已知限制与演进方向两新章节
- codebase-index.md:新增章节索引文件(11 行索引表)
- codebase-after.md:PR #2 应用建议后的更新版
- learning.md / codebase-suggestions.json:最新 AI 提炼产物
- 记录实际使用模型:claude-internal v1.1.9(DeepSeek-V3.1-Terminus)
- 保持 tgit → [internal] 等脱敏风格
--other=phase0-p44-acceptance-report-refresh
* docs(validation): replace codebase-after full content with before/after diff
A4 Step 4 中 codebase-after.md 展示方式由全文改为 unified diff,
更直观反映 MR 建议应用后的变更:新增"主要模块"章节(+8 行),
包含 import-local/import-mr/import-iwiki/codebase 四个关键模块说明
--other=phase0-p44-acceptance-report-refresh
* docs(roadmap): add Phase 6 — Phase 5 hardening
Phase 5 shipped the team-level codebase aggregation pipeline; in
shipping it we deliberately deferred several reliability concerns to
keep each step deliverable. Phase 6 captures those deferrals as a
focused hardening pass — no new capability surface, just turning the
Phase 5 deliverables into something safe to run in production
indefinitely.
Six sub-steps, three independent (P6.0 / P6.1 / P6.5) and three
chained (P6.2 → P6.3 → P6.4):
- P6.0 Real TGit listOrgRepos (replace stub)
- P6.1 Cache lifecycle (LRU + size cap + GC command)
- P6.2 Section-level diff with HTML-anchor in-place updates
- P6.3 pending-review CLI (review / apply / reject)
- P6.4 Domain-drift auto-apply workflow
- P6.5 Global codebase doc lint (cross-file consistency)
Appendix C dependency table updated with all six rows.
Phase 5 leftovers explicitly out of scope here (二级业务域 / 跨仓重复
检测 / search-index 联动 / agent 检索效果量化) are listed under the
new "遗留至 Phase 7" block.
--other=phase6-roadmap
* feat(import): Phase 5 — team-level codebase aggregation
Lift teamai-cli's codebase knowledge base from a single-repo, local view
into a team-wide, multi-repo aggregation that can be initialized in one
command, kept in sync incrementally, and audited end-to-end. The work
ships in five sub-steps but is a single coherent feature; merging as one
commit per the project's MR rules.
What's new
==========
P5.0 Business-domain dictionary (src/domains/*)
Zod schema + YAML store + AI batch clustering + single-repo
recommendation + interactive review CLI + jsonl audit log. Library
only; no CLI wiring at this step.
P5.1 Single remote repo import
`teamai import --from-repo <url>` shallow-clones into
~/.teamai/cache/repos/<provider>/<owner>/<repo>, reuses the existing
generateCodebaseMd scanner, and writes a per-repo summary at
docs/team-codebase/repos/<slug>.md. Three-tier auth (HTTPS+token /
HTTPS-anonymous / SSH); tokens are always redacted in error output.
P5.2 Batch import + domain aggregation
`--from-repo-list <yaml>` drives a whitelist with per-entry
{ url, domain, auth, priority } plus org entries (deferred). Failures
on one repo don't block siblings. Pure-template aggregator emits
docs/team-codebase/domains/domain-<name>.md and a top-level
docs/team-codebase/index.md from the per-repo files. Default output
root moved to docs/team-codebase to avoid colliding with the
existing teamai-cli self-codebase at docs/codebase.md.
P5.3 Incremental sync + domain drift
`--incremental` skips the full clone when the cache is hit and
LAST_SYNC is present, falling back to a fresh shallowClone if fetch
fails. After scan, a fresh recommendDomain pass is compared against
the existing assignment; divergent recommendations (different
domain + confidence > 0.5 + delta > 0.4) land in
domains.history.jsonl as a drift event without auto-reassigning.
AI failures never block the main flow. CI scheduling examples
shipped under examples/ci/ for GitHub Actions and Coding CI.
P5.4 Org bootstrap + iWiki dual output
`--from-org <org> --bootstrap` lists repos via gh api (paged with
/orgs/<o>/repos -> /users/<o>/repos fallback), AI-clusters them,
and walks the user through reviewDomains to produce both
domains.yaml and repo-whitelist.yaml before chaining into
importFromRepoList for the first full sync. TGit listOrgRepos is a
stub (Phase 6).
`--from-iwiki --iwiki-dual` extracts business APIs / external
knowledge / glossary into docs/team-codebase/external-knowledge.md
guarded by HTML comment anchors so future syncs replace bodies in
place. `--require-review` defers section writes to
.teamai/pending-review.jsonl. A small source-conflict helper flags
multi-source updates within a 24h window.
Filesystem layout introduced
============================
docs/team-codebase/
index.md # business-domain map + repo index
domains/domain-*.md # per-domain aggregate
repos/<slug>.md # per-repo detail (from --from-repo)
external-knowledge.md # iwiki-extracted sections
.teamai/
domains.yaml # business-domain dictionary
domains.draft.yaml # AI cluster draft
domains.history.jsonl # decision audit
repo-whitelist.yaml # repo allowlist
source-marks.jsonl # multi-source conflict tracking
pending-review.jsonl # deferred high-risk changes
~/.teamai/cache/repos/ # shallow-clone cache + LAST_SYNC
Surface area
============
CLI flags added to `teamai import`:
--from-repo / --from-repo-list / --from-org / --bootstrap
--depth / --ssh / --domain / --concurrency / --skip-aggregate
--incremental / --max-repos / --exclude-archived
--include-pattern / --exclude-pattern / --skip-import
--iwiki-dual / --require-review
Tests
=====
~95 new unit tests across 14 new test files. Full suite passes
1165/1170 (the one remaining failure in types.test.ts pre-dates this
change). tsc clean (only the pre-existing recall.test.ts error is left).
Out of scope (tracked in Phase 6)
=================================
- TGit listOrgRepos real implementation (currently a stub)
- Cache LRU + 5GB cap + GC command
- Section-level diff with in-place anchor updates
- pending-review CLI to consume the deferred changes
- Domain-drift auto-apply workflow
- Global codebase doc lint
--other=phase5-team-codebase
* feat(agents): multi-CLI subagent sync + security hardening
- introduce YAML intermediate spec for team agents with renderers
for claude / claude-internal / codebuddy / codex / codex-internal / cursor
- add agents path for codex / codex-internal / cursor in toolPaths
- pull renders per target tool format (.md / .toml); push reverses
native files back to YAML, warns and skips on conflicts
- legacy agents/*.md still synced to claude-family for back-compat
Security fixes:
- clone.ts: drop token-in-URL, use http.extraHeader + sanitizeGitUrl
- ai-client.ts: execFileSync with shell:false, timeout, CLI whitelist
- path-safety.ts: assertSafePath + assertSafeResourceName
- import-local.ts: path traversal guard on --dir / output
- push.ts --skill / status.ts --agent: assertSafeResourceName
- env-commands.ts: env list masked by default, add --reveal flag
CSIG fixes:
- parseAgentYaml returns ParseResult instead of throwing
- fileContentEqual catch logs the error instead of swallowing
* feat: Phase 6 — Phase 5 hardening pass
Phase 5 shipped the team-codebase aggregation pipeline; in shipping
it we deliberately deferred a handful of reliability concerns to keep
each step deliverable. Phase 6 closes those gaps -- no new capability
surface, just turning the Phase 5 outputs into something safe to run
in production indefinitely. Six sub-steps, three independent and
three chained, merged here as one commit per project MR rules.
What's in the box
=================
P6.5 Global codebase doc lint (src/codebase-{lint,cmd}.ts)
A deterministic, AI-free cross-file lint over docs/team-codebase
and the .teamai/ controls. 12 categories spanning anchor
integrity, repo/whitelist consistency, sync staleness, frontmatter
completeness, multi-source conflict, etc. `--fix` is intentionally
narrow: only mechanical low-risk actions (orphan-md → archived,
schemaVersion backfill, index counts refresh). High issues that
aren't fixable end up in the skipped list. Exit 1 when any high
remains so CI can gate merges. `--json` for downstream tooling.
P6.0 TGit listOrgRepos real implementation
Replace the Phase 5.4 stub. Uses TGit's GitLab-style OpenAPI:
GET https://git.woa.com/api/v3/groups/<encoded-path>/projects
with token from the existing gfGetOAuthToken() helper. Multi-level
group paths are URL-encoded whole. Pagination loops to maxRepos
(200 default). Field mapping matches the GitHub side; primaryLanguage
is left blank because the list endpoint doesn't return it. Errors
are clean: 404 → "group not found or no access", other HTTP →
"TGit API HTTP <code>: <body>", missing token → explicit hint
about ~/.netrc / TAI_PAT_TOKEN. Token never reaches a log line or
Error message.
P6.1 Cache lifecycle (LRU + size cap + GC command)
The Phase 5 shallow-clone cache only grew. P6.1 adds an explicit
metadata file at ~/.teamai/cache/repos/.cache-index.json; every
successful clone/fetch in importFromRepo now refreshes its row
via touchCacheEntry() (wrapped in try/catch + log.debug so cache
bookkeeping never blocks the import). GC algorithm: stale-evict
> 30d, then if over cap (5GB default, override via
TEAMAI_CACHE_MAX_BYTES or --max-bytes) evict by ascending
last_used until totalBytes ≤ cap*0.8. Eviction order is safe:
fs.remove first, only then splice from index; failures land in
skipped[]. getCacheStatus auto-heals index entries whose physical
directory is gone. CLI: `teamai cache --status | --gc [--dry-run]
[--max-bytes N] [--stale-days N] [--json]`.
P6.2 Section-level diff + in-place anchor updates
The Phase 5 --incremental flag skipped clone but still rewrote
docs/team-codebase/repos/<slug>.md whole, producing churn even
when the source repo had no real change. P6.2 turns those
summaries into anchored sections so unchanged content survives
byte-equal across sync runs. Every `## title` block is wrapped in
<!-- managed-by: import --from-repo, section: <slug>,
source: ..., syncedAt: ... -->
## <title>
<body>
<!-- /managed-by: <slug> -->
Section slugs derive mechanically from the title; duplicate slugs
in one file get -2 / -3 suffixes so split / parse stay aligned.
generateCodebaseMd is intentionally NOT touched -- the AI still
emits one whole markdown blob, and the --workspace path that
maintains the teamai-cli self codebase (docs/codebase.md) is
untouched. Anchors only apply to per-repo team-codebase outputs;
importFromRepo runs the AI output through mergeWithAnchors() per
slug:
- same body hash → kept (old syncedAt + source preserved)
- body changed → rewritten with fresh body + new meta
- present in fresh only → added (appended)
- present in old only → removed (dropped)
The frontmatter rule that actually delivers byte-equal: if all
sections are kept, the prelude is also taken from old, otherwise
from fresh. Without this tie-break the fresh `lastUpdated: <ISO>`
in frontmatter would mtime-bump the file every run.
P6.3 pending-review CLI
Phase 5.4 wrote .teamai/pending-review.jsonl when --require-review
fired but provided no consumer. P6.3 adds the consumer:
teamai review # list (sorted by risk desc)
teamai review <id> # show details
teamai review <id> --apply # apply + drop + audit
teamai review <id> --reject [--reason ...]
teamai review --all-apply [--max-risk medium|low]
Schema upgraded to {id, ts, kind, target, payload, source, risk}
with backward-compat: loadPendingReview() normalises old rows
on read, computing id from sha1(file|section|ts).slice(0,12) and
inferring risk from a small hardcoded set of high-risk sections.
iwiki-dual.ts now writes through appendPendingReview() so new rows
always land in canonical shape. --apply only handles
kind=codebase-section -- it calls patchManagedSection (P6.2),
writes a fresh syncedAt, drops the row, appends an audit event.
Other kinds gracefully degrade to "not auto-applicable". Atomic
write through .tmp+rename so partial writes can't corrupt the
jsonl.
P6.4 Domain-drift auto-apply workflow
P5.3 detected drift and wrote history.jsonl, but gave the user
no way to act on it. P6.4 turns drift into an actionable backlog:
detectDomainDrift now dual-writes -- besides history.jsonl, every
new event lands in pending-review.jsonl as kind=domain-drift,
deduped 24h per url so a re-drifting repo stays one open item
instead of growing. CLI:
teamai domains drift # list
teamai domains drift <repoUrl> --apply
teamai domains drift <repoUrl> --lock
teamai domains drift --apply-all [--threshold 0.8]
Apply does the actual reassignment (splice old → push new,
auto-create the new domain after a TTY confirmation; non-TTY
refuses), updates confidence/signal from the recommendation,
audits via appendHistory(reassign), drops the row, then calls
regenerateAggregate so domain-*.md and index.md catch up. Lock
sets RepoEntry.locked=true and clears stale drift items for the
url. apply-all walks confidence-desc, applies above threshold,
failures don't abort the batch.
CLI surface added
=================
teamai cache --status | --gc
teamai codebase --lint [--fix]
teamai review [id] [--apply | --reject | --all-apply]
teamai domains drift [url] [--apply | --lock | --apply-all]
Tests
=====
~150 new unit tests across 14 new test files. Full suite passes
1344 / 0 failing on this branch (Phase 5's pre-existing
recall.test.ts / types.test.ts failures were fixed in main between
Phase 5 and now and stay green here too). tsc clean.
Out of scope (tracked as Phase 7 in roadmap_jael.md)
====================================================
- Two-level domain hierarchy (e.g. AI/inference, platform/CI)
- Active cross-repo duplicate detection
- codebase.md ↔ search-index/recall integration
- agent retrieval effectiveness metrics
--other=phase6-team-codebase-hardening
* chore: stop tracking local drafts (roadmap, validation, .codebuddy)
These files exist in the working tree to support local iteration on
the team-codebase pipeline (personal roadmap notes, internal phase
acceptance reports, and the .codebuddy plan tree), but they should
not land in the upstream open-source repository. Add them to
.gitignore and untrack them via `git rm --cached` so they:
- stay on disk for local use
- stop showing up in `git status` for daily work
- disappear from the diff against upstream when sending PRs
This is a tracking change only -- no code or test behaviour is
affected.
* docs(readme): trim subagent phase notes and add Phase 5/6 commands
Three adjustments based on mentor feedback:
1. Remove the "Recall via subagent (Phase 1)" subsection from both
README.md and README.zh-CN.md. The phase-numbered design note
was useful during development but reads as roadmap detail on
the public README. The downstream paragraph that explains what
teamai recall actually returns -- the [<type>] tags plus the
four-category index table -- is kept; that one is product
behaviour, not phase trivia.
2. Tighten the public-facing tool list to the openly distributed
editors (Claude Code, Codex, Cursor, CodeBuddy IDE, OpenClaw,
WorkBuddy) and the matching ~/.claude/skills, ~/.codex/skills,
~/.cursor/skills, ~/.codebuddy/skills paths. No code changes:
the underlying tool registry, sync paths, usage tracker, agent
format dispatch, and AI-client probing are all left untouched,
so existing setups keep working as before -- only the public
README is shorter.
3. Add the recent commands that landed in PR #6 / #8 to the table:
teamai import --from-repo / --from-repo-list / --from-org /
--from-iwiki [--iwiki-dual]
teamai cache --status | --gc
teamai codebase --lint [--fix]
teamai review [id] [--apply | --reject | --all-apply]
teamai domains drift [url] [--apply | --lock | --apply-all]
Documentation only; no code or test changes.
* fix(p5-p6): address audit findings — 2 blockers, 1 major, 5 medium
Independent review of the Phase 5 / Phase 6 codebase surfaced eight
issues; this commit fixes all of them. No new dependencies.
Blockers
========
1. iwiki anchor prefix mismatch broke `teamai review --apply`.
iwiki-dual writes `<!-- managed-by: import --from-iwiki, ... -->`
but the parser in section-patcher locked the prefix to
`--from-repo`, so any pending-review item produced via
--iwiki-dual --require-review threw `section not found` on apply.
Fix: relax the parsing regex on both sides (parseSections and
patchManagedSection) to accept `--from-(?:repo|iwiki)`. Writers
stay as-is so the source-of-truth is still recoverable from the
anchor metadata.
2. `--from-org` silently dropped private repos.
`&type=public` was hardcoded into the GitHub list-org-repos URL,
so on enterprise / internal orgs (mostly private) bootstrap
produced a near-empty draft without error. Removed the query
parameter -- relying on the caller's auth visibility (gh CLI or
GITHUB_TOKEN) is the right default and matches GitHub's `type=all`.
Major
=====
3. tryEndpointPrefix returned success when the first page was empty.
Combined with the bug above, an internal org with all-private
repos returned [] from `/orgs/<x>` and never tried `/users/<x>`.
Fix: when items.length === 0 && page === 1, return false so the
outer code falls back to the user endpoint. Applied to both the
gh CLI branch and the fetch branch.
Medium
======
4. ReDoS hardening on section-patcher anchor regexes.
`[^>]*?` could be coaxed into exponential backtracking by hostile
input. Replaced with `[^>\n]{0,256}?` -- bounded character class
plus length cap. Applied to all four open/close anchor regexes
(parseSections + patchManagedSection, both directions).
5. 10 MB hard cap on YAML / JSON config reads.
loadDomains, loadCacheIndex, loadRepoList, and loadPendingReview
now stat() before readFile and reject anything over 10 MB. Stops
a malformed or hostile config file from blowing up memory.
6. Final path-safety check before per-repo writeFile.
importFromRepo now calls assertSafePath() (an existing helper from
PR #7) on the resolved repos/<slug>.md path. Defence-in-depth on
top of the existing slug sanitisation; refuses to write outside
the configured reposDir even if a future code path generates a
weird slug.
7. SSRF guard + 50 MB response cap on outbound HTTP.
gh-org and gf-org's fetch path now sets `redirect: 'manual'` and
throws on any 3xx, and reads the body as a stream that cancels
the reader and throws once total bytes exceed 50 MB. The gh CLI
branch is unaffected -- gh handles redirects itself.
8. Backup before mergeWithAnchors fallback.
When the existing repo file has corrupt / unclosed anchors,
parseSections used to silently throw and importFromRepo fell
back to a full-rewrite, losing every prior syncedAt timestamp.
It now writes the old file to <repoMdPath>.bak (single overwrite,
no accumulation) before doing the fallback wrap, so the prior
state is recoverable.
Tests
=====
- iwiki-review-apply.test.ts (new): end-to-end -- iwiki-dual writes
to pending-review.jsonl with --require-review, then `teamai review
<id> --apply` is asserted to actually mutate external-knowledge.md
(string contains the new body, anchor still says `--from-iwiki`).
This was the regression that the parsing-regex fix unblocks.
- gh-org.test.ts (new): three cases -- private repos visible without
type=public; first-page empty on /orgs/ falls back to /users/;
/orgs/ 404 also falls back to /users/.
- section-patcher.test.ts: added cases for splitToSections /
parseSections / patchManagedSection on iwiki-flavoured anchors.
- domains-store / cache-index / repo-list / review-store: each
gained a test that writes an actual 11 MB file to a tmpdir and
asserts the loader rejects it (not mocked).
- import-repo-merge.test.ts: corrupted-anchor case asserts the
.bak file appears with the original content.
96 test files / 1356 tests pass / 0 failures (12 added by this
commit). tsc has only the pre-existing recall.test.ts error (carried
over from main, unrelated). Line length still ≤ 120 across all
touched files.
--other=fix-p5-p6-audit
* fix(test): add missing `type` field to recall.test.ts SearchIndexEntry mock
upstream CI's `Type check` step (npx tsc --noEmit) blocked PR #28 with:
src/__tests__/recall.test.ts(62,7): error TS2741: Property 'type'
is missing in type '{ filename, title, author, date, tags,
tokens, votes }' but required in type 'SearchIndexEntry'.
SearchIndexEntry was extended in Phase 1 with a required `type:
KnowledgeType` field for the multi-bucket index, but the mock
factory in the recall vote test didn't follow. Local vitest doesn't
type-check the source so the bug never surfaced; upstream CI does
run tsc strictly and the typecheck step gates everything else
(unit tests, build, e2e), which is why PR #28 failed at the very
first step.
The test only exercises recallVote's counter path -- the `type`
field is never read -- so 'learnings' is just a representative
default; behaviour is unchanged.
Verified locally:
- npx tsc --noEmit → 0 errors (was 1)
- npx vitest run recall → 9/9 passing (unchanged)
- npx vitest run → 1360/1360 passing (unchanged)
- npm run build → success
* fix(test): resolve cross-platform path assertion failures in review-cmd.test.ts
--story=fix-github-actions-test-failures
Replace absolute path assertions with expect.stringContaining() to handle different tmp directory paths on macOS (/private/var) vs Linux (/var)
---------
Co-authored-by: jaelgeng <jaelgeng@tencent.com>
21 lines
641 B
TypeScript
21 lines
641 B
TypeScript
import { defineConfig } from 'vitest/config';
|
|
|
|
export default defineConfig({
|
|
test: {
|
|
include: [
|
|
'src/__tests__/e2e/**/*.test.ts',
|
|
'src/__tests__/*-e2e.test.ts',
|
|
'validation/*.test.ts',
|
|
],
|
|
testTimeout: 60_000,
|
|
hookTimeout: 30_000,
|
|
// E2E tests spawn child processes and touch the real filesystem.
|
|
// Run test files sequentially to avoid race conditions (parallel
|
|
// file-level execution causes intermittent "Cannot find module
|
|
// dist/index.js" on GitHub Actions CI runners).
|
|
fileParallelism: false,
|
|
// Retry once: flaky tests recover, real bugs stay failed.
|
|
retry: 1,
|
|
},
|
|
});
|