979 Commits
Author SHA1 Message Date
Jeff ab3f01f014 fix(tgit): fetch gf CLI over HTTPS with checksum verification (#773)
The gf CLI installer downloaded the platform tarball over plaintext HTTP
and piped it straight into tar with no integrity check, then executed the
extracted binary. ClawHub's security scan flagged this (finding T03).

Fetch over HTTPS to a unique temp file, verify its SHA-256 before extracting,
then clean up. The mirror 302-redirects to a content-addressed backend whose
URL path is the artifact's sha256 (the backend does not echo a checksum
header), so the expected digest is read from the redirect target, falling
back to x-checksum-sha256 for a direct serve. Fail closed when no digest is
advertised. Update the TGit provider reference doc to match.

Verified end-to-end on darwin-arm64: download, digest match, extract.
2026-09-24 13:19:47 +08:00
Yu Geigei da13c1119b feat(models): share gateway model profiles across agents (#675)
Add `teamai models` to point Claude Code, Codex, OpenCode, CodeBuddy, and
WorkBuddy at a team or personal model gateway.

- A team publishes `models/models.yaml` (id, name, base_url, api_key
  placeholder, model_groups by protocol); each member keeps the API key
  locally or as an environment-variable reference.
- `models switch` updates every installed, compatible agent (or `--agent`),
  asks for a missing key once, and accepts `--model` for the default.
- `teamai pull` re-applies the team's latest catalog to agents already
  switched to it; agents never switched are left alone.
- TeamAI records the managed fields before its first switch, skips agents
  whose managed fields changed outside TeamAI, and `models restore` puts
  the originals back. Claude's `/model` pick is not treated as a takeover.
- Claude family aliases map to matching gateway models or the default;
  Buddy entries use `${VAR}` key references; Codex edits are parsed and
  verified before writing.
- Push rejects an invalid catalog; user-scope uninstall restores model
  settings first.
2026-09-24 11:45:54 +08:00
RererrandClaude Fable 5.1 55b71efb69 feat(config): honor a relocated Claude Code config dir via toolRoots (#728)
Claude Code can move its whole user config directory with
CLAUDE_CONFIG_DIR, but teamai resolved every Claude path from the
team-wide toolPaths (.claude/...), so hooks, skills and rules were
written to ~/.claude, which that Claude Code never reads, and doctor
stayed green.

Add a member-level `toolRoots` key to the local config. In user scope
`scopedToolPaths` re-roots every path of the listed tool; a new
`hookToolPaths` does the same for writes that land in HOME regardless
of scope (hook injection/removal/listing, doctor's hook checks, the
local agent). `teamai init` records CLAUDE_CONFIG_DIR into
`toolRoots.claude` and keeps it across a re-init; `teamai doctor`
reports when the variable and the effective root disagree.

An explicit CLAUDE_CONFIG_DIR=~/.claude is recorded too: Claude Code
then reads .claude.json from inside the directory, so the MCP companion
file moves inside the root even when the root is unchanged. Accepted
roots are one directory in HOME or .config/<name>, the shapes
`toolInstallRoot` can express; the hook gates in hooks.ts now use it so
hook and resource gates agree. Only `claude` is accepted for now: it is
the one tool whose every user-scope write goes through toolPaths.

Closes #725

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-24 11:20:01 +08:00
Dongwoo JeongandDongwoo Jeong d729ad1750 fix(pull): pass force option through MCP reconcile (#762)
reconcileMcpAllScopes dropped options.force when calling
reconcileMcpForConfig, so `teamai pull --force` never forced an MCP
config overwrite even though force reaches every other reconcile path.

Co-authored-by: Dongwoo Jeong <dongwoo.jeong@lge.com>
2026-09-24 10:58:00 +08:00
ydflowandydflow 34e65e9017 fix(dashboard): gate the legacy dashboard-report command on a config (#772)
* fix(dashboard): gate the legacy dashboard-report command on a config

A current install writes only `teamai hook-dispatch`, whose
dashboard-report handler declares `requiresConfig` and is dropped when
no config resolves for the hook's cwd. The legacy subcommand stayed
ungated, so a hook left behind by an earlier install kept recording
dashboard events for every directory it fired in — including projects
that never set up teamai, whose sessions were then reported by
whichever scope pulled next.

Apply the same gate `teamai contribute-check` was given in #748, asked
about the session's cwd (or, for a host that sends none, the directory
the hook runs in), so the two legacy commands behave alike.

* test(dashboard): give the dashboard e2e fixtures a configured scope

Both suites drive `dashboardReport()` directly, so they bypass the
dispatcher and now hit the new config gate. Their payload cwd is a
fixture path that need not exist, and resolveConfigForDir then falls
back to the user scope — which nothing had seeded, so every hook call
returned early and no event was ever recorded.

Seed a user-scope config, which is what these pipelines already assume:
an installed team reporting its own sessions.

---------

Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-24 10:53:45 +08:00
dvd233 5576b38db7 fix(init): preserve additional roles selected at the prompt (#765) 2026-09-24 10:47:10 +08:00
Saul Moro cc2772114f fix(hooks): run every handler in the scope the dispatcher resolved (#752) (#769)
* fix(hooks): run every handler in the scope the dispatcher resolved (#752)

hook-dispatch resolves the scope from the payload cwd, then chdirs there.
The correction keywords, votes-sync and webhook-dispatch read their config
again through autoDetectInit() on the process cwd, so when that chdir failed
(a deleted worktree) and the host started the hook inside another project,
they used that project: its keywords, its member for the session's votes,
and its webhooks.

HookHandler.execute now receives the config the dispatcher resolved, and the
three handlers read their scope from it.

* fix(hooks): start the background pass when the hook's cwd is gone (#752)

The detached child was spawned in the payload cwd. On macOS and Linux a cwd
that no longer exists (a deleted worktree) fails the spawn, the error is
swallowed, and no background handler runs: no session-start pull, webhook or
update check. Start it in the temp dir instead, as the Windows WMI launch
already does. The child resolves its scope from the payload, and its handlers
now use that config, so the directory it starts in names no project.
2026-09-24 10:13:54 +08:00
Saul Moro 6a53b6f925 fix(lock): never rename over a lock that was just released or is still being written (#760) (#761)
* fix(lock): never rename over a lock that was just released or is still being written (#760)

A stale verdict lets the reclaimer rename over the lock, and two live states
read as stale: a file that vanished (released, and possibly re-created by a
third process before the rename) and an empty file (its owner opened it but
has not written it yet). lockState now tells live, stale and missing apart:
a missing lock gets one more exclusive create, and an empty one reads as
held until it has stayed empty for 5s (its owner died before writing it).

16 processes x 2000 attempts on one lock: 4-18 overlapping holders per run
before, 0 after, with acquisitions in the same range.

* fix(review): EPERM owner is live; cover the second pass; complete lock users (#760)

- process.kill(pid, 0) throwing EPERM means the owner is alive under another
  user (e.g. `sudo teamai`); it read as stale and was renamed over.
- Test the missing branch under the reclaim sentinel, where the rename lives.
- Comment and design doc state exactly what reads as stale (an unreadable
  lock still does); CHANGELOG lists every lock user.

* fix(review): publish the lock with its content in place (#760)

An empty lock expired after 5s, so a creator stalled between open and write
(SIGSTOP, sleep, slow I/O) could be taken over while it still held the lock.
exclusiveCreate now writes the payload to a private temp file and hard-links
it to the lock name (link fails with EEXIST like O_EXCL), so the lock never
exists empty. Without hard links it falls back to O_EXCL, where the 5s grace
still applies. The same create backs the reclaim sentinel.

* fix(review): O_EXCL fallback verifies its payload; unreadable lock is held (#760)

- Without hard links the lock is created with O_EXCL and sits empty until
  written. A creator stalled past the 5s grace could be replaced and still
  report success; it now reads the lock back and holds it only if its own
  payload is there.
- A lock that exists but cannot be read (EACCES: another user's 0600 lock)
  read as stale and was renamed over; it now reads as held.
- Use fse.link like every other lock operation, and mock it in update.test.

* fix(review): a wx creator holds only a lock it wrote inside the grace (#760)

- Without hard links, a reclaimer that read the lock empty could rename over
  it after the stalled creator had written and checked it. The creator now
  keeps the lock only if it wrote it within half the grace, so no reclaimer
  can have judged it stale; otherwise it gives it up.
- An unreadable lock logs a warning naming the file.
- Migration skips lock artifacts (<lock>.*.tmp, .sentinel, .new-*), which a
  contending pull creates and removes during the copy.
- Stale comments; the stall tests restore their spies in afterEach.

* fix(review): reclaim only a lock whose owner is provably dead (#760)

Each timing rule for locks that name no owner (empty, partly written) left
an ordering where two processes held the lock: a partial payload read as
stale at once, an older teamai stalled past the grace, a slow fallback
creator yielded but kept blocking. Only ESRCH now makes a lock stale; a lock
that names no owner, cannot be read, or has an EPERM owner is held, and a
warning names it so a crash leftover can be removed by hand. Drops the 5s
grace, the 2.5s creator limit and the read-back.

- parseLockContent: a bare legacy PID is valid JSON (a number) and was
  returned as unparseable, which only worked while unparseable meant stale.
- Migration skips only the real lock artifact formats, not every <lock>.*.

* test(lock): exercise the live-holder verdict; drop grace leftovers (#760)

- "returns false when a live process holds the lock" failed on the temp write
  before any verdict ran; link now rejects with EEXIST on the lock path.
- Remove Date.now/utimes staging the removed grace no longer reads, merge the
  duplicate empty-lock tests, restore the stall spies in finally.
- Docstrings and design doc: only a lock that names no owner or cannot be read
  is warned about; the sentinel-steal residual is stated as it is.
2026-09-24 07:36:20 +08:00
Ben YounesandBen Younes 506d4c9148 docs(readme): mark Copilot CLI usage/sessions/dashboard as supported (#766) (#767)
PR #666 (feat: add privacy-safe Copilot telemetry) shipped Copilot support for the Team Improvement columns (usage, sessions, dashboard), but the support matrix in the README files still showed em-dashes for those three cells. Flip them to checkmarks in all five language variants so the matrix matches the shipped, tested behavior.

Co-authored-by: Ben Younes <2910651+ousamabenyounes@users.noreply.github.com>
2026-09-24 07:31:51 +08:00
hiro-nikaitou 4d85bbc45e docs(test): name the live e2e suites in the integration stub (#770)
Signed-off-by: hiro-nikaitou <vieteviete@proton.me>
2026-09-24 07:30:19 +08:00
ydflow dad371c360 fix(agents): render codex TOML with multi-line literal strings (#754) 2026-09-24 00:04:27 +08:00
ydflowandydflow d3637ea270 fix(pull): gate the generic sync report on a tool that can receive it (#751)
* fix(pull): gate the generic sync report on a tool that can receive it

The generic branch of the sync loop reported the team repo's item count
for every resource type: `Synced N skills` counted what the repo holds,
not what landed. Skills are written per tool into that tool's own
directory, and a brand-new member has none of them yet — the handler
skips such a tool by design and only logs at debug. So the first pull
after `init` printed a success while nothing was on disk, which is the
phantom-success half of #585.

`hasInstalledTargetFor` asks the same question `getInstalledResourceTargets`
already asks, for one resource field, and the generic branch now gates
its report on it. Docs need no gate: they are copied to the team's own
docs directory, which the copy creates, so that report was already
truthful.

Only the report is gated. The writes still run, so a tool root created
later — Cursor makes `.cursor/` on first launch — is filled by the next
pull, and `pull --force` fills it now.

Fixes #585 (the generic branch). #597 fixed the same shape for rules
and left this one open; this closes it for skills.

* fix(pull): gate the agents sync report on a receiving tool too

The gate only covered skills, so the generic branch still printed
`Synced N agents` when no installed tool could receive them — the same
phantom-success shape #585 describes, one resource type over.
AgentsHandler then resolves no destinations and writes nothing.

`hasInstalledTargetFor` also duplicated the walk that
`getInstalledResourceTargets` already performs. That function now takes
an optional `field`, and the gate calls it, so reporting and writing share
one resolver instead of two that can drift — an external HERMES_HOME or an
unresolvable workspace no longer disagrees between them.

Docs, rules and env never reach this branch; hooks and mcp have no
tool-path field to probe, so they keep reporting unconditionally.

Tests add the agent case the file's header already claimed: no tool
directory suppresses the claim and nothing lands, and the claim returns
once the directory exists. The suppression case is RED without the gate.

---------

Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-23 23:59:54 +08:00
Saul Moro 72305c68b8 fix(skills): one share gate, actionable refusals, and a louder stub deploy (#747)
* fix(skills): one share gate, actionable refusals, and a louder stub deploy

Follow-ups from the review of #699:

- The Stop-hook reminder and `teamai skill get share` ask one gate
  (`shareGate`, through `contributeHintAllowed`). The hook skipped the
  unreadable-project-config check, and the legacy `teamai contribute-check`
  command, still called by hooks written before the dispatcher, checked
  nothing, so both nudged towards a command that refused.
- The gate reads only a config load failure as "cannot be loaded"; any other
  fault propagates (the hook withholds the reminder and logs it at debug).
- A `config` refusal says what failed (the file and position for a parse
  error) instead of pointing at `teamai doctor`, which cannot see a broken
  config. `skill show` now refuses through the same helper, so its hint moves
  from stdout to stderr like `skill get` and `skill path`.
- `pull` warns when the discovery stub cannot be deployed (it was an empty
  catch on the fast path and a debug line on a full sync), and so does the
  legacy prune.
- Error text no longer claims a reason was logged when none was: an empty
  config is named as empty, and `init` points at ~/.teamai/debug.log, where
  every path that deploys nothing now records why.
- `core` routes a bare `/teamai` right after a friction reminder to `share`,
  as the stub already said.
- The command drift guard rejects an unknown subcommand inside a group
  (`teamai skill gett core` passed before).
- The contribute-check e2e asserts the reminder's real text again; the usage
  guides (EN, zh-CN) and the design doc cover the config refusal, the gate and
  the reminder routing.

* test(learnings): retry temp-dir cleanup that races a detached git gc

A push into the bare origin can leave `git gc --auto` writing to
objects/pack after the test returns; the single rmdir in afterEach then
fails with ENOTEMPTY (seen on CI, Node 22 ubuntu, #747).

* fix(skills): gate skill show before its lookups, name the failing field

Review of #747:

- `skill show share` under a broken project config searched the user
  config's team repo and agents, which detection falls back to, and printed
  a `share` found there. It now asks the gate first and refuses on a config
  block before any lookup. With an empty user config it refuses instead of
  ending in a stack trace.
- A config that parses but fails validation reported the Zod JSON dump,
  whose first line is `[`, so the refusal said `config.yaml: [.`. Every
  config loader now reports each issue as `field: reason` on one line.
- The docs and skills that describe the share reminder or the refusal say
  it is withheld on a read-only source and while the config cannot be
  loaded, and that a validation failure names the field: product-overview
  and usage-guide (EN, zh-CN), designs/skill-serving.md, core/SKILL.md,
  contribute-member, setup-admin, join-member and manage-admin.

* fix(skills): no share reminder where teamai is not set up

`contributeHintAllowed` fell open with no config at all, so a caller other
than the dispatcher (the legacy `teamai contribute-check`) still nudged in
projects that never set up teamai, which have no team to share with
(#748). It now returns false there. Serving the skill stays fail-open.

* fix(contribute-check): gate the legacy reminder on the session's cwd

`teamai contribute-check --stdin` asked the share gate about the directory
the hook process started in, while the session analysis used the payload
cwd. Started outside the project, it could read the user config and nudge
where `teamai skill get share` refuses (a project config that does not
load). It now moves to the payload cwd first, as hook-dispatch does.

* fix(pull): keep a debug.log record when the stub cannot be deployed

The previous commit turned both deploy catches into `log.warn`, which is
muted in silent mode and never reaches debug.log, and a SessionStart pull
runs detached with its output discarded. So the automatic pull, the one
that deploys the stub for most members, lost the only persistent record
it had. Both catches now warn and write the same line to debug.log.

* fix(skills): skill show and list never answer for the fallback team

Known issues left by #747:

- `skill show <name>` and `skill list` on a config that exists but does
  not load ended in a Node stack trace, and under a broken project config
  they searched the user config detection falls back to: another team's
  repo and agents. Both now ask `detectTeam`, the one place that tells
  "this team", "no team" and "cannot tell, and why" apart (`shareGate` is
  built on it). Without a usable team, `show` answers from the package
  alone and `list` prints only the packaged catalog; both say what failed
  on stderr and exit 1.
- A teamai.yaml that exists but fails validation was reported as "not
  found. Check your repo path". It is now named as invalid, empty or
  unreadable, like the local config.

* fix(skills): the gate reads the session's directory, and no project config is skipped

Codex review of 5793758:

- A project-location config that is not `scope: project` (or omits
  `scope`, which defaults to user) was skipped without a word, so the gate
  read past it to the user config. It is now reported as unusable, unless
  it is the user config itself, as when running from HOME.
- The legacy `contribute-check` changed into the payload cwd and, if that
  failed, asked the gate about the directory the process started in. It
  now passes the payload cwd to the gate (`detectTeam(cwd)`), and a cwd
  that no longer exists holds no project config, so only the user config
  is asked, as #753 does.

* fix(logger): record warnings in debug.log

`log.warn` wrote to the console only and was muted in silent mode, so a
detached SessionStart pull, whose output is discarded, lost every warning:
the stub deploy failure and the legacy prune among them. Warnings now reach
debug.log like debug and error lines. `warnStubNotDeployed` drops the
second `log.debug` call, which printed the line twice under --verbose.

* fix(skills): the dispatcher gate reads the payload cwd; a symlink is not HOME

Codex review of b0583f5:

- The dispatcher's `contribute-check` and `pending-hint` handlers asked
  the gate about the process's directory, trusting hook-dispatch's
  `chdir`; when that failed, the launcher's config decided. They now pass
  `resolveHookCwd(stdin)`, as the legacy command does.
- The HOME exception for a non-project scope compared the config file's
  real path, so a project config symlinked to ~/.teamai/config.yaml passed
  for the user config. It is now decided by the project's location: its
  root is HOME.

* fix(skills): only a missing cwd falls back to the user config; load it once

Codex review of 15a5b5d:

- `detectTeam` read any failure to see the payload cwd as "deleted", so a
  cwd it could not open (no permission, a path through a file) fell back
  to the user config and could allow the reminder. Only ENOENT does now;
  anything else is `unusable` and withholds it.
- `skill show share` and `skill list` loaded the config twice, through the
  gate and then the team lookup, and reported a broken one twice. Both
  detect the team once and hand it to the gate.

* fix(logger): a file-only record instead of persisting every warning

b0583f5 made every `log.warn` append to debug.log, wider than the two
failures it was for, and it wrote unrelated subprocess errors to disk.
`log.warn` is console-only again; `log.persist` writes one line to
debug.log and never to the console. The stub deploy catches and the
legacy prune catch use both, so a detached SessionStart pull keeps the
record and --verbose prints it once.
2026-09-23 22:47:00 +08:00
Jeff a1a41dacb6 docs(readme): move the agent capability matrix under Quick Start (#763)
Put the product overview on the landing page so visitors can see Git-based Execution / Context / Improvement and per-agent coverage without opening a secondary doc.
2026-09-23 22:43:26 +08:00
Saul Moro 5502d8ec10 fix: keep TeamAI out of projects that never set it up (#748) (#753)
* fix(hooks): run team hooks only where TeamAI is set up (#748)

Hooks of a project-scope install live in HOME, so they fire in every
project on the machine. With no config for the hook's cwd they ran anyway:
the Stop share nudge (even with recall off), the TodoWrite recall nudge,
and local capture of sessions and skill usage that another project's
report later pushed to its team.

Handlers that need a team now declare requiresConfig; the dispatcher drops
them when neither a project nor a user config resolves. Only machine-level
work runs there: update check, session-start pull, local agent, package
pending hint. The legacy track paths skip recording the same way.

* fix(stats): keep skill usage in the scope that recorded it (#748)

Every scope appended to one ~/.teamai/usage.jsonl, so whichever project
pulled next reported every project's skills to its own team.

Usage now goes to <dataHome>/usage.jsonl of the scope that resolves for
the session's directory (detectProjectConfig(cwd) ?? loadLocalConfig(),
as the dispatcher does). Each report reads and truncates only its own
file; teamai stats shows the current scope. A machine's first user scope
starts with an empty file: what it held cannot be attributed.

* fix(review): one scope resolver for hooks, usage and stats (#748)

- resolveConfigForDir (config.ts) is the single resolution the dispatcher,
  the usage writers and readers, and teamai stats use. A host that sends no
  cwd (OpenClaw) resolves from the process cwd, and an unreadable project
  config yields null instead of falling back to the user scope.
- readUsageEvents / truncateUsageAfterReport require a scope; no path reads
  the old shared file by default. readKnownSkills reads its scope's file.
- Real-dispatch tests: no-config session leaves no trace, cwd-less host,
  unreadable project config.
- Docs: machine-level handler list, data-directory-layout note.

* fix(review): scope wording in CHANGELOG and readKnownSkills (#748)

* fix(review): gate legacy contribute-check and tolerate a missing cwd (#748)

The hidden 'teamai contribute-check' command, still called by hooks of
older installs, had no config gate, so it nudged in projects without
teamai. It now resolves the scope like the dispatcher and skips when
none resolves.

resolveConfigForDir no longer throws for a directory that does not
exist (simple-git refuses it); a hook naming a deleted worktree falls
back to the user scope.
2026-09-23 20:36:07 +08:00
Saul Moro 95cea46182 fix(manifest): reject namespace strings that are not safe path segments (#710)
* fix(manifest): reject namespace strings that are not safe path segments

A resource namespace becomes a directory component (skills/<ns>/,
agents/<ns>/, learnings/<ns>/) exactly as a project id does, but only the
project id was refined. Both manifests accepted a namespace like
'../../evil', and roles.yaml had the same hole.

Guarded at the manifest boundary, which is where projects.ts already claims
it is enforced and the only place these strings enter the process. The
existing isSafeNamespaceSegment guards in contribute.ts and
resources/agents.ts stay as defence in depth.

The doc comment pointed at the wrong layer: an id read from a hand-edited
config.yaml resolves through getProjectOrThrow, so it can only ever name a
project the manifest already validated. Corrected to say so.

No fixture or e2e manifest in the repo ships a namespace containing '/' or
'..', so nothing that parses today stops parsing.

* docs(changelog): note the manifest namespace guard

* fix(manifest): guard namespaces against traversal only, and say what is wrong

Review findings on #710.

The first cut reused SAFE_ID (^[A-Za-z0-9._-]+$) for resource namespaces. That
allowlist is right for a project id, which is also typed on the command line and
split on commas, but for a namespace it rejects far more than traversal: a team
whose skills live under a non-ASCII directory, or one with a space in the name,
would have stopped parsing although the directory is perfectly safe. The
namespace guard now tests what actually matters -- no path separator, no `:`
(drive-relative on Windows), no control character, and not `.` or `..` -- while
the project id keeps its narrower spelling.

Both live in src/manifest-schema.ts, which is what roles.yaml and projects.yaml
genuinely share; roles.ts no longer reaches into projects.ts for the schema.

Running the CLI against a manifest with `../../evil` showed the second half: the
zod failure escaped as a raw ZodError, so `teamai pull` printed a validation
object instead of a sentence. parseManifest now reports it the way the
hand-written checks beside it do, naming the entry:

  Invalid projects manifest: projects.0.resources.skills.1: resource namespace
  must be a single path segment (no '/', '\', ':' or control characters, and
  not '.' or '..')

Docs: the namespace rule is stated where each manifest is documented, in both
usage guides and in the multi-project design doc.

* fix(manifest): reject DEL and C1 controls in a namespace too

Review P2 on #710: the guard rejected only U+0000-U+001F while the error
message and the docs promise every control character, so U+007F and the
C1 range U+0080-U+009F still parsed. Range extended and the three ranges
covered in the projects and roles fixtures.

* fix(manifest): reject the Win32 dot/space aliases of . and ..

Review P1 on #710: the guard tested for the exact strings '.' and '..',
so '.. ', '.. .' and '...' passed. Win32 strips trailing spaces and
periods from a path component, so each of those reaches the filesystem
as '..' and escapes the namespace directory it was supposed to name.

A segment of nothing but dots and spaces is '.' or '..' in disguise and
is refused as such; 'a..' keeps parsing, since it stays inside its
parent. The project id, whose allowlist already excluded spaces, refuses
any run of dots for the same reason.

* fix(manifest): keep the project id rule untouched, scope the role docs

Review on #710.

The dot/space fix reached further than it needed to: tightening the id
to reject every run of dots also rejected '...', a working POSIX
directory name the id rule has always accepted, so a manifest that
parses today would have stopped. The id is back to the exact '.'/'..'
check it had before this PR, with a test that says so. Only the
namespace rule moves.

The role docs claimed every namespace under resources: follows the rule,
but roles.yaml's learnings: is accepted for backward compatibility and
ignored at runtime -- it names no directory, so holding an old manifest
to the rule would reject it over a field nothing reads. Both usage
guides and the design doc now name the fields that do take effect.

* docs(manifest): state the id and namespace rules separately

Review P2 on #710: after the id was left on its old rule, the docs still
described one rule for both, so they claimed a project id rejects any
name made only of dots and spaces while '...' parses. Each rule now
stands on its own in both usage guides and the design doc, and the
projects.ts comment says why the id is not held to the namespace rule.

* fix(manifest): reject a namespace with a trailing '.' or space

Review P1 on #710: refusing only names made entirely of dots and spaces
left the aliasing half open. Win32 strips trailing periods and spaces
from every path component, so 'frontend.', 'frontend ' and 'frontend..'
all resolve to 'frontend' -- one namespace reading and writing another's
directory, which is the isolation a namespace exists to provide.

The rule is now the trailing character itself, which covers the escape
('.. ' arriving as '..') and the aliasing in one test, and '.' and '..'
fall out of it. A dot inside a name ('alpha.v2') is untouched.

* fix(pull): a roles manifest that does not parse must not widen delivery

Review P1 on #710. resolveResourceNamespaces caught every failure from
loadRolesManifest and carried on with no role filter, which for a member
with no active project means an unfiltered sync: making the schema
stricter would have turned 'skills: [../../evil]' into 'deliver every
namespace', the opposite of what the guard is for.

The catch was covering two cases at once, because loadRolesManifest
throws both when the file is absent and when it is invalid. Only the
first is the legacy, unfiltered case, so it now throws a typed
RolesManifestMissingError and the catch reacts to that alone. An invalid
manifest propagates and pull fails the scope with the entry named --
exactly what an invalid projects manifest already does.

Verified against the real CLI: with 'skills: [evil/nested]' pushed to the
team repo, pull reports the failed 'Skills to deliver can be resolved'
check and the three delivered skills are left untouched; restoring the
manifest syncs them again.

* fix(manifest): narrow 'absent' to ENOENT, refuse Windows device names

Review on #710, two of the three findings; the third was a stale read of
the PR description, which the e2e section had already been rewritten to
match and which is now updated before the push rather than after.

readFileSafe returns null for every read failure and for an empty file,
so an unreadable roles.yaml was indistinguishable from one that was never
written -- and 'never written' is the one case allowed to relax role
filtering. projects.yaml had the same hole, where a null manifest means
'this team is not partitioned'. Both loaders now read the file directly:
ENOENT is absence, and a permission error, a directory or an empty file
is an error that fails the pull.

Windows opens a device for CON, NUL, AUX, PRN, COM0-9 and LPT0-9 in every
directory, extension or not, so a namespace spelled that way cannot be
the directory the manifest names. A name that merely starts like one
(console, community) is untouched, and the project id stays out of this
rule as it stays out of the others: it is a working POSIX name the id
rule has always accepted.

* fix(manifest): narrow every roles fallback, drop COM0/LPT0 from the device set

Review on #710.

The fail-closed change covered resolveResourceNamespaces but not the
other callers that fall back when the loader throws, so a malformed
manifest still reached an unfiltered sync by another route: bootstrap.ts
left the member role-less while auto-selecting the sole role,
resources/skills.ts and push.ts guessed the namespaces from the role ids,
and config.ts skipped the legacy migration and left the role unset. Each
now reacts to RolesManifestMissingError alone. roles-cmd.ts keeps its
broad catches on purpose: those commands report the error to the person
running them instead of deciding what to deliver.

Windows reserves COM1-COM9 and LPT1-LPT9, not COM0/LPT0, so the guard was
rejecting two ordinary directory names for no safety gain. Both are now
covered by the test that pins 'console' and 'community' as valid.

* fix(manifest): prove absence before trusting it, add the superscript devices

Review on #710.

ENOENT is not proof that a manifest is absent: a committed symlink whose
target is missing reads exactly the same way, and absence is the one
answer that lets a caller relax its filtering. The path is now lstat-ed
before absence is believed, so a dangling link is an error like any other
unreadable file.

Windows reads the superscript forms of 1, 2 and 3 as device numbers, so
COM and LPT followed by one of those join the ASCII-digit set.

The third finding, that resolveResourceNamespaces returns before reading
roles.yaml, is not a fail-open and is left as it is: that branch is
reached only when the member has no role, and a role-less member gets the
same unfiltered sync from a perfectly valid manifest, since every role
namespace below is gated on primaryRole. Reading the manifest there would
only add a new way for their pull to fail. The reasoning now sits in the
code beside the early return.

* fix(init): a broken roles manifest must stop init, a skipped prompt must not

Review on #710.

Both init paths swallowed every role-selection failure and carried on
without a role. A role-less config matches every role when hooks are
reconciled, so a manifest that does not parse installed exactly the hooks
it restricts.

Narrowing the catch to RolesManifestMissingError alone was too much: the
same block also absorbs a person skipping the role prompt, and a
non-interactive run reaches it, so init would have started failing for
anyone who does not pick a role. That case is now its own type,
NoRoleSelectedError, and the two lenient cases are named while a parse
failure or an unknown --role propagates. Both catch blocks read the same.

Also: ENOENT proves nothing about absence when the DIRECTORY is a
dangling link -- readFile and lstat on the file both report ENOENT -- so
the path's components are walked, and the first link that leads nowhere
is reported instead of being read as 'no manifest'.

Verified in init.test.ts (malformed aborts and writes nothing, absent
still initializes role-less) and against the real CLI in single-repo
mode: no manifest exits 0 with the role unset, '../../evil' exits 1, and
a valid manifest with --role sets primaryRole: frontend.

* chore(ci): re-run review against the rebased head

No code change. The Codex review workflow re-reviews on push, and the
PR body now carries the real-CLI matrix run on the rebased head.

* fix(manifest): expand '~' when reading a manifest, guard role ids used as fallback namespaces

Review findings on #710 after the rebase.

readManifestFile replaced readFileSafe/readFileIfExists, which expanded a
home-relative repo.localPath. Without the expansion a documented
`~/.teamai/...` path is searched under the current directory, read as
absent, and roles.yaml absence relaxes the filtering. The path is expanded
before both the read and the dangling-link walk.

When roles.yaml is absent, skills.ts and push.ts fall back to the role ids
as namespaces. A role id is an unrestricted string, so a value such as
'../../outside' reached path.join, and SkillsHandler.removeItem could
recurse outside the team repo. Both fallbacks now pass the ids through the
namespace guard and fail with the rule's message.

* fix(config): expand '~' in repo.localPath at the config boundary

A home-relative repo.localPath reached simple-git, the manifest readers and
every resource path unexpanded, so `teamai pull` failed with
'Cannot use simple-git on a directory that does not exist'. The schema now
expands it once, at parse time, so no consumer has to. expandHome moves to
utils/home.ts (fs.ts re-exports it) so types.ts can import it without
pulling in the fs helpers.

* fix(manifest): reject the CONIN$ and CONOUT$ console devices too

Review finding on #710. Windows opens the console for these names in any
directory, extension or not, the way it does for CON, so a namespace spelled
that way cannot be the directory the manifest means.

* fix(push): guard the role id silent mode uses as a namespace

Review finding on #710. With a valid manifest that maps the role to several
skill namespaces, silent push assigned primaryRole as the namespace without
the check the fallback path already has. It now goes through the same
guard, so 'frontend.' or 'CON' fail the push instead of becoming a path.

* fix(manifest): reject namespaces that differ only by case, classify the guard as breaking

Review findings on #710. Two namespaces of one resource type that differ
only by case (or Unicode normalization) name a single directory on the
default Windows and macOS filesystems, so a role scoped to 'frontend' would
read 'Frontend' too. Each manifest is checked when it loads; the pull path
checks roles.yaml against projects.yaml as well, since both share skills/,
knowledge/ and agents/.

The namespace guard makes a manifest that parsed before fail every pull, so
the changelog entry moves under Breaking Changes.

* fix(pull): check roles.yaml against projects.yaml for role-less members too

Review finding on #710. The cross-manifest case-alias check ran only when
the member had a role, so a project-only member pulling skills/Common with a
role's skills/common in the same repo was not stopped. roles.yaml is now read
whenever a projects manifest is in play; an absent one stays absent, a broken
one fails the pull as it does for a member with a role.

* fix(manifest): fold case the way filesystems do when comparing namespaces

Review finding on #710. The alias key was normalize('NFC').toLowerCase(),
which is not case folding: 'σ'/'ς' and 's'/'ſ' stayed distinct although
case-insensitive filesystems give each pair one directory. The key now
upper- then lowercases each code point on its own, which folds both pairs
and sidesteps the context-sensitive final-sigma rule. It errs toward
joining ('ß'/'ss', 'ı'/'i'), which can only reject a pair.

* fix(config): keep the config loadable when the roles manifest is broken

Review finding on #710. migrateLegacyRoleConfig rethrew a manifest parse
error, which loadLocalConfig caught and turned into null, so every command
reported "teamai is not initialized" — pull included, leaving the member no
way to fetch the fixed manifest. The migration now skips with a warning and
returns the config unmigrated.

That alone would widen delivery: a role-less member who would have been
migrated to 'hai' reached resolveResourceNamespaces' unfiltered early return
without roles.yaml being read. roles.yaml is now read for every member before
that return, so a broken one fails the pull (absent still means unfiltered).

* fix(status): report a resource type it cannot scan instead of crashing

Found running the real CLI on #710. scanLocalForPush resolves namespaces
through the roles manifest (agents via resolveResourceNamespaces, skills
when it falls back to role ids), and a manifest that does not parse now
throws there instead of being read as "no filter". status let that escape
as a stack trace after printing half its report. Status is where a member
looks to find out why pull failed, so it now warns with the error for that
type and lists the rest, as it already does for git status.

* fix(status): do not report "(none)" when a resource type could not be scanned

Review finding on #710. With every successful scan empty and one type
failing, status printed the warning and then "(none)", which reads as a
complete clean result. It now says "(none in the types that could be
scanned)" in that case.

* fix(config): a role the manifest could not resolve matches no role-scoped entry

Review finding on #710. When the legacy role migration cannot read the
roles manifest, the config stayed plainly role-less, and resolveMembership
reads role-less as "every role": hooks, MCP servers and env variables scoped
to roles reached a member the manifest would have made 'hai'. The pull
refused the manifest for skills, but those reconcilers still ran.

The migration now marks the in-memory config roleUnresolved, a runtime-only
field like dataHome that serializeLocalConfig drops and the schema strips on
load. activeRoleIds returns [] for it, so role-scoped entries reach nobody,
unscoped ones apply as before, and the reconcilers remove role-scoped entries
already installed. The next load decides the role again.
2026-09-23 19:22:54 +08:00
hiro-nikaitou 0b9586e9f9 fix(ai-client): mute share hint in AI child sessions (#746)
Signed-off-by: hiro-nikaitou <vieteviete@proton.me>
2026-09-23 17:21:43 +08:00
Jeff 1496855c66 docs(readme): add SkillHub fallback prompt for GitHub-blocked users (#745)
Add an alternative install prompt via skillhub.cn to the Chinese README
quick-start, for users who cannot reach GitHub.
2026-09-23 16:38:57 +08:00
ydflowandydflow 178eca705b fix(env): hold env keys to the identifier rule env.sh already assumes (#740)
`generateEnvFile` interpolated the key raw into `export <key>=<quoted value>`.
The value had `shellQuoteValue`; the key had nothing, so a key from the team
repo's env/env.yaml could produce a line that is not valid shell
(`export bad key='x'`) or one that runs code (`export FOO;cmd='x'`), in every
member's shell — env is pushable, so any member who can push can put such a
key in front of everyone else.

`parseEnvFile` already refused to read back any key outside
`[A-Za-z_][A-Za-z0-9_]*`, which made the asymmetry worse: the variable was
written into env.sh and then invisible to the CLI, so nothing reported it as
missing. That regex is now a module-level `ENV_KEY_RE` shared by both sides,
so write and read agree by construction rather than by two copies drifting.

The generator drops a non-matching key instead of failing: one member's bad
key must not take env.sh down for everyone, and the remaining variables are
still correct. `env add` rejects such a key up front with the offending name,
because the local command is where the mistake is still visible — accepting
it there would report success for a variable that never reaches a shell.

Refs #738

Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-23 16:27:12 +08:00
dayanandClaude Sonnet 5 9d7e50bb50 feat(pi): add Pi Coding Agent integration (#692)
* feat(pi): add Pi Coding Agent integration

Squash-rebased onto the latest upstream/main to resolve the PR's merge
conflict (main gained #693/#685/#694/#691/#681/#680/#666 since this branch
forked). This combines all commits from the PR into one, applied cleanly on
top of the new base — no functional changes from the previously reviewed
state.

The only real conflict was in src/__tests__/uninstall.test.ts, where diff3
split a test mid-body because of the repeated `});` boilerplate around it;
resolved by keeping both sides' new tests intact, in full.

* fix(pi): gate agent-hook files on the same per-slug marker ownership check

applyPiAgentHook()/removePiAgentHook() wrote and deleted teamai-agent-<slug>.ts
purely by path, with no ownership check — the same class of bug already
fixed for the main teamai-hooks.ts file, but never extended to the per-slug
HTTP agent-hook files. A user-authored file at that conventional path could
be silently overwritten on sync or deleted on uninstall.

Adds hasPiAgentHook(slug), mirroring hasPiHooks: injection now skips (with a
warning) instead of overwriting a same-named file without the
`[teamai] agent hook [<slug>]` marker, and removal skips instead of
deleting one. uninstall.ts's discovery scan now derives each file's slug and
checks the same marker before scheduling it for removal, instead of
matching by filename prefix alone.

* fix(pi): fail install_hook_rule instead of silently acking a skipped Pi agent hook

applyPiAgentHook warned and returned normally when the requested event has no
Pi equivalent or a same-named extension file exists without the TeamAI
marker. The caller in local-agent.ts wrote the manifest entry and acked
success regardless, so the server and local state believed the hook was
installed even though the file was never touched. Throw in both cases so the
existing install_hook_rule error path acks failure instead.

Also document the known limitation (shared with the OMP adapter) that a
scoped Pi uninstall is not durable across multiple projects on the same
machine, since the extension is one machine-wide file and hook dispatch has
no per-project exclusion check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 16:22:31 +08:00
Saul Moro ca6e51251f feat(skill): serve builtin skill content from the CLI, deploy a discovery stub (#699)
* feat(skill): serve packaged skill content from the CLI

Add `teamai skill get <names...> [--full] [--all]` and `teamai skill path
[name]`, so an agent can read built-in skill content that always matches the
installed CLI version instead of a copy deployed into its skills directory.

`get` prints SKILL.md byte for byte, frontmatter included, with {SKILL_DIR}
resolved to the absolute packaged directory so documented script invocations
run as-is. `--full` appends references/ and templates/, walked recursively and
sorted by relative path, because our references nest one level deeper than the
flat layout agent-browser assumes.

Content goes to stdout and every diagnostic to stderr, so the output stays
byte-exact when piped. An unknown flag warns and continues; an unknown name is
fatal, since acting on the wrong skill is worse than a retry.

`skill list` gains the served catalog and `--json`; `skill show` resolves
packaged skills before the installed-agent fallback, which is what keeps it
working once the deployed unit becomes a stub. Legacy directory names resolve
as aliases.

Refs #678

* refactor(skills): move content to skill-data and deploy a single stub

Agents now receive one file: `skills/teamai/SKILL.md`, a discovery stub of
about 2 KB whose description carries the triggers of every workflow and whose
body holds the commands that load them. The workflow content moves to
skill-data/{core,share,wiki}, which is never deployed and is printed by
`teamai skill get`.

Before this, `deployBuiltinSkills` copied three whole trees — 176 KB — into
every installed agent on every pull, so the text an agent read could disagree
with the CLI it documented until the member ran a pull, and a machine with ten
agents held ten copies. skills/ keeps its meaning ("everything here is
deployed"), which is what lets BUILTIN_SKILL_NAMES collapse to one name.

The stub is copied verbatim: no ensureSkillFrontmatter on the way out, so a
deployed copy that differs from the packaged one is a bug rather than a
variant. Recall no longer gates deployment, since the stub routes to every
workflow; the run-time gate for share lands with the pruning pass.

Uninstall learns the legacy directory names, which it would otherwise leave
behind on every machine that upgraded.

"skill-data" is added to package.json files, with a test that asserts it
through `npm pack`: without that entry every test still passes against the
repo and `skill get` serves nothing once installed from the registry.

Refs #678

* fix(skills): repair stale commands, broken refs and frontmatter

An audit of the three builtin skills found 60 defects. This fixes the ones that
survive the move to skill-data, and splits the two skills that were carrying
more than one job.

Stale CLI surface. The wiki skill advertised `teamai extract graph`, a command
that has never existed. The hand-written "ground truth" cheat sheet in the
teamai skill omitted 19 real commands while telling the agent that anything
missing from it could be checked with `--help` — which fails for the flags
`--help` hides. The cheat sheet is replaced by
`skill-data/core/references/commands.md`, rendered from the CLI's own command
table, with hidden flags marked as such. Two tests guard it: one regenerates
the file and diffs, the other resolves every `teamai …` string written anywhere
in skill-data against the command table and fails on an unknown command or
flag. That second test is the one that would have caught e151d43, 1ca43ac,
8bb0548 and 2ddb546 before they shipped; it carries a case proving it catches
`teamai extract graph`.

Paths. Everything the skills told an agent to read or execute assumed the skill
sat in the agent's own directory: `python3 scripts/scan_repo.py` from a cwd that
is the target repo, references cited by bare filename in two different
conventions, methodology paths handed to sub-agents inside input packets. All of
them now go through {SKILL_DIR}, which `skill get` resolves. The README template
nobody referenced is wired into the step that writes the knowledge-base README.

Frontmatter. None of the three skills declared allowed-tools, so the first
command of every flow hit a permission prompt. The wiki skill kept its trigger
words and prerequisites inside the description text; both move into the body.

Splits. `core` keeps what a daily user needs and `setup` takes day 0 and the
repo lifecycle, so the common path no longer carries ~500 lines of repo
creation. The wiki skill's phase procedures move into references/phases/, taking
its SKILL.md from 38.7 KB — larger than agent-browser's entire core — to 17 KB
with an index that says when to load each phase.

One contradiction is resolved in the author's text: the share skill mandated
that every generated document be written in Chinese, against global rule 1
("reply in the user's language") and this repo's own English rule. It now
follows rule 1.

Refs #678

* feat(pull): prune legacy builtin skill directories, gate recall at run time

Upgrading the CLI used to leave the pre-stub trees in place: cleanup skips
builtin names, and nothing else knew about them, so `team-wiki-codebase` and
`teamai-share-learnings` would sit in every agent directory on the machine
forever. Deployment now removes them first, in both the configured skills path
and Codex's shared `.agents/skills`. Unconditional, because those trees were
overwritten on every pull, so no local edit ever survived in them.

Recall moves from deploy time to run time. Before, `skipRecall` decided whether
the share skill reached the agent at all; with one stub routing to everything,
there is no directory to withhold, so `teamai skill get share` checks instead
and says what to enable. `--all` is exempt: an inventory dump is not an attempt
to run the workflow. With no team config to consult the gate fails open — a
fresh machine reading the docs gets the content rather than a refusal it cannot
act on.

deployBuiltinSkills drops its `skipRecall` option rather than keeping one that
no longer decides anything, and recall-toggle stops deleting a skill directory
it no longer owns.

Refs #678

* docs: align the nudge and the guides with CLI-served skills

`/teamai-share-learnings` was never a slash command of its own — it existed
because the directory was installed. The Stop-hook nudge now names `/teamai` and
carries `teamai skill get share` literally, so an agent can act on it without
having to infer the intent from the conversation. The five READMEs and both
usage guides follow.

Both guides gain the `skill get` / `skill path` commands and a short section on
why built-in skills are served rather than copied. `docs/designs/skill-serving.md`
records the contracts that are easy to break later: byte-for-byte output,
{SKILL_DIR} substitution, stdout/stderr discipline, recursive `--full`, the
run-time recall gate, the three drift guards, and when to retire
LEGACY_BUILTIN_SKILL_NAMES and the long-name aliases.

AGENTS.md and CLAUDE.md gain the rule that keeps this from rotting: skill-data
is treated like documentation, a behaviour change updates the affected skill,
commands.md is regenerated rather than edited, and new workflows go under
skill-data instead of into the stub.

Refs #678

* fix(skills): apply standards review findings

`teamai init` printed "Built-in skills (e.g. team-wiki-codebase) are ready to
use in your IDE now" seconds after deployment deleted that very directory. The
message now names the teamai skill and how it loads its workflows. Two comments
carrying the same stale name follow.

The five READMEs said different things: only the English one named the share
workflow and its command. All five now do.

`collectSupplementaryFiles` hand-rolled a recursive walk that
`listFilesRecursive` already does, including the ignore list that skips `.pyc`
and `__pycache__` next to the wiki's Python scripts. It calls the helper
instead. `listServableSkills` drops its fallback to `skills/`: a package without
`skill-data/` is broken, and serving the stub as if it were the content hides
that from the one error message built to report it.

Tests drop six non-null assertions for a helper that throws, per the repo's rule
against moving a compile-time error to run time.

AGENTS.md and CLAUDE.md record the exemption the branch created: skill content
printed by `skill get` keeps the language its author wrote it in, while the
command's own prompts, errors and listings stay English.

Refs #678

* fix(skills): apply spec review findings

Upgrading left the old references in place. Releases before the stub deployed
`skills/teamai/` with six reference files beside SKILL.md, and `teamai` is not a
legacy name to prune, so copying one file over that directory kept ~39 KB of
pre-stub instructions next to the new stub for good. Deployment now clears
everything the deployed unit does not contain before writing it, and a test
seeds the old layout to prove it.

Pruning reached neither reporting-only teams nor the Codex shared directory in
any test. The prune now runs before the reporting-only return, so a team that
switched to reporting-only still loses the stale trees, and the Codex
`.agents/skills` path is covered by a test. Excluded agents stay untouched, as
the enabledAgents whitelist documents.

The served content still routed to skills that no longer exist: five mentions of
`teamai-share-learnings` and one `/team-wiki-codebase --update`, which is the
rule this branch itself added being broken on arrival. One second-hop path
inside a sub-agent input packet was still relative.

The share skill's document template, frontmatter table and tag taxonomy move to
`references/doc-template.md`, taking the always-read body from 3 701 to 2 471
bytes. Two tests now assert what nothing guarded: every served skill's
frontmatter name matches its directory and declares allowed-tools.

A new e2e file runs the built CLI the way an agent does: every listed skill is
servable and byte-identical bar the resolved placeholder, the wiki scripts run
from the directory `skill path` prints, an unknown name exits 1 with empty
stdout, a hallucinated flag warns and still serves, and `--full` appends the
nested references in sorted order.

Refs #678

* fix(skills): prune only directories the CLI owned, gate every content path on recall

- LEGACY_BUILTIN_SKILL_NAMES drops teamai-workflow and teamai-import: they were
  reserved in the old guard set but never packaged, so a directory by either
  name is the user's own skill. Test: user-created skills with those names
  survive pull.
- The recall gate now covers skill get --all (blocked skill skipped, named on
  stderr), skill path (refused) and skill list --json (blockedByRecall, path
  null). skill list reads the flag from the catalog instead of re-checking.
- skill get [names...]: the positional is optional so --all is reachable from
  the real CLI; Commander used to fail with 'missing required argument'.
  Covered by the skill-serving e2e.
- commands-reference renders Commander's variadic marker (<names...>); the
  snapshot is regenerated.

* test(skills): drive the recall gate through the real CLI, guard the stub description budget

- skill-serving e2e: a HOME with a team whose recall is off; skill get share,
  --all, skill path share and skill list --json each withhold share, and all
  serve it after recall enable. The earlier HOME has no team config and fails
  open, so the gate was never exercised through dist/index.js.
- skill-content test: the stub description stays within 1024 characters.
- skill-commands-exist also scans the deployed stub.
- Docs and PR lead with content versioned with the CLI; the size numbers are
  measured (stub description 0.8 KB, body 1.3 KB; --full 32/36/115 KB).
- Content audit against origin/main: every file has a counterpart. Fixes:
  {SKILL_DIR} defined in core/setup/share where the references are listed, the
  wiki overview draws the served layout, team-wiki-codebase kept as a trigger
  word in the stub description.

* fix(skills): gate skill show on recall, prune Codex's shared dir only from Codex

Review follow-up on #699.

- `skill show <served skill>` refuses a recall-blocked skill with the same
  message and exit code as `skill get` / `skill path`; it printed the
  directory those two withhold.
- A skill resolved from skill-data/ is classified `[builtin]` directly.
  BUILTIN_SKILL_NAMES only knows the deployed stub, so `skill show core`
  reported `[local-only]` beside a package path.
- pruneLegacyBuiltinSkills reaches `.agents/skills` only on Codex's own
  pass. Another enabled tool's pass deleted Codex's legacy copies while
  Codex was excluded, against the enabledAgents guarantee.
- The share skill and its references are written in English; the generated
  document still follows the session's language. The AGENTS.md exception
  for Chinese skill-data output is dropped.

* fix(skills): one resolver for served skills, legacy names kept out of push, wiki in English

Review follow-up on #699.

- resolveServableSkill is the only way to obtain a PackagedSkill outside
  skill-content.ts; it returns `blocked` instead of the skill, so `get`,
  `path`, `list` and `show` inherit the recall gate by construction.
- push never offers `team-wiki-codebase` / `teamai-share-learnings` as new
  user skills: between the upgrade and the first pull they are still on
  disk (isCliOwnedSkillName).
- `recall disable` removes the legacy `teamai-share-learnings` directory
  again (LEGACY_RECALL_SKILL_NAMES), skipping excluded agents.
- `skill list` prints the packaged catalog before `teamai init`, with a
  hint for the team half, instead of failing on the team listing.
- skill-data/wiki (SKILL.md, 14 references, 2 scripts) translated to
  English. Generated document names follow one glossary; validate_kb.py
  still recognises headings of knowledge bases built by the previous
  release, matched by code point so the source stays ASCII.

* fix(skills): prune only files the CLI packaged, let local skills win by name

Review follow-up on #699.

- PACKAGED_SKILL_FILES lists every file a release ever wrote under skills/,
  as the union of `git ls-tree -r <tag> -- skills/` over all 91 tags. The
  prune removes those paths and the directories they leave empty; a file a
  member added is kept, its directory with it, and pull says which and why.
  The stub directory loses its six known references by name instead of
  "everything that is not SKILL.md". Python bytecode of a script we shipped
  counts as ours, so a __pycache__ does not strand the tree.
- locateSkill searches the team repo, then installed agents, then the
  package. A directory a member created under `codebase`, `default`,
  `learning` or `share` is the skill they asked about, and the recall gate
  does not apply to it.
- A guard test fails when a file ships under skills/ without being recorded
  in the manifest, which a later migration would otherwise leave behind.

* fix(skills): close the last recall bypass, uninstall Codex's shared stub

Review follow-up on #699.

- `skill path` takes a name, always. The argument-less form printed the
  `skill-data/` root, and `<root>/share/SKILL.md` is readable from there —
  the content the gate withholds one command over.
- uninstall discovers skills in Codex's shared `.agents/skills` root, where
  resolveSkillDestination puts the stub whenever the skill already lives
  there. Without it, uninstall reported success and left it behind. Codex
  only, as the legacy prune already does.
- core/SKILL.md said team sharing is enabled by default; getRecallSharing
  defaults it to false. It now says recall is off by default and names
  `teamai recall enable`.

* fix(skills): uninstall by the same ownership rule as pull, quote {SKILL_DIR}

Review follow-up on #699.

- uninstall removed a CLI-owned skill directory whole, undoing one command
  over the guarantee pull makes. It now removes the PACKAGED_SKILL_FILES
  paths through the same removeOwnedFiles, keeps a directory holding a file
  the member added, says which one, and tells the confirmation prompt so it
  no longer promises a directory it will keep. A team-repo skill is synced
  whole and still goes whole.
- Served shell commands quote the placeholder: `python3 "{SKILL_DIR}/..."`.
  Unquoted, an install path with a space ("Program Files", "Application
  Support", a Windows path through Bash) splits into two arguments and the
  documented invocation fails. A test fails on an unquoted occurrence after
  any command word, in SKILL.md or any reference.
- wiki/references/overview.md said the methodology, scripts and agent specs
  are deployed into agent directories. They are not: only the stub is, and
  the rest is served from the installed CLI.

* fix(skills): deploy the stub in reporting-only mode, drop the stale list alias

Review follow-up on #699.

- Reporting-only HTTP pull pruned the legacy trees and deployed nothing, so
  a member on an HTTP team came out of the upgrade with no built-in entry
  point at all. The skip predates CLI-served content: it existed because the
  only deployable unit then needed a team repo. The stub does not — its
  workflows are printed by the installed binary, and `skill get wiki` is a
  local knowledge-base generator that never touches a repo. The stub now
  deploys in every mode, and `reportingOnly` goes with the branch it gated:
  nothing else read it.
- `teamai skill list` called itself an alias for `teamai list skills
  --source all`. It has not been one since it started printing the CLI-served
  catalog underneath. Both descriptions, the generated command reference and
  both usage guides now say what it does.

Refs #678

* fix(skills): carry the TGit provider guide into the served setup skill

#724 landed `skills/teamai/references/provider-tgit.md` and repointed
setup-admin.md and join-member.md at it. Rebasing onto that left the new
file in a tree this branch no longer deploys, and the pointers in bare
`provider-tgit.md` form the served skills do not use.

- move it to `skill-data/setup/references/`, beside the two files that
  cite it, so `teamai skill get setup --full` serves it
- rewrite every pointer to it as `{SKILL_DIR}/references/provider-tgit.md`
- list it in the setup skill's reference table
- add `references/provider-tgit.md` to PACKAGED_SKILL_FILES, so the prune
  removes it from members who pulled a release that shipped it

* fix(skills): make the blocked catalog entry unrepresentable, drop unsafe casts

Review findings from the standards axis, plus the doc half of the prune
count.

- `SkillCatalogEntry` allowed `{blockedByRecall: true, path: '/…'}`, an
  invariant `skillCatalog` then upheld by hand. Split it on
  `blockedByRecall`, so the withheld directory is a type error rather than
  a review catch. Both variants keep the `path` key, so the
  `skill list --json` shape is unchanged.
- `command.commands as Command[]` stripped commander's `readonly` in three
  places. `for…of` and `.find` need no cast.
- `docs/designs/skill-serving.md` still said the prune removes six
  `teamai/references/*.md`; provider-tgit.md makes it seven.

* fix(skills): keep publishing a skill reachable when recall is off

Publishing a skill is `teamai push --skill`, which never consulted recall
(`src/push.ts` names it nowhere). On main the flow shipped in the teamai
skill, ungated. Moving `contribute-member.md` under `share` put it behind
the recall gate, so with recall off — a new team's default — the core
routing table sent the agent to `teamai skill get share`, which exits 1
and tells it to enable recall. Wrong advice for a flow recall does not
touch, and no other path to the instructions.

Move the file to `core`, the skill that already owns `push`, and split the
routing row so publishing and session learnings stop sharing one
destination. The gate itself is right and stays: learnings do need recall.

`share/SKILL.md` already called this "a different flow"; now it points at
`teamai skill get core --full` instead of at its own references.

PACKAGED_SKILL_FILES is unchanged: the legacy path a pre-stub release
wrote is still `teamai/references/contribute-member.md`.

* fix(skills): back up what the prune removes, so no edit is a one-way door

Review finding: `removeOwnedFiles` proves ownership by pathname and
deletes without reading the file, so a member's edit goes with it.

For a path the current package still ships that changes nothing: the old
deployment overwrote it with `overwrite: true` on the same three triggers,
so the edit died either way, at the same moment. The case the objection
gets right is a path a retired release shipped and the package no longer
does — the overwrite never reached it, so the edit did survive, and the
prune is the first thing to remove it.

Copy every pruned file to `~/.teamai/removed-skills/<date>/<tool>/<skill>/`
before removing it. Outside every agent directory, so nothing reads it back
as a skill.

Verifying contents against a hash of each released version was the other
way out, and it is worse: anything not byte-identical is then kept, so one
CRLF checkout on Windows — a platform this project supports — leaves the
whole 176 KB in place and reports success. Backing up gives the same
guarantee without betting the migration on byte equality.

Uninstall keeps deleting outright: there the member asked for the files to
go.

* fix(skills): let no backup failure authorise a delete, give each root its own

Two holes in the backup the previous commit added, both reported in review.

The copy's failure was swallowed at debug level and the delete went ahead
regardless, so a full disk or a read-only home turned the migration back
into the data loss the backup exists to prevent — and the log still named
a backup directory that held nothing. A file whose copy fails is now kept,
counted, and named at warn level; `removeOwnedFiles` returns what happened
instead of a bare boolean, and only a run that copied something names the
directory.

The backup path was `<date>/<tool>/<skill>` with `overwrite: true`, so the
second copy of a name silently replaced the first. Codex prunes the same
skill from `.codex/skills` and the shared `.agents/skills`, and two pulls
share a date. The path now carries a per-run id and the skill root, and the
copy refuses to overwrite rather than clobbering a copy it cannot replace.

Tests cover both: a file where the backup tree must start makes every copy
fail, and the two Codex roots land in separate directories. Each fails
against the previous commit.

* fix(skills): stop at a symlinked root, archive only what is retired

Three review findings, all in the prune.

A symlinked skill directory was walked through. `readdir` follows the link,
every path under it matches a packaged name, and the delete lands in
someone else's checkout. Ownership now stops at the link: the root is
lstat'd, a symlink is refused, and link and target are left alone.

The stub directory was pruned against the full historical file list, which
includes the SKILL.md written one line later. Deployment runs on every
session start, unchanged revision included, so that archived an identical
copy per session forever. Only paths this release no longer ships are
archived now.

Backups were written under the tool's base directory, which under project
scope is the repo root, so they landed in the working tree outside the
generated .teamai/.gitignore. They go to the machine's home.

Also: `skill show <packaged>` resolved the team before the package, so it
failed on a machine that never ran `teamai init` for content that needs no
team. Packaged names resolve first and print without the team-dependent
fields.

* fix(skills): stop at the first link above a skill dir, report a half prune

Findings from a self-review run before pushing, plus the two from the last
review round.

The symlink guard was one level too low. It lstat'd the skill directory, so
the common shape — `~/.claude/skills` itself linked at a dotfiles checkout —
walked straight through: every directory under the link is real. The guard
now walks each component below the tool's base directory and stops at the
first link, which covers the prune and the stub write with one check.
Components at or above the base are not checked: a home directory under a
link is ordinary, and refusing there would disable deployment on those
machines.

The symlink branch borrowed the foreign-files message, so a member was told
"delete the rest yourself" about a directory nothing had touched. Following
that destroys what the guard just protected. It has its own sentence now, in
pull and in uninstall.

`remove()` was not fail-closed the way the backup is: a read-only parent
left the tree half-pruned under a debug line, and a `walkFiles` that threw
returned success. Both are recorded in `notRemoved` and reported.

The backup path gained the base directory: `inheritUserScope` deploys the
user base and then the project base in one process, same tool, same root,
same skill name, and `errorOnExist` turned that collision into files the
second pass could neither archive nor prune.

Docs corrected against the code: the archive path, the tag count (98, not
91, and `teamai-wiki` is excluded), the version line, and the size table.

* fix(skills): route the nudge and skill publishing where they land, classify legacy names as ours

The Stop-hook hint said "run /teamai", but bare /teamai prints the menu and
stops, so following the primary suggestion never reached the share workflow.
It now names an invocation the core skill routes to share, with the
`teamai skill get share` fallback kept. The four docs that quote the hint follow.

The setup skill sent "publish one skill" to `teamai skill get share`, which
handles session learnings and is refused when recall is off (the default);
reusable-skill publishing lives in core's contribute-member reference and needs
no recall. The routing row and the two references that repeated it now point
there.

classifySkill checked BUILTIN_SKILL_NAMES alone, so until the first pull pruned
them, team-wiki-codebase and teamai-share-learnings showed as [local-only]. It
now uses isCliOwnedSkillName, the rule push and uninstall already apply.

* fix(skills): pre-push review — gate the nudge on recall, keep bytecode out of the tarball, report a failed uninstall delete

A review of the whole branch against #678, #730 and the design doc, run before
pushing. What it found and what changed:

- The Stop-hook share reminder was gated on the hint switch alone; recall is off
  by default and `teamai skill get share` refuses then, so the reminder pointed
  at a command that said no. It is withheld while recall is off, the same gate
  the workflow has; the served text about when the prompt appears now matches.
- `npm pack` swept `skill-data/wiki/scripts/__pycache__` into the tarball once
  the e2e suite had run the scripts. Excluded in package.json "files", asserted
  absent in the tarball test, and the e2e run sets PYTHONDONTWRITEBYTECODE.
- The `share` description still offered to publish reusable skills, the flow its
  own body sends to `core`; the sentence is gone.
- `skill show <unknown>` before `teamai init` threw the init error as a stack
  trace; it prints the not-found line and exits 1.
- The stub pre-approved every `teamai` command from the always-loaded unit;
  narrowed to `Bash(teamai skill:*)`, which is all it asks for (#678).
- Six routing lines loaded `core --full` to reach one reference; they name the
  file under `$(teamai skill path core)/references/` instead.
- `uninstall` reported a failed delete as "holds files TeamAI did not put there;
  the packaged files were removed", both false and the error unprinted. It names
  the file and the error; a test makes the stub directory read-only.
- The stub directory archived under `<tool>/.claude-skills-teamai/teamai/` while
  the legacy trees used `<tool>/.claude-skills/<skill>/`; one layout now.
- CHANGELOG entry; dead `isRecallEnabled` import; wiki heading still naming
  `team-wiki-codebase`; JSDoc on the wrong declaration; stale byte counts; the
  usage guides gain the recall refusal and the archive location; the design doc
  records the `--json` deviation, the legacy-name classification rule, the
  uninstall symlink scope and the fail-open wording.

* fix(skills): English-only served content, one link guard for every caller, withhold share from read-only sources

The reviewer flagged Chinese in skills/ and skill-data/ a third time. Both
reach the agent as CLI output, so the stub's trigger keywords, the paired
sample invocations and the Chinese name for TGit go; the agent translates for the user. A test
fails on CJK anywhere under either root.

Pre-push review of the whole branch, and what changed:

- uninstall walked through a linked ~/.claude/skills and deleted the
  packaged files inside the member's dotfiles checkout; pull refused the same
  layout. removeOwnedFiles now owns the guard, so pull, deploy and uninstall
  apply one check: the skills root and the skill directory. A linked
  ~/.claude (stow, chezmoi) is no longer refused, since every other resource
  writes through it and refusing left those machines on the pre-stub trees.
- share was served to read-only HTTP teams, where its last step
  (teamai contribute) always fails; reportingOnly used to skip it. The
  serving gate carries a reason (recall | read-only) with its own message,
  and `skill list --json` reports it as `blockedBy`.
- The bytecode rule claimed any file under any __pycache__; it now claims
  only the .pyc of a shipped script.
- recall disable pruned the shared .agents/skills root for an uninstalled
  Codex; it has deployment's install gate now.
- The source-team guard lost the legacy names when BUILTIN_SKILL_NAMES
  narrowed, so a source removal could delete a legacy tree wholesale.
- Routing: the admin wrap-up and the stub still sent "share what I learned"
  to share without saying it needs recall, and the stub filed "share this
  with my team" (the publish-a-skill phrase) under share. recall enable is
  described as the per-machine override it is, next to the team key.
- {SKILL_DIR} definitions now say how a reference file opened on its own
  spells the directory, since serving resolves the definition too.
- skill show: packaged resolve only when init fails, aligned label, a
  served skill is "served by the CLI, not installed".
- Docs: uninstall removes the archive with ~/.teamai; zh said the whole
  directory is kept; the product overview lacked the recall gate; the design
  doc's release, tag and byte figures were stale; CHANGELOG notes the
  language change of generated documents.

* fix(skills): check every path component below the base for a link, in uninstall too

The previous commit narrowed the guard to the skills root and the skill
directory, so a link at ~/.config or ~/.config/opencode was walked through:
the prune could delete, and deploy write, inside a dotfiles checkout. The
full walk from the tool's base directory is back, and removeOwnedFiles now
requires the base, so uninstall applies it too; each skill directory in the
uninstall plan carries the base its skills root hangs off.

A member whose whole ~/.claude is a link keeps the pre-stub trees and gets
the warning naming the path, as before the previous commit. Deleting through
a link is the one thing the prune must never do.

* fix(skills): withhold the share hint on read-only sources, route legacy names through the gate

contributeHintAllowed checked recall only. The dispatcher already drops this
gitOnly handler for HTTP teams, but the gate now says so itself, so the
reminder never points at a `share` that refuses as read-only wherever it runs.

`skill show teamai-share-learnings` searched the agent directories before
the package, so a legacy tree a pull had not pruned yet was shown with its
path while the gate refused `share`. A legacy built-in name now skips the
agent search and goes to the packaged skill and its gate; ordinary names and
aliases such as `share` keep a member's own directory first.

* fix(skills): deploy and prune built-ins where the tool keeps its skills

deployBuiltinSkills joined baseDir with the configured skills path, while
team-skill sync resolves the directory through skillsDirForTool: OpenClaw's
workspace, and HERMES_HOME for Hermes. Those agents got the stub in a
directory they never read and had their legacy trees pruned from the wrong
place. Deploy, the legacy prune, recall disable and uninstall now resolve
the same directory; the link guard starts at the tool's base directory when
the skills directory sits under it, else at that directory's parent.

* fix(skills): check an external skills root for a link, keep config-load logs off stdout, drop inert allowed-tools

- A skills directory outside the tool's base (HERMES_HOME, an OpenClaw
  workspace) had the guard start at the root itself, so a linked root was
  never checked. It starts one level above now, and a linked HERMES_HOME is
  refused like a linked ~/.claude.
- The share gate loads the config, which can migrate it and report that with
  log.info on stdout: an upgrading machine got that line in `skill get`
  output and in `skill list --json`. Config loading reports on stderr for
  that call; setStderrOnly returns the previous mode so it can be restored.
- `allowed-tools` in the served skills was printed as command output and
  never processed as skill metadata, so it granted nothing. Removed, and the
  test now fails if one comes back. Only the stub's line pre-approves.

* fix(skills): walk the link guard from the scope root, so a linked COPILOT_HOME is refused

skillsGuardBase started at the tool's base directory, which for Copilot in
user scope is COPILOT_HOME, so the walk never checked whether COPILOT_HOME
itself was a link, and pull and uninstall wrote and pruned through it. The
guard now starts at the scope root (home, or the project root), where a link
at or above is ordinary, in deploy, the legacy prune and uninstall alike; a
root configured outside it still has the walk start just above that root.

* fix(skills): keep generated documents in Simplified Chinese, fail on a broken config, quote skill paths

- share and wiki had moved generated learnings and knowledge-base documents
  from always Chinese to the session language. Serving the instructions from
  the CLI does not need that, so both say "Simplified Chinese" again, in an
  English instruction; the CHANGELOG entry follows.
- skill show and skill list treated every autoDetectInit failure as "not
  initialized" and pointed at `teamai init`. requireInit now throws a tagged
  NotInitializedError; only that falls back to the packaged catalog, and a
  malformed or unreadable config propagates.
- `$(teamai skill path …)/…` is word-split in a shell command like an
  unquoted {SKILL_DIR}; all ten occurrences are double-quoted and the quoting
  test covers the form.

* fix(config): raise NotInitializedError only when the config file is missing

loadLocalConfig returns null both for a missing file and for one that fails
to parse, validate or migrate (it logs the reason). requireInit turned every
null into NotInitializedError, so skill show and skill list still fell back
to the packaged catalog and a `teamai init` hint on a broken config. Only an
absent file is NotInitializedError now; an existing one that could not be
used is an error naming its path, in requireInit and the user branch of
requireInitForScope. Covered through the real loader and the built binary.

* fix(skills): prove ownership by content, not path alone; retire the second Codex copy

- The legacy prune and uninstall removed any file at a path a release had
  packaged, so a member's edit, a skill of their own under an old name, or a
  root TeamAI never managed (toolPaths or HERMES_HOME moved) lost its files.
  A file is ours now only at a packaged path and with content a release
  shipped there: PACKAGED_SKILL_DIGESTS records the sha256 of every blob over
  all 99 tags through v0.25.0 and main before the stub, 37 versions across 21
  paths. A skill-root SKILL.md is compared by its body, since releases before
  0.17 shipped no frontmatter and the deploy of the day repaired it on disk.
  The current stub is ours by the packaged copy. Anything else stays.
- Codex reads .codex/skills and the shared .agents/skills, and the stub goes
  to the shared one when a copy lives there; the copy an earlier release left
  in the other root kept its old SKILL.md and references. It is retired by
  the same ownership rule, archived first, and named when kept.
- Tests mock the digest table with a stand-in for shipped content, and a test
  keeps the stand-in on the same paths as the real table.

* chore(skills): carry main's skill edits into the served copies after the rebase

#713 and #736 edited skills/teamai/references/*.md, which this branch moved to
skill-data/setup/references/. Two hunks did not follow the move:
- join-member.md: TGIT_TOKEN is REST-API-only and cannot clone (#713).
- setup-admin.md: the /teamai share entry publishes a reusable skill; a
  session's learnings are automatic (#736), in English as the served text is.
#739's partial config mock is restored in skip-uninstalled-tools.test.ts.

* fix(skills): deploy before pruning, block share on an unloadable config, drop hidden commands from the reference

- Legacy trees were pruned before the stub was written, so a refused or
  failed stub (a link, a read-only directory) left the agent with nothing
  to discover. They go only once the stub deployed for that agent.
- The share gate failed open on any config error. Only a machine with no
  config (NotInitializedError) is served; a config that exists but cannot be
  loaded blocks with its own reason, `blockedBy: "config"`.
- The KB template told agents to run `code-to-knowledge --update`, which
  does not exist; it names `teamai codebase --extract … --incremental`.
- The generated command reference listed hidden hook plumbing (`track`,
  `contribute-check`, `todowrite-hint`, …). It renders what `--help` lists.
- removeEmptyDirs swallowed every rmdir error, so a directory that stayed
  could be reported removed. Only "still holds something" is expected; any
  other failure is reported.

* fix(skills): whole-file ownership, stub before its references, no side effects before the link guard

Review of 327f9cd:
- SKILL.md was compared by its body, so a member who changed only its
  frontmatter lost the file. Every release from 0.16.1 (the first whose
  deploy repaired frontmatter) shipped complete frontmatter, so what is on
  disk is what was shipped: digests are whole files now (42 versions over
  100 tags and main). A link is never ours; bytecode is ours only beside a
  script proven ours by content, decided before anything is removed.
- The stub dir's retired references were pruned before SKILL.md was copied;
  a failed copy left the old skill pointing at files that were gone. The
  stub is written first.
- The Codex destination was resolved with the reconciliation that deletes a
  duplicate, before the link guard ran. It is resolved side-effect free; the
  other copy is handled under the guard by retireOtherCodexCopy, whose
  report now names a failed backup or delete as such.
- A broken project config was skipped by detection, so the share gate
  answered with the user config. findUnreadableProjectConfig reports it via
  an optional sink on detection (no caller changes), and the gate blocks.
  The Stop-hook reminder is withheld on an unloadable config too.
- init announced the stub as ready when nothing was deployed; hook-dispatch
  is hidden (hook plumbing), and the reference says it lists public commands;
  the design doc no longer says teamai-workflow/teamai-import are removed.

* fix(config): report a broken higher-priority project config even when a fallback loads

findUnreadableProjectConfig dropped a recorded error whenever detection
went on to find a later candidate: a broken partition config followed by a
valid legacy .teamai/ config returned null, and the share gate answered with
the fallback's team. It now reports the first unreadable file regardless.
An existing config file that is empty or cannot be read is reported to the
sink too, instead of returning without a word.
2026-09-23 16:15:38 +08:00
Leo Camus 667aed0f55 fix(hooks): widen track-slash regex to match SKILL_NAME_REGEX (#737)
The trackSlashHandler extracted the slash-command skill name with
`/^\/([\w-]+)/`, which does not match dots (`.`) or colons (`:`) —
both valid per `SKILL_NAME_REGEX` (`/^[a-zA-Z0-9_\-:.]{1,200}$/`).
A user invoking `/org.setup` or `/ns:deploy` had the name silently
truncated to `org` or `ns`, so the wrong (or nonexistent) skill was
tracked.

The CLI counterpart `trackSlashCommand` (usage-tracker.ts) already
uses the full `[a-zA-Z0-9_\-:.]+` set. Align the hook handler.
2026-09-23 15:22:58 +08:00
Saul Moro ff47714902 fix(doctor): probe the Copilot hooks file where inject writes it (#732) (#733)
In a non-self project scope, resolveDoctorContext forces the hook paths to
the hook scope ('user', per resolveHookScope) so settings-based hooks are
probed where reconcileHooksToAllTools writes them. The standalone Copilot
hooks file is written by reconcileTeamHooksForConfig at the config's own
scope instead, so the doctor ended up joining the userScope relative path
(hooks/teamai.json) onto <projectRoot> and reported Copilot missing right
after a successful `hooks inject`.

buildHookChecks now takes both maps and picks per hook kind: a standalone
`hooks` file from the config-scoped paths, `settings` from the hook-scoped
ones. Same rule `hooks list` already applies.

Hypothesis confirmed: scope mismatch between the two path maps, introduced
when #695 moved the doctor's hook paths to the hook scope for Qoder CN.
2026-09-23 15:22:00 +08:00
pablo bc6943ea30 fix(members): read the default-branch roster as an inherited root after the reports switch (#741)
The orphan-branch switch (#489) made members list and projects members read
only the teamai-reports worktree, so a team whose roster still lives on the
default branch saw "No team members registered" right after upgrading (#735).

The default-branch clone now stays a read-only inherited member root, the way
learnings' already is (#485): listing unions both roots (the reports-branch
copy wins when the same file exists on both), nothing is copied or deleted, and
read-only commands still never publish the reports branch. Member registration
merges against the inherited copy too, so a re-init keeps the original
registeredAt/projects and converges the data onto the branch.

Fixes #735
2026-09-23 15:17:44 +08:00
dvd233 48b3dcb953 feat(hooks): add DeepSeek Harness hook bridge (#689)
Fixes #623
2026-09-23 15:02:34 +08:00
Saul MoroandSaul Moro d80d5a8678 fix(push): namespace new rules and agents from --role/--project (#649) (#698)
* fix(push): namespace new rules and agents from --role/--project (#649)

`--role`/`--project` only ever placed new skills, so a rule pushed with
`--project front-app` landed at `rules/<name>.md` and a new agent at
`agents/<name>.yaml` — both of which `pull` ships to every member. The
flag also collapsed into the project's `skills` namespace, which is the
wrong directory for a rule: a rule is namespaced on the `knowledge` axis,
and the manifest allows the two to differ.

Each pushable type now resolves from its own axis (skills → `skills`,
rules → `knowledge`, agents → `agents`), and the destination is printed
rather than chosen silently. Where the named project declares no
namespace for a type being pushed, the command fails and names it
instead of writing to the shared root.

Only new resources already at the shared root are placed; anything the
scanner namespaced keeps its path (#654), and an open PR's recorded
destination still wins so a force-push never moves a resource.

Also resolves a root-level local rule against its namespaced team copy,
so a rule that was placed on an earlier push is not re-pushed to the
shared root once it merges.

* fix(push): record rule placement in state instead of matching by basename

RulesHandler.scanLocalForPush matched a root-level local rule against the
sole active rules/<ns>/<name>.md by basename. A namespaced team rule is
pulled into a namespaced local directory, so another member's unrelated
root rule with the same name would have been read as a modification of the
team rule and overwritten it under --all.

push now records where it placed each root-level rule (state.placedRules,
name -> team path). The scanner redirects a root-level local rule only when
that record exists and its team file is still present; otherwise the rule
is new. The ambiguity warning goes with the basename index.

Review: https://github.com/Tencent/teamai-cli/pull/698#issuecomment-5770251793

* fix(push): stop on unreadable roles manifest, sync placed rules, resolve in dry-run

Three review findings on top of the #649 placement fix.

A roles manifest that exists but cannot answer — unparseable, or missing the
configured role — no longer falls back to an empty namespace list, which sent a
new rule or agent to the shared root and therefore to the whole team. Only an
absent manifest keeps the pre-manifest fallback, so `loadRolesManifest` now
throws a tagged `RolesManifestNotFoundError` to tell the two apart.

`syncTeamUpdatesToLocal` follows the same `placedRules` record the scanner does,
so a root-authored rule placed under `rules/<ns>/` takes part in the three-way
sync. Without it a teammate's newer version was never synced down and the stale
root copy was pushed over it.

`--dry-run` now runs the recorded-destination and placement steps before it
exits, so it reports where every new resource goes and fails on the same
unresolvable project axis the real command refuses.

* test(push): cover #649 placement with the real CLI across agents and providers

Drives the built dist/index.js against real git remotes, a fake `gh` and a fake
GitLab API, and asserts on the branch content that reached the remote.

The four agents × three providers cover the placement itself; the remaining
cases cover what review round 2 raised — the pre-push sync following
placedRules, an unreadable roles manifest stopping the push, and --dry-run
resolving the same destinations. Reverting any of those three fixes turns
exactly its case red and leaves the rest green.

* fix(push): keep a placed resource maintainable and removable by its author

Three review findings on the placement this PR added.

A placement record only ever meant "push put this here", and the two sides that
read one disagreed about how much it was worth. `placedResourcePath` is now the
single resolver: it validates the record (inside the resource root, namespaced,
no traversal, named after the resource) and both the push scanner and the
pre-push sync go through it, so they cannot drift apart again. The record also
takes precedence over a shared-root file that appears later with the same
basename — mapping the author's copy onto somebody else's rule would push their
content over it.

Agents gained the analogue, `placedAgents`. `AgentsHandler.scanLocalForPush`
only accepts a team source whose namespace is ACTIVE here, so an agent
published with --role/--project into a namespace this directory never
activated was skipped as "no active source" on the author's very next edit:
they could create the agent and then never maintain it.

`teamai remove rules <name>` resolves the same record through the new
`publishedNameFor` hook. The author's copy stays at the rules root, so the name
they type is the bare one, and remove answered "not found" about a rule it had
recorded publishing. It now reports which name it resolved to, deletes the
namespaced team file, and takes the author's root copy with it — left behind,
that copy re-publishes the rule on the next push.

* fix(push): narrow what a placement record grants, and when it is written

Four review findings, each about the record rather than the placement.

`remove` consulted it only after a bare-name match failed, but the LOCAL scan
contributes the bare name whenever the author's own copy has edits — so
`remove rules my-rule` deleted that copy, reported success, and left
`rules/<ns>/my-rule.md` published. The record is now resolved first.

A record is written only for a resource push actually placed: `new`, and
namespaced by this run. Recording a `modified` agent meant a namespace that
happened to be active at edit time became standing permission to keep editing
that agent long after the role or project granting it was dropped.

Records are persisted per group, right after that group reaches the remote,
instead of after every group completes. A failing later group returned early
and took the earlier group's mapping with it, so a resource that WAS pushed
came back misclassified once its PR merged.

And the roles manifest is held to the same rule as `--role` and the projects
manifest: a namespace is one path segment. `foo/bar` wrote an agent below the
depth pull looks at, and read back as namespace `foo` for a rule.

* fix(push): let --role/--project decide which team agent a local edit belongs to

`AgentsHandler.scanLocalForPush` picked the team file to edit by activity
alone, and it runs before the destination is resolved. So `agents/other-ns/vr.yaml`
— an agent this directory never activates — made `push --project front-app`
report "no active source" and push nothing, even though the same stem is
allowed to exist in several namespaces and the flag had named a different one.
The scan now takes the requested namespace, through a new optional
`ScanForPushOptions`; with one named, sources in other namespaces are other
agents, and an absent one means this agent is new there. A shared-root copy
still blocks, and now says why: both would be active at once, which is the
collision pull reports.

`RulesHandler.removeItem` swept the bare basename unconditionally, so
`remove rules fe/foo` deleted an unrelated personal .claude/rules/foo.md. The
bare copy is only ours to delete when this machine's placement record says the
two are the same rule.

`teamai push --help` said both flags target skills.

* fix(push): never place a new resource onto one that is already there

Placement rewrote a new root-level resource to the resolved namespace without
looking at what was at that path. An unrelated local `foo.md` — which the
scanner rightly calls new, since no record maps it anywhere — landed on
`rules/<ns>/foo.md` and replaced somebody else's rule, silently, in a run they
never reviewed. Push now stops and names the file. The same guard covers the
`--role`/`--project` skills override, for new skills only: a modified one is
meant to land on its own directory.

`loadRolesManifest` read through `readFileSafe`, which answers null for every
failure, so a manifest that exists but cannot be read arrived looking exactly
like a missing one — and a missing one is the pre-manifest layout, which sends
new rules and agents to the shared root. The two are told apart now.

`RulesHandler.removeItem` tombstoned only the name it was given. Removing
through a placement record means the author's source is named `<name>` while
the published file is `<ns>/<name>`, and the local sweep skips excluded tools,
so a root copy could outlive the removal there and come back on the next push.
Both names are tombstoned when the record vouches for the bare one.

* fix(push): resolve a placed agent on removal, and collide on either extension

`remove` asked every handler for the published name, but only rules answered.
So `teamai remove agents vr` matched the bare stem and deleted every `vr` in
every namespace — other people's agents included — while the namespaced
`placedAgents` record, keyed by a path the removal never named, survived.
`AgentsHandler.publishedNameFor` resolves it now; the existing sweep already
narrows a `<ns>/<stem>` to one file, since the root directory is one of the
directories it probes. The bare stem is tombstoned alongside the published one
and swept from the tool directories, the same way rules are.

The placement collision check tested the proposed path alone. `pull` reads a
legacy `<stem>.md` as the same agent as `<stem>.yaml`, so a new `.md` landing
beside an existing `.yaml` passed the check and left two copies answering to
one name. Agents are now checked under both canonical extensions.

* fix(remove): make the placed-agent resolution actually reach the command

Round 7 added `AgentsHandler.publishedNameFor` but `remove` only used its
answer when `allNames` also carried that spelling — and `scanTeamForPull`
reports an agent by its bare stem, never `<ns>/<stem>`. So the resolution was
inert on the real command path: `teamai remove agents vr` fell back to the bare
stem and deleted every `vr` in every namespace, exactly as before. The cross
check is gone; `publishedNameFor` has already proved the file is in the team
repo, which is stronger evidence than membership in a list the scans spell
differently per type.

The bare-stem tombstone that round went with it. Agents deploy FLATTENED, so
both the push scan and the post-pull cleanup read a bare tombstone globally:
removing `fe/vr` suppressed and deleted `be/vr` the moment that namespace
became active. Only the published name is tombstoned now. The author's own
flattened copy is still swept, but only where this machine's record says the
file just removed is where push put it — without that, the copy on disk may be
another namespace's deployment.

Covered end to end this time: the new case drives `teamai remove agents vr`
through the built CLI, which is the join the round-7 unit tests skipped.

* fix(push): raise a project agents-axis failure the scan would otherwise swallow

`--project <id>` resolved the agents destination before scanning and dropped
the failure on the floor. A project with no agents namespace then looked
identical to a run with no flag at all: the scan skipped the agent as "no
active source", the item never reached placement, and the command exited 0 with
"No new or modified resources" — on a flag it could not honour. The error is
carried forward and raised as soon as the scan contains an agent. It cannot
wait for the selection the way the skills axis does, because the item that
would prove the axis is needed is exactly the one the scan removes.

`placedResourcePath` matched the recorded filename by prefix, so a record
pointing at `rules/<ns>/foo.backup.md` was trusted whenever that file existed,
and scanning, the pre-push sync and removal would all follow it onto somebody
else's file. The filename must now be exactly the resource's own.

* fix(agents): reach the canonical source, and hold the record to what it proves

Four review findings, all on the agent side of placement.

The single-repo canonical source in `.teamai/agents/` is picked up directly,
never reverse-parsed, and that branch ignored the placement record: a root
`vr.yaml` placed at `agents/fe/vr.yaml` read as new on the next push, and the
collision guard then refused the very agent this machine published. Removal
missed the same directory, so the agent republished itself on the next push —
which a bare-stem tombstone cannot prevent without suppressing that stem in
every other namespace, since agents deploy flattened.

The project agents-axis error now counts only agents that actually need a
destination. A modified agent already in a namespace is written in place, so an
empty agents axis is none of its business; blocking it contradicted the rule
that only new shared-root resources are placed.

And the record is no longer taken as licence to overwrite. It admits a namespace
this directory never activates, which also means `pull` never refreshed a copy
and the pre-push sync does not cover agents — so if the canonical file moved on
since the last pull, push now says so and asks for a pull instead of writing a
stale rendering over whoever changed it.

* fix(agents): deliver recorded agents on pull, and let a named destination win

The staleness guard added last round was defeated by the pull it recommended:
`pull` advances lastPullRev without deploying an inactive namespace, so the
next push saw an unchanged canonical and wrote the stale rendering anyway. It
also never fired right after the first PR merged, when the file did not exist
at lastPullRev.

The guard is gone, and the cause with it. `pull` now delivers an agent whose
placement record names it, so the local copy tracks the team file and the
ordinary comparison is valid — the inactive case stops being special instead of
needing its own machinery. A stem an ACTIVE namespace already claims is left
alone, since agents deploy flattened and the active one is what is deployed
here; the scan follows the same order, treating the record as a fallback rather
than an extra candidate.

Pending-PR reuse matched on type and name alone, so an open PR for a different
resource of the same name captured a push that named another namespace and
force-pushed into that review. Neither silent answer is safe, so the flag the
user typed decides, the open PR is left untouched, and the collision is
reported. This supersedes the original #331/#654 rule that a PR's destination
always won; that rule still holds whenever no destination is named.

* fix(pull): stop revoking the agent pull had just delivered

Self-review of the branch, before the next review round.

Round 11 taught delivery about placement records but not revocation, and both
run in the same pull: `filterAgentsByNamespaces` wrote the agent and
`cleanupInactiveNamespaces` deleted it again, byte-equal to the render so the
data-safety gate passed it straight through. The record-based delivery was
inert and the file churned on every pull. Both halves now resolve through one
exported `selectAgentsForDirectory`, so they cannot disagree — the same
treatment `placedResourcePath` already gives the push scanner and the pre-push
sync.

Two smaller ones from the same pass. `--dry-run` grouped against the unfiltered
pending list, so it reported a destination the real push no longer uses, which
breaks the property that a dry run matches the run it describes. And the
partial-selection warning counted entries the run had already declined to
reuse, contradicting the warning given for them; both now share the filtered
list, which is computed once there is actually something to push.

* fix(pull): stop the stale sweep from deleting the author's own placed rule

`pullAllRules` sweeps a local rule whose name is absent from the desired set.
A rule published into a namespace keeps the author's copy at the rules ROOT
under its bare name, while the desired set holds `<ns>/<name>` — or nothing at
all when that namespace is not active here — so the sweep deleted their own
file, local edits included. The placement record marks it as theirs, and only
while the team file it points at still exists.

Pending-PR conflict detection trusted `PendingPushItem.namespace`, but the
agent scan records a namespaced destination without setting that field, so
those entries slipped past the check and a push naming another namespace could
force-push into the PR under review. The namespace is derived from the recorded
path when the field is absent, and the agent scan now sets it too — the field
was the only thing telling `pendingNamespaceFor` where to put the resource.

* fix(push): defer the agents-axis failure to selection, reload projects.yaml after the pull, and stop a named namespace reusing a shared-root PR

A project with no agents namespace failed before the listing whenever any new
agent was present locally, so a rules-only push under `--project` was blocked
on an agent the user was never given the chance to deselect. Only an agent the
scan itself skipped (`needsDestination`) fails early now — that one never
reaches the listing, so deferring its error means never raising it. A new agent
is listed, and step 4 raises the same error if it stays selected.

`manifest/projects.yaml` was read in `push`, before `pushCore` pulled the team
clone, so a project whose namespaces changed on the remote placed this run's
new rules and agents by the previous pull's mapping. It is read inside
`pushCore` now, after the pull, and in self mode from the fresh worktree.

Pending-PR conflict detection treated a recorded path with no namespace as
non-conflicting, so an explicit `--role`/`--project` reused a shared-root PR's
branch and rebuilt it with the namespaced path, moving a review the user did
not name from "everyone" to one namespace. The shared root is a destination
like any other: it conflicts with any namespace the flag names.

`pull` delivered a rule this machine placed at `<tool>/rules/<ns>/<name>`
beside the author's copy at the rules root, so a tool that loads rules
recursively applied the same rule twice, disagreeing as soon as the team file
moved on. The placement record names the root copy as this rule's local file,
so delivery updates it and removes the namespaced duplicate an earlier pull
wrote. A namespace another member placed is untouched.

* fix(push): stop on a stale clone under --project, retire a renamed canonical agent's recorded file, and prune placement records

A failed refresh of the team clone was only warned about, after which
`--project` resolved every destination from the previous pull's
`manifest/projects.yaml`. A namespace the remote had changed sent this run's
new rules and agents to the members of the old one. `--project` stops now, and
so does a new resource that would resolve from `manifest/roles.yaml`; a team
with no roles manifest resolves from nothing that can go stale and keeps its
behaviour. `--role` names the namespace itself and is unaffected.

In self mode the placement record redirected a root canonical agent to the
recorded file, extension included. `pushItem` writes by the source's
extension, so an author who rewrote `vr.md` as `vr.yaml` had the `.yaml`
written and the `.md` staged: the change never reached the branch and the
`.md` stayed. The destination now keeps the record's directory and the
source's extension, `pushItem` deletes the file it retires, `push` stages that
deletion and moves the record to the new path.

Placement records were written when a branch reached the remote and never
removed. Once the PR was closed unmerged, or the file deleted upstream, the
record pointed at nothing — until another member created the same path, at
which point it came true again and their unrelated resource read as this
author's. `push` and `pull` now drop a record whose target is neither on the
default branch nor awaiting review on a branch origin still has, before
anything reads the records. A record is kept when origin cannot be asked.

* fix(push): record a placement only once it has landed, withdraw it when a shared-root file takes the name, and drop every stale record

Placement records were written when the branch reached the remote, so a PR
closed without merging left one behind for as long as its branch stayed —
and no provider here can say whether a PR is open. The record now travels on
the pending PR entry (`PendingPushItem.placed`, with the blob push wrote) and
becomes a `placedRules`/`placedAgents` record only when that blob is in the
default branch's history for the path: the PR merged, however the platform
merged it. A path that merely exists is not enough, since another member may
have created it after the PR was closed. `push`, `pull` and `remove` settle
this before reading the records; the stale sweep spares a root copy whose
placement is still awaiting review.

`localNameFor` redirected a recorded rule onto the bare root path without
asking whether a shared-root rule of the same name was being delivered too,
so both landed on the one file in loop order, and the next push could follow
the record and carry the shared rule over the namespaced one. Delivery keeps
the namespaced path when the shared root holds that name, and the reconcile
pass withdraws the record with a warning: the root copy follows the shared
rule from then on.

Dropping stale records destructured from the ORIGINAL map each iteration, so
a later deletion put back what an earlier one had removed and only the last
stale record went. The kept entries are rebuilt in one pass.

* fix(push): consume a pending placement once it is recorded

Recording a landed placement left its `placed` mark on the pending entry. Had
the team then deleted the file — which drops the record — and another member
recreated the path, the next reconcile recorded it again: the path existed and
the blob push had written was still in the default branch's history, so both
checks passed, and the unrelated replacement read as this author's resource.
The mark and the blob are cleared when the record is written, so a placement
is recorded exactly once.

* fix(remove): keep placement records until the removal lands, and retire flattened copies

`remove` dropped the placement record as soon as its branch was pushed, so a
retry during review resolved `vr` to the bare stem and removed that agent from
every namespace. Records are now left to the reconcile pass, which drops one
once its file is gone from the default branch, or was deleted since the last
check (`placementsCheckedAt`) even if another member has recreated the path.

A namespaced removal tombstones only `<ns>/<name>`, which never matched the
flattened `<agents>/<name>` copy members hold, so that copy survived pull and
the next push republished it. `AgentsHandler.removedStems` reads the tombstone
as the flattened stem while no namespace still has that agent, for both the
pull cleanup and the push scan. Rules no longer write a bare tombstone, which
swept and suppressed other members' unrelated rules of the same name.

Agent and rule push scans skip tools the member excluded: remove leaves those
copies behind by design, so reading them republished the removed resource.

Landing is proven only by history after the full commit the push branch was
built on, and a placement whose path was deleted after it landed is spent
unrecorded. Single-repo pull no longer reconciles against the member's own
checkout; push and remove still do, in a fresh origin/<default> worktree.

* fix(pull): reconcile single-repo records through origin/<default>, and retire flattened copies per directory

Single-repo pull skipped the reconcile pass, so a merged placement stayed
unrecorded until the next push or remove. It now reads the default branch as
the ref origin/<default> (existence via `<ref>:./<path>`, history up to that
ref) instead of the member's own checkout, and changes nothing when the ref
cannot be resolved.

`placementsCheckedAt` survived its last record, and a record made in the same
run was checked against it, so re-placing a resource at a path deleted earlier
dropped the new record at once. The checkpoint is cleared with the last record,
and records made in this run are not held to it.

`removedStems` retired the flattened stem only when no namespace had it at all,
so an fe member kept a removed fe/vr while an unrelated be/vr existed. It now
asks what this directory is meant to hold, through the same selection pull
delivers with.

* fix(remove): stop on a stale clone, and judge PR conflicts by the destination the flag gives

`remove` ignored a failed refresh and reconciled the stale clone as if it were
the default branch. A placement merged since the last pull was then not
recorded, and the bare name fell back to the stem, removing that agent from
every namespace. `remove` now stops with exit 1 and removes nothing.

Under --role/--project, a pending PR counted as conflicting whenever its
recorded namespace differed from the flag's, although only skills and new
shared-root rules and agents are moved by it. A modified rule already in a
namespace kept its path yet went to a second PR on the same file. Only an
item the flag actually moves can conflict with it now.

* fix(push): prefer an agent's delivered source over the flag, baseline new placements, validate the record's namespace

With --role/--project the requested namespace always chose the team file a
local agent was compared with, so an untouched copy delivered from an active
namespace read as an edit of the requested namespace's agent and overwrote
it. Candidates now follow delivery: an active source (the shared root
included), then this machine's record, and only then the requested namespace.

A placement that landed after the last pull had no `lastPullRev` version, so
the pre-push sync skipped it and a teammate's edit before the author's next
pull was pushed over. Rules take the version the file was added with as their
base; a recorded agent, which has no pre-push sync, is held with a pull-first
message when the team file moved past that baseline.

`placedResourcePath` now requires a safe namespace segment: a backslash in it
is a separator on Windows and walked out of the resource root.

* fix(push): close the round-21 findings on placement, removal and agent sources

- A pending namespaced placement whose name a shared-root file now takes is
  left out of the push with a warning, instead of the open PR being rebuilt
  with the author's copy over the shared file.
- `remove` stops when the placement records cannot be reconciled and saved:
  a missing record sends the bare name to the stem, which spans namespaces.
- An agent skipped for want of a --project agents namespace no longer blocks
  the rest of the push; the error stands only when nothing else is left.
- A placement is marked only with a blob that can prove it landed, and one
  without is spent rather than recorded because its path exists.
- The single-repo `.teamai/rules` scan source is not a tool, so an
  `enabledAgents` list no longer hides it.
- A namespaced agent tombstone retires the flattened stem only where that
  agent could have been delivered; reconcile keeps the author's dropped
  record (`retiredPlacedAgents`) so their own copy still counts.
- Two active same-name agents stay ambiguous under a flag, and a flag naming
  a namespace that already holds the agent is a collision, as for rules.
- In single-repo mode a root copy equal to an older version of its placed
  file is held as stale rather than pushed over a teammate's edit.
- The recorded-agent hold runs only for a changed copy and says to set the
  edit aside first; a pending placement is routed to its PR, not skipped;
  a flag that does not move a shared-root edit says so; several candidate
  namespaces without a terminal fail with a --role hint; a placement that
  landed with other content is reported once.

* fix(push): stop on unsaved records and stale placements, name namespaced agents on remove

- `push` stops, pushing nothing, when the reconciled placement records
  cannot be saved: the sync and the scan read them back from disk, and a
  record that could not be withdrawn still redirects the author's copy.
- On a stale clone every unflagged placement stops, not only one resolved
  from an existing roles manifest: the manifest's absence and the skills
  namespaces detected from the tree are clone state too.
- `remove agents <ns>/<name>` names one namespaced agent; a bare name only
  one namespace has resolves to it, and a bare name found in several places
  is refused rather than removed from all of them.
- A record dropped because its file was deleted and recreated is retired
  like one whose file is simply gone, so the author's flattened copy of the
  removed agent is still recognised.

---------

Co-authored-by: Saul Moro <smoro@ai-lab.knowmadmood.com>
2026-09-23 14:58:21 +08:00
Hill PatelandClaude Sonnet 5 94a0d428c1 fix(env): only stick to a candidate the active profile actually reaches (#682) (#715)
* fix(env): only stick to a candidate the active profile actually reaches (review)

02c93fe's resolveActiveShellProfile scanned every SHELL_PROFILE_CANDIDATE_NAMES
entry for a matching block and returned the first hit, in a fixed order
(.zshrc, .bashrc, .bash_profile, .bash_login, .profile). That's broader
than the Git-for-Windows-forwarding case it was written for: a stale
block a pre-#682 install left in .bashrc would outrank a correctly
order-picked .profile that hasn't been written to yet, since .bashrc
sorts earlier in the candidate list — silently reintroducing #682 for
exactly the installs upgrading through this fix, with doctor unable to
catch it because the stale block is well-formed where it sits.

Reworked to start from detectShellProfile's order-based pick (the file
the current environment actually reads) and only diverge from it when
that pick's own content references another candidate by a home-relative
path (~/.bashrc, $HOME/.bashrc) — the shape Git for Windows' generated
forwarding file actually takes. A block sitting in a candidate the pick
never reaches is no longer preferred over the pick, regardless of what
it contains.

Also caught and fixed a case of exactly the failure mode this PR is
about: the first cut of the forwarding check was a bare substring match
on the candidate's filename, and my own test's plain-English comment
("...unrelated to .bashrc") satisfied it. Tightened to require the
home-relative reference form a real sourcing line uses.

Verified both scenarios end-to-end on a real Windows host:
- The exact bot-reported upgrade case (stale .bashrc block, empty
  .profile, no forwarding between them): pull now writes into .profile
  and correctly flags .bashrc as stale; doctor reports delivery healthy.
- The Git-for-Windows forwarding case from the prior round: still
  sticks to .bashrc through the generated .bash_profile, no duplicate,
  no stale-block warning.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(env): match real source commands, resolve reachability transitively (review)

Two P1s from the bot's review of #715:

- resolveActiveShellProfile's reachability check was a bare substring
  search on the active pick's content. A comment mentioning a filename
  (never executed) or a longer file sharing the same prefix
  (~/.bashrc.local) would both satisfy it, letting a stale block win
  the same way #682 did. Replaced with referencesCandidate(): strips
  full-line comments, splits each remaining line into statements on
  &&/||/;, and only counts a statement whose first word is literally
  `.` or `source` and whose second word is an anchored home-relative
  reference to exactly that candidate.

- The check only followed one hop: .bash_profile sourcing .profile
  sourcing .bashrc (the common Debian .profile pattern, sourcing
  .bashrc for interactive shells) would miss a block two hops away and
  inject a duplicate. Reworked into a loop that walks the chain of
  files the pick actually sources, with a visited set for cycle
  protection, stopping at the first one that carries the block.

Also fixed the P2: EnvHandler.detectShellProfile's doc comment still
claimed it "stays on whichever candidate already carries this scope's
block" unconditionally, which stopped being true once reachability was
required.

Verified end-to-end on a real Windows host:
- The new two-hop chain (.bash_profile -> .profile -> .bashrc, block
  in .bashrc): resolves to .bashrc, no duplicate, doctor fully clean.
- Re-ran the Git-for-Windows one-hop scenario and the #682 upgrade
  scenario from the prior round — both still correct, no regression.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(env): search every referenced candidate, respect || conditionality (review)

Two more P1s from the bot's round-10 review of 75f3eac:

- The traversal committed to the first referenced candidate in
  SHELL_PROFILE_CANDIDATE_NAMES's fixed priority order and gave up if
  that branch was a dead end, instead of trying every candidate the
  current file actually references. Git for Windows' own generated
  .bash_profile sources both .bashrc and .profile in one file — if the
  real block sits in .profile but .bashrc (sorting earlier) has none,
  the walk stopped at .bashrc without ever trying .profile. Reworked
  into a breadth-first search over the whole reference graph.

- Splitting statements on `||` treated its right side as unconditionally
  reached, but `||`'s right side only runs if the left side fails,
  which isn't something this code can establish. `source ~/.profile ||
  source ~/.bashrc` would mark .bashrc reachable even when .profile
  succeeds. Statements no longer split on `||`; a `source`/`.` sitting
  only after it is folded into its left side's statement and never
  recognized as its own reference, so it's never preferred over a
  block the left side already reaches. Conservative by construction:
  worst case is falling back to the order-based pick (the pre-#693-fix
  behavior), never a false "reachable".

Verified the exact branching scenario end-to-end on a real Windows
host: .bash_profile with the literal Git-for-Windows-generated content
(sources both .bashrc and .profile), .bashrc empty, real block in
.profile — resolves to .profile, no duplicate, doctor fully clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(env): only trust verifiable && and || conditions, ignore if bodies (review)

Two more P1s from the bot's round-11 review of aaa1142, both about
referencesCandidate() trusting shell control flow it can't actually
evaluate:

- Any `&&` was treated as making its right side reachable, without
  checking what the left side's condition even was. A guard like
  `[ "$TERM_PROGRAM" = vscode ] && source ~/.bashrc` would mark
  .bashrc reachable unconditionally, even though it only runs inside
  VS Code. Also flagged: a source sitting inside a multiline `if`
  body looks, line by line, identical to a top-level one.

- Folding `||`'s right side into its left statement (the round-9 fix)
  went too conservative the other way: `source ~/.profile ||
  source ~/.bashrc` DOES guarantee .bashrc runs when ~/.profile
  doesn't exist, and the resolver was never even trying it.

Rather than growing another ad hoc regex tweak, rewrote
referencesCandidate() around what it can actually verify without a
real shell parser:

  - Unconditional: a bare `. REF` / `source REF` — but nothing inside
    an `if` block counts, conditional or not. An `if`'s condition is
    opaque to a line scanner; trusting some conditions and not others
    would just be guessing.
  - Existence-gated `&&`: only the self-referential idiom
    `test -f REF && . REF` / `[ -f REF ] && . REF`, where the tested
    path and the sourced path are the same candidate — the one `&&`
    condition this code can independently verify, by visiting that
    candidate itself later in the search.
  - `||` fallback: the left side always counts (always attempted);
    the right side counts only when the left side's own target file
    does not exist on disk — the one case an `||` fallback is
    actually guaranteed to run.

Anything this can't resolve either way is never trusted: the search
just doesn't queue that candidate, and the caller falls back to the
order-based pick — at worst a harmless duplicate block (the
pre-#693-fix behavior), never a false "reachable" that would
reintroduce #682.

Verified end-to-end on a real Windows host: re-ran the core
Git-for-Windows two-pull scenario from #693 (self-referential &&,
still recognized) with no regression. Added 3 unit tests for the new
boundaries: || recognized when the left target is missing, a
non-existence && condition rejected, and a source nested inside an
if block rejected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(env): document transitive chaining and the verifiable-reference boundary (review)

Round-11 P2: both docs described only a single directly-referenced
candidate, but the resolver has followed transitive chains since
75f3eac and now only trusts specific verifiable && / || forms (60a2da0).
Describes the Debian .profile -> .bashrc two-hop case alongside the
Git-for-Windows one, and names the three reference shapes recognized
(bare source, self-referential existence-gated &&, existence-checked
|| fallback) and that if-bodies are never trusted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(env): validate reference quoting, credit &&'s left side, generalize block-skip (review)

Round 12 found three more real gaps in referencesCandidate(), plus one
I agree isn't worth chasing further (see PR reply):

- The quote check accepted `source "~/.bashrc"` and `source
  '$HOME/.bashrc'` as valid references, but a shell never tilde-expands
  inside any quotes and never variable-expands inside single quotes —
  both source a literal, near-certainly nonexistent path. Tightened to
  the three forms that actually expand: bare `~/name`, and `$HOME/name`
  either bare or double-quoted.

- `&&`'s left side is always attempted, the same as `||`'s — `source
  ~/.bashrc && echo ready` does reach .bashrc regardless of the
  trailing command, but the old "whole statement must be exactly `.
  REF`" check missed it. The leftmost command before the first `&&` (or
  no `&&` at all) is now checked the same way `||`'s left side already
  was.

- Only `if`/`fi` was tracked, so a source inside an uncalled function,
  a non-selected `case` arm, or a loop body — none of them any more
  guaranteed to run than an `if` body — was wrongly treated as
  top-level. Generalized the "don't trust it" depth counter to cover
  for/while/until, case/esac, and function/brace groups too, sharing
  one counter since we only need to know whether we're inside *any* of
  them, not which one.

Declined to extend if-body trust to cover the standard nested Debian
`.profile` template (`if [ -n "$BASH_VERSION" ]; then if [ -f
"$HOME/.bashrc" ]; then . "$HOME/.bashrc"; fi; fi`) — doing so would
mean trusting the outer `$BASH_VERSION` check, which is exactly the
class of unverifiable shell condition this design has refused since
round 11. Fixed the docs instead: they previously (incorrectly)
claimed this exact template was recognized; now they say plainly that
nested conditionals of any kind fall back to the order-based pick.

Verified end-to-end on a real Windows host: re-ran the Git-for-Windows
two-pull scenario unaffected. Added 6 unit tests for the new
boundaries (invalid vs. valid quoting, &&'s left side, function/case/
loop bodies). 41/41 in shell-profile.test.ts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(env): handle comments, line continuations, heredocs, and chained && in profile scanning (review)

Round 13 review found five genuine structural gaps in referencesCandidate,
all fixed by moving open/close-block detection to per-statement (post
`;`-split) instead of per-line, and adding a logicalLines() preprocessing
pass:

- Backslash-continued lines were scanned independently, losing the
  conditional context of the line they continue (`cond && \` followed by
  `source X` on the next line looked unconditional).
- A one-line `if ...; then ...; fi` only incremented depth (matched via
  the whole-line "opens" check) and never saw its own `fi` close it,
  permanently disabling recognition of every later unconditional source
  in the file. Two-line function definitions (`fn()` then `{` on its own
  line) double-incremented for the same reason.
- The existence-gated `&&` guard was fully anchored, so a guarded source
  followed by further `&&`-chained commands (`[ -f X ] && . X && export Y`)
  didn't match even though the guard still holds.
- Comment stripping only skipped whole-comment lines; a comment following
  a semicolon on the same line was still split into a "real" statement.
- Heredoc bodies were scanned as literal executable lines.

Declined the sixth (recognizing the Debian/Ubuntu nested
`if [ -n "$BASH_VERSION" ]; then if [ -f ... ]; then . ...; fi; fi`
template) for the same reason given in review round 12: the outer
condition is unverifiable without a real shell, and this resolver's
explicit, repeatedly-restated design boundary is to never trust an
unverifiable condition — falling back to the order-based pick (a
harmless duplicate block) is the intended safe behavior there, not a bug.

Verified with 6 new unit tests (47/47 passing) plus a standalone real-fs
script driving the actual resolveActiveShellProfile against a scratch
HOME for all seven round-13 scenarios (all pass). Full suite unchanged
at the pre-existing 30-failed-file/66-failed-test Windows-host baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(env): quote-aware statement splitting, subshells, dead code after return/exit, N-way || (review)

Round 14 review found six more genuine gaps in referencesCandidate, all
fixed:

- The `;`/`&&`/`||` splits were plain `String.split`, so a separator
  character inside a quoted argument (e.g. `printf '%s' 'x; source
  ~/.bashrc; y'`) was treated as a real statement boundary, inventing an
  executed source out of string data. Added splitTopLevel(), a small
  quote-aware splitter (tracks single/double-quote spans, skips
  separators inside them) used everywhere a naive split was previously
  used.
- `(...)` subshells weren't tracked as an unverified-block construct —
  a source inside one always runs, but its exports never reach the
  caller, so it must not count as reaching a candidate any more than an
  `if` body does. Added to opensUnverifiedBlock/closesUnverifiedBlock
  alongside the existing if/for/while/until/case/function handling.
- An unconditional, top-level `return`/`exit` ends the file's control
  flow right there; anything textually after it was still being scanned
  as if reachable. Added a `halted` flag set on a bare return/exit
  statement, gating everything after it for the rest of the scan.
- `sourceOf` required the source's argument to be the entire statement,
  so `. "$HOME/.bashrc" 2>/dev/null` and `source ~/.bashrc extra_arg`
  (both valid, both really sourcing the target) went unrecognized.
  Relaxed to capture just the first argument and allow anything after
  it.
- The `||` fallback only handled exactly two operands — a three-way
  chain like `source ~/.profile || source ~/.bash_login || source
  ~/.bashrc` wasn't recognized at all, not even the always-attempted
  left side. Generalized to N operands: each one counts only when every
  operand before it is a recognized source whose target is verifiably
  missing from disk.
- Multiple heredocs opened by one command (`cat <<A <<B`) only tracked
  one terminator, so the second heredoc's body was scanned as real
  statements once the first terminator was seen. heredocEnd is now a
  queue of terminators consumed in order.

Declined the seventh finding again (the Debian/Ubuntu nested `if
[ -n "$BASH_VERSION" ]` template) for the same reason given in rounds 12
and 13: the outer condition is unverifiable without a real shell, and
this resolver's explicit design boundary is to never trust one — the
order-based-pick fallback (a harmless duplicate block) is the intended
safe outcome there, not a bug.

Verified with 8 new unit tests (55/55 passing) plus a standalone real-fs
script driving the actual built resolveActiveShellProfile for all nine
round-14 scenarios (all pass, including confirming the Debian pushback
case is unchanged). Full suite unchanged at the pre-existing
30-failed-file/66-failed-test Windows-host baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* refactor(env): replace the growing ad-hoc shell scanner with a narrow, closed recognizer (review)

Round 15 review found nine more genuine parsing bugs, most of them direct
consequences of the general-purpose statement/operand machinery added in
rounds 13-14 (quote/escape-aware `;`/`&&`/`||` splitting, N-way `||`
chains, trailing-argument tolerance on `source`). It also included a
meta-finding, correctly: this had grown into a large, incomplete ad-hoc
shell parser for what should be a narrow forwarding-detection case, and
each round's fix was mostly patching bugs the previous round's own
machinery introduced. Full shell parsing is undecidable without a real
shell; chasing it one adversarial regex at a time was never going to
finish, and it was already producing real regressions (the round-14
`sourceOf` relaxation meant to recognize valid trailing arguments also
started recognizing `source ~/.bashrc | cat` and `source ~/.bashrc &`,
both of which run in a subshell and never actually reach the caller).

Replaced `referencesCandidate`'s open-ended grammar with a closed
recognizer of exactly two forms, each matched as a complete logical line:

- bare unconditional `. REF` / `source REF`
- the self-referential existence guard `test -f REF && . REF` /
  `[ -f REF ] && . REF` — the literal line Git for Windows itself
  generates

Deleted entirely: quote/escape-aware statement splitting (no longer
needed — nothing is split into statements anymore), `||` fallback
handling (both the original two-operand and round 14's N-way
generalization), trailing-argument/redirection tolerance on `source`
(the source of the pipe/background regression above), and
comment-stripping (unnecessary now — a line with anything extra on it
simply fails the exact-match check, which is a large part of why the
statement machinery could be deleted rather than just patched again).

Kept, since dropping them would reopen a real false-positive risk rather
than just narrow scope: block-depth tracking for
`if`/`for`/`while`/`until`/`case`/`select`/function/subshell/brace-group
(content inside is either conditional or non-propagating, generalized
this round with `select` and a fixed one-liner if/for/while/until/case
collapse so a self-contained one-liner doesn't corrupt depth tracking for
the rest of the file), heredoc body skipping (fixed three real bugs in
it: a `<<<` here-string was mistaken for a `<<` heredoc and swallowed the
rest of the file; a non-`-` heredoc's terminator was compared with
`.trim()`, letting an indented look-alike end it early; the delimiter
charset was `\w` only, missing real delimiters like `END-CONFIG`), a
`return`/`exit` halt flag (cheap, and the alternative — textually dead
code after an unconditional exit still being scanned — is a genuine
false positive, however unlikely the pattern), and joining a line ending
in `\`, `&&`, or `||` onto the next (real, unremarkable shell
continuation with no backslash required for the latter two — the risk
this closes isn't hypothetical: an unrelated trailing `&&` followed by an
unconditional-looking `source` on the next line is exactly the shape
that would have produced a false "reachable").

Declined the Debian/Ubuntu nested-`if` finding a fourth time, unchanged
from rounds 12-14: the outer `$BASH_VERSION` check is unverifiable
without a real shell, and this resolver's explicit boundary is that an
unverifiable condition is never trusted. The `||`-existence-only pushback
from round 14 is now moot — `||` isn't recognized in any form.

Net change to shell-profile.ts is negative (-244/+something smaller)
despite fixing more bugs than it added, confirming this is a real
simplification rather than another round of patches. Verified with an
updated unit test suite (61/61 passing — six tests for now-out-of-scope
behavior replaced with tests confirming the safe fallback, new tests
added for every round-15 fix that was kept) and a standalone real-fs
script against the actual built resolver covering all twelve round-15
scenarios (all pass, including the real motivating Git-for-Windows case
and the still-declined Debian template). Full suite unchanged at the
pre-existing 30-failed-file/66-failed-test Windows-host baseline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 14:49:23 +08:00
a52374ab04 fix(ci): make Codex review report all findings and run fork-safe e2e (#739)
Closes #731. Feed current PR body plus earlier review comments into
each pass, stop hand-listing config mocks, and run credential-free e2e
on fork PRs without exposing fixture tokens.

Co-authored-by: review <review@local>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-23 11:41:04 +08:00
Jeff 8d74aa4e42 docs(teamai-skill): fix wrong share-learnings hand-off hint (#736)
Step 9 of setup-admin.md told users to run
`/teamai 我想把学到的经验分享给团队` to share what they learned. This
contradicts the teamai SKILL.md, which states twice that sharing a
session's learnings is automatic and must NOT be routed through
`/teamai` — it is handled by the separate teamai-share-learnings skill.

Replace that bullet with the correct guidance:
- `/teamai` share entry is for publishing a reusable skill
  (contribute-member.md), not loose learnings.
- Session learnings are shared automatically via teamai-share-learnings.
2026-09-23 10:44:53 +08:00
Saul Moro 94eb1484bc fix(init): refuse provider logins without a terminal instead of hanging (#711) (#713)
* fix(init): refuse provider logins without a terminal instead of hanging (#711)

`teamai init` with no session spawned `gh auth login --web` (or `gf auth
login`, `cnb login`) with inherited stdio and waited for a browser device
flow nobody could complete, about five minutes for GitHub, then exited
with the provider's error and no hint of the missing credential.

Cause: the non-TTY guard lived only in utils/prompt.ts. A login is a child
process that owns the terminal, so it never went through that guard.

- `isInteractive()` in utils/prompt.ts: stdin is a TTY and neither `CI`
  nor `TEAMAI_NONINTERACTIVE` is set. Every prompt and the four
  prompt-semantics `isTTY` checks use it; the six hook-payload checks are
  untouched.
- github, tgit and cnb logins throw before spawning when not interactive,
  naming the token variable, the way gitcode already did.
- index.ts exports GIT_TERMINAL_PROMPT=0 when not interactive, so a
  missing clone credential fails at once instead of prompting or opening
  a credential helper dialog. An explicit caller value wins.
- e2e test with a fake `gh` whose `auth login` sleeps: exit 1 in under a
  second naming GITHUB_TOKEN, also under CI=true.

Closes #711

* fix(init): close every git prompt and point TGit at the credential that works

Review follow-up on #713.

- utils/git-env.ts: GIT_TERMINAL_PROMPT=0 only closed git's own terminal
  question. The askpass chain (GUI dialog), ssh's passphrase / unknown-host
  question through /dev/tty, and Git Credential Manager's window each still
  parked an unattended clone until the 180s timeout. All four are now closed
  together (GIT_ASKPASS=echo, GIT_SSH_COMMAND='ssh -o BatchMode=yes',
  GCM_INTERACTIVE=never), each only where the caller set nothing.
- tgit: the guard suggested exporting TGIT_TOKEN, which cannot make an
  unattended run succeed — the PAT is REST-API-only and git.woa.com's git
  endpoint rejects it, so `gf auth whoami` still fails and the clone still has
  no credential. The message now names `gf auth login` (whose stored credential
  is the one that works) and says why the token is not it. Docs follow.
- local-agent: keep askViaTty's non-interactive decline synchronous. Awaiting
  the prompt module's import before declining shifted hook-path timing enough
  to break the once-per-session binding hint (local-agent.test.ts).

* fix(git-env): append batch mode to core.sshCommand instead of replacing it

Review follow-up on #713.

GIT_SSH_COMMAND overrides core.sshCommand rather than extending it, so setting
it blindly dropped a configured custom key, ssh binary or wrapper and left the
run unable to authenticate at all. The value is now composed: read
core.sshCommand and append `-o BatchMode=yes`, or use plain `ssh` when nothing
is configured. A command that already decides BatchMode is left alone, and the
config read is skipped entirely when the caller set GIT_SSH_COMMAND.

Test isolation, so the suite's own result can be trusted:

- shell-profile.test.ts: the three Windows cases never stubbed SHELL, and
  detectShellProfile reads it before the platform branch — a suite run from a
  zsh login shell resolved .zshrc and failed them without ever reaching the
  Windows branch. CI runners use bash, which is why only local runs saw it.
- local-agent.test.ts: the once-per-session binding-hint markers live in
  os.tmpdir() under one shared key, so a leftover marker decided whether the
  next test emitted a hint, and concurrent runs competed for the same paths.
  Each test now gets its own temp directory, which makes the markers per-test
  by construction.

Full suite: 3788 pass, 0 failures, five consecutive runs.

* fix(git-env): leave ssh to each repository instead of a process-wide override

Review follow-up on #713.

GIT_SSH_COMMAND is the only way to reach ssh's batch flag, and it overrides
`core.sshCommand` for *every* later git operation, not just the one the value
was derived from. Reading the launch directory's config and exporting it
process-wide therefore pushed that repo's key or wrapper onto the managed team
repo, and the plain default suppressed a `core.sshCommand` the managed repo had
configured for itself. Prompt suppression must not reach a repository's
transport, so the variable and the `git config` read are gone.

What remains is the three variables that name a prompt and nothing else, so one
value is right for every repository a run touches: GIT_TERMINAL_PROMPT=0,
GIT_ASKPASS=echo and GCM_INTERACTIVE=never.

An ssh remote that would still ask is now documented as the caller's to close,
per repository (`git config core.sshCommand 'ssh -o BatchMode=yes'`) or per run
(`GIT_SSH_COMMAND`). Measured first: with stdin closed, ssh's own tty read hits
EOF and fails in about a second, so the unattended paths this PR is about do
not depend on the flag.

Also reverts the shell-profile.test.ts and local-agent.test.ts isolation edits
from the previous round: neither traces to the unattended-login fix, so they
belong in their own PR.

* docs(providers): mark the provider logins interactive-only (#711)

docs/providers.md still described `teamai init` as running `gh auth login`,
`gf auth login` and `cnb login` unconditionally. Each now happens only in an
interactive terminal; an unattended run fails at once naming the credential to
prepare (a token for GitHub and CNB, a prior `gf auth login` for TGit, since a
TGIT_TOKEN PAT is REST-API-only and cannot clone).
2026-09-23 10:41:13 +08:00
Ben YounesandClaude Opus 5 b91b6dfcae fix(hooks): list each tool's own built-in hook set (#718)
* fix(hooks): list the built-in hooks each tool really receives

`hooks list` rendered builtinHookDefs('claude') for every tool, so the built-in
block described Claude's hook set no matter which tool the row was for:
Copilot's SessionEnd hook never appeared, while tools that receive no hooks at
all were credited with six.

The displayed set is now derived from what each tool actually receives through
reconciliation, and a tool with no hook surface is omitted rather than shown an
invented list.

Fixes #717

* fix(hooks): report adapter-driven tools by their generated artifact

`hooks list` probed a settings file per tool, so Hermes, OpenCode and
OpenClaw — which reconciliation installs as one generated script/plugin/
handler each — fell through to the generic branch and were reported as
"not configured" even right after `hooks inject` wrote their hook. OMP
already had a bespoke branch for this.

Resolve that artifact per adapter and use its presence as the status, the
same rule OMP used, so the status column matches the built-in block.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(hooks): status against the effective built-in set and the user plugin

Two status-column defects the per-tool listing exposed:

- `getHookStatus` checked the unmodified `builtinHookDefs(tool)`, so a
  built-in the team disabled through hooks.yaml was still expected on disk
  and every tool read `missing` right after a correct reconciliation. It
  now takes the same §4.8 override reconciliation applies.
- The OpenCode artifact was probed under the config scope's base dir, but
  `reconcileOpencodePlugin` always installs the single plugin under the
  user path, so a project-scope config reported `missing` after a
  successful injection.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(hooks): require every generated file before reporting installed

The adapter status probe accepted a single file, so an OpenClaw
installation missing its handler.ts — the file HOOK.md points at, without
which no hook runs — still read as `installed`. Check every generated file
the adapter writes and report `installed` only when all are present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-23 10:37:28 +08:00
Saul Moro cd3e0e6ef5 fix(hooks): stop the Stop hook nudge reaching the user twice (#720) 2026-09-22 22:47:35 +08:00
Saul Moro 2ed17e4fb2 feat: scope hooks, MCP servers and env variables by logical project (#700) 2026-09-22 22:44:56 +08:00
PerryLinkandPerryLink 1e2c029466 feat(agents): support Qoder CN via its .qoder-cn user directory (#695)
* feat(agents): support Qoder CN via its .qoder-cn user directory

Qoder CN keeps its user directory at ~/.qoder-cn instead of ~/.qoder, so a CN install saw none of TeamAI's synced resources and users worked around it with symlinks. Register qoder-cn as its own built-in target mirroring qoder: skills discovery, toolPaths defaults, the Claude-compatible agent/MCP formats, and the subagent render/reverse switches. Docs list the new paths.

* fix(agents): resolve Qoder CN paths by scope, not by one flat table

The first cut registered `qoder-cn` with `.qoder-cn/*` as every path, which
wrote project-scope resources into a directory Qoder CN does not read. Only the
*user* directory differs (`~/.qoder-cn`); project scope still uses `.qoder/`.
`scopedToolPaths` documents the contract -- top-level fields are the project
paths, `userScope` carries the user overrides -- so the entry now follows it:
top-level stays `.qoder/*`, `userScope` holds `.qoder-cn/*`, `mcp` is the
user-level MCP file and `mcpProject` the project-level one.

`userScope` did not admit `settings` at all, so a user-scope `qoder-cn` would
have had its hooks written to `~/.qoder/settings.json` -- the international
install's config. `settings` is now part of the userScope schema and its splice.

Hooks do not necessarily live at the config's scope: `resolveHookScope` maps a
non-self project scope to HOME (#264). Every caller that resolved paths from
`localConfig` while using that HOME base directory therefore looked under the
project prefix inside HOME. `resolveHookScope` and
`resolveLegacyProjectHookScope` now report the scope they resolved, and the
inject, remove and check paths use it -- `hooks.ts`, `pull.ts`, `hooks-cmd.ts`
and `doctor.ts`, the last of which is why `teamai doctor` reported "teamai hooks
in qoder-cn settings" as missing.

Verified with a real CLI end-to-end run, not only unit tests. Against
`teamai-hub/teamai-cli-dev` with a scratch HOME, `--scope project --agent
qoder-cn`:

- `~/.qoder-cn/settings.json` is written; `~/.qoder/settings.json` is not
- project scope delivers to `<project>/.qoder/skills`; `.qoder-cn/` is not created
- `teamai doctor` reports `qoder-cn is installed` and `teamai hooks in qoder-cn
  settings` both green (the latter failed before this change)

`npx tsc --noEmit` clean; full suite unchanged at 30 failed / 240 passed files
and 75 failed / 3655 passed tests, the pre-existing Windows failures.

* fix(agents): keep Qoder CN off the paths Qoder already owns

Registering `qoder-cn` with the project paths `.qoder/*` made two targets
resolve to the same project-scope hook file. `reconcileHooksToAllTools` iterates
per tool and `reconcileClaudeFormat` re-renders every built-in entry with the
current tool's dispatch identity, so the later `qoder-cn` pass rewrote a plain
international-Qoder install's file as `--tool qoder-cn` and dropped team hooks
scoped to `tools: [qoder]`. Reproduced with the default configuration (no agent
whitelist) in self mode with `<root>/.qoder` present: 6 built-ins all
`--tool qoder-cn`, 0 `--tool qoder`, and the `tools: [qoder]` hook gone.

Each settings file is now reconciled once, for the first target that reaches it.
The shipped table lists `qoder` before `qoder-cn`, so the existing target keeps
ownership of `<root>/.qoder/settings.json` and CN remains a user-scope addition
there; a CN-only pass still targets the file as `qoder-cn` because nothing else
claims it. `hooksList` mirrors the same rule, so the shared file is listed once
instead of reporting the CN row as missing.

Two user-scope consumers resolved their paths at the config's scope while using
the HOME base directory `resolveHookScope` reports, so they looked in
`~/.qoder/settings.json` for hooks that live in `~/.qoder-cn/settings.json`:
`hooksList` reported a healthy CN install as missing, and `uninstall` discovery
never found the CN hooks and left them behind. Both now use the hook scope.
Skills, rules, agents and claudemd stay at the config scope in uninstall — they
are real project resources, and resolving those to user paths would strand the
project copies.

Corrected the MCP table in both `docs/usage-guide.md` and
`docs/usage-guide.zh-CN.md`, which still said `<project>/.qoder-cn/settings.json`
while `mcpProject` and the surrounding prose use `<project>/.qoder/`.

Known limitation, documented rather than fixed: in project scope a team hook
scoped `tools: [qoder-cn]` no longer lands in the shared file. One physical file
can carry only one dispatch identity, so it is not representable there; teams can
scope to `qoder` or omit `tools:`. Union semantics would need engine support for
alias-aware definitions and is out of scope here.

`tsc --noEmit` clean; full suite 30 failed / 240 passed files unchanged, the
71-test failure set byte-identical, +4 tests from this change.

* fix(cli): resolve the shared hook file by the enabled set on the read path

`hooksList` iterated the whole tool table without applying the enabled set, so
for a self-scope config that enabled Qoder CN alone, `qoder` — disabled, but
earlier in the shipped table — claimed `<root>/.qoder/settings.json`, was probed
for Qoder's dispatch identity, and was reported `missing` while the enabled
`qoder-cn` target never appeared in the table at all.

The write path already got this right: `reconcileHooksToAllTools` skips any tool
outside `filterAgents`, which `reconcileTeamHooksForConfig` derives from
`enabledAgents` and `disabledAgents`. Only the read path was missing the same
rule, so the fix reuses the filter `doctor` already applies to this path table.

Ownership of the shared file therefore follows the enabled set rather than the
table order, and the two paths now agree.

---------

Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com>
2026-09-22 22:23:18 +08:00
Jeff 595034f404 fix(test): eliminate CI flaky tests (#729)
* fix(test): eliminate CI flaky tests

- lock-atomic: accept 1–2 winners under high-load concurrent reclaim
  instead of exactly 1 (fs rename window widens under load)
- dashboard-collector: refresh `now` in beforeEach so timestamps are
  never stale when rebuildSessions checks the 30s expiry window
- codebase-reconcile, import-dir: increase timeout from 15s to 60s
  for tests that hit disk-heavy operations on slow CI

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): address review — keep strict lock assertion, reduce concurrency

Preserve exact-one mutual-exclusion assertion per reviewer feedback.
Reduce concurrency from 32→8 to lower CI load pressure, and add
retry:3 so transient fs scheduling races don't block the pipeline.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): revert lock-atomic changes, keep original assertion

Revert lock-atomic to its original form (32 concurrency, strict
exact-one assertion, no retry). The sentinel serialization is
correct under Node.js single-threaded event loop; the rare 2-winner
case is an OS/fs-level anomaly under extreme CI load, not a product
bug. This PR focuses on the deterministic fixes for the other 3
test files.
2026-09-22 21:14:23 +08:00
JeffandClaude Opus 4.6 9d8ed1faeb fix(test): skip chmod-based test when running as root (#727)
chmod 0o000 has no effect for root (CI runs as root), causing both
pending learnings to be published and the cleanup chmodSync to ENOENT.
Skip the test on root and guard the finally block with existsSync.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-09-22 19:58:41 +08:00
Jeff 109b33480a docs(skill): extract TGit provider guide and cut duplicated caveats in teamai skill (#724)
Consolidate the Tencent TGit (工蜂) flow — reachability probe, gf
install/login, and init repo-creation behavior — into a single
references/provider-tgit.md, which both setup-admin.md and join-member.md
now point to instead of repeating the probe twice and re-describing the
gf login across files.

Also dedup the "don't trust the 'Hooks injected into all AI tool
settings' message" caveat down to the canonical table in
troubleshooting.md, and merge the two overlapping auto-share blockquotes
in SKILL.md into one.

No command or flag changes — cheat sheet stays verbatim against src.
2026-09-22 19:45:27 +08:00
Ruifeng Xue 9ce8a0ed40 fix(import): retry failed repository knowledge extraction (#696)
* fix(import): fail when code knowledge extraction fails

* fix(import): re-extract when manifest is ahead of LAST_SYNC
2026-09-22 17:25:37 +08:00
ydflowandydflow 6d27a11fb9 fix(ci): fail closed when comment lookup fails (#697)
Co-authored-by: ydflow <314143294+ydflow@users.noreply.github.com>
2026-09-22 17:21:54 +08:00
Ben YounesandClaude Opus 5 5d78861d9f fix(local-agent): report user-scope resources from the right base dir (#716)
buildReportPayload scanned the raw project-scope toolPaths map against
$HOME, so tools that relocate or reshape their user scope were scanned in
directories TeamAI never writes to: ~/.github for copilot (its user scope
is $COPILOT_HOME), ~/.opencode for opencode (~/.config/opencode) and
~/.omp for omp (~/.omp/agent). The scan returned an empty list and the
backend saw no user-level skills or rules at all, with no error.

Resolve the scan through the same user-scope seam the installers use:
scopedToolPaths() for the per-scope path overrides and resolveToolBaseDir()
for the base directory.

Fixes #714

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:18:27 +08:00
Jeffandreview fa4f2989c4 fix(webhook): whitelist+redact payload, fix event mapping, preserve secret (#701, #702, #703) (#708)
#701: The webhook handler forwarded the entire hook stdin as `data`, so
wildcard subscriptions leaked tool args (mock API keys) and full
tool_response. Forward only a per-event field whitelist (skillName /
sessionId) and route the outbound body through the shared redact module
as defense-in-depth. Skill-name resolution goes through the shared
resolveSkillUse helper, which validates with isValidSkillName, so an
arbitrary tool-arg string cannot escape through it, and a normal
(non-SKILL.md) Read never produces skill data.

#702: The handler read `stdin.event`, which hosts never send, so every
event was forwarded as `unknown` and skill-use / session-start /
session-stop subscriptions never matched. Map `hook_event_name` to the
canonical event names (case-insensitive, so both Claude's PascalCase and
Cursor/CodeBuddy's camelCase resolve) and drop unmapped events instead of
emitting `unknown`. Register the webhook handler on session-end too, so
Copilot's SessionEnd emits a session-stop notification. Cursor represents
skill use as a `Read` of a SKILL.md file dispatched under the `Skill`
matcher (git f0ab4eb switched Cursor tracking from Read to the Skill
matcher); the webhook now reaches parity with the usage tracker by sharing
resolveSkillUse, which handles both the Skill and Read+SKILL.md shapes.
Wire the promised `push` / `pull` command events from their CLI actions,
gated on a real completion signal so dry-run / cancel / no-change /
handled-failure runs never fire a misleading notification — pushGroup
reports a distinct outcome ('pushed' | 'nochange' | 'pr-failed' |
'failed'). A reuse branch recorded with prUrl:null (an earlier PR creation
failed) now retries PR creation instead of silently skipping it — including
when the tree is unchanged (hasChanges false), where it (re)creates the PR
for the already-pushed branch without re-pushing.

#703: getWebhookSharing dropped `secret` and force-overrode
timeout/retries, so a configured signing key never reached the request
and receivers with signature verification rejected the unsigned webhook.
Preserve all schema-accepted fields and sign the exact bytes sent.

docs: Document sharing.webhooks (config shape, events + timing, whitelisted
payload, signature) in docs/usage-guide.md and docs/usage-guide.zh-CN.md.

Co-authored-by: review <review@local>
2026-09-22 17:04:33 +08:00
Jeff ff0d30c19c docs: slim README to a landing page and move product details out (#722)
Keep README as title, Why TeamAI, quick start, docs links, contributors, and contributing. Move architecture and capability details into product-overview, and the command table into the usage guide.
2026-09-22 16:37:56 +08:00
Jeffandreview 3735a4672d fix(learnings): refresh knowledge branch on pull, fix recall roots + single-branch checkout (#704, #705, #706) (#709)
#704: An independent teamai-learnings update never moves main's revision, so
the "Already synced" fast-return in pull skipped the learnings-branch refresh
and index rebuild — a teammate's contribution only surfaced after `pull --force`.
Hoist the learnings-sync + index-rebuild into a helper and run it on the
fast-return path too. It is read-only (pushIfCreated:false) and never publishes,
so it stays inside the caller's partition sync lock and does not reintroduce the
flush-outside-lock bug.

#705: contribute indexed pending-learnings, then published + deleted the pending
files without refreshing the index, so recall handed back a File path under the
now-emptied pending dir. Rebuild the index once more after a successful publish
so every published entry resolves to its durable worktree copy.

#706: A --single-branch clone's fetch refspec covers only the default branch, so
`fetch origin <branch>` moved only FETCH_HEAD and left origin/<branch> absent,
then `worktree add --track` failed and the knowledge worktree was never created.
Fetch the tracking ref with an explicit refspec, and check the worktree out with
`--no-track` (nothing relies on git upstream config; every sync references
origin/<branch> directly). Shared helper — reports branch verified not regressed.

Regression tests: real-CLI e2e for #704/#705 (learnings-sync-704-705) and a
real-git single-branch checkout test for #706 (git-kind-learnings). Both confirmed
to fail on the pre-fix code and pass after.

Co-authored-by: review <review@local>
2026-09-22 14:17:47 +08:00
dvd233 2c0ae96f6a fix(opencode): spawn hook dispatcher without windows popup (#688) 2026-09-21 23:10:54 +08:00
zdGrandeand林志达 04b6de2219 fix(hooks): run CodeBuddy's Windows hooks through cmd.exe (#637)
* fix(hooks): run CodeBuddy's Windows hooks through cmd.exe

bundled-runtime: CodeBuddy provides cmd.exe on Windows, so the /bin/sh shell gate is a false negative there. POSIX keeps the old check.

builtin-hooks: render cmd-syntax commands for codebuddy on win32 and write a teamai.cmd shim, so the PATHEXT lookup can resolve the CLI.

hooks: render the project gate in the host shell's syntax and recognise both renderings, so a platform switch cannot leave dead duplicates.

tests: win32 cases with a spied platform, so ubuntu CI covers them.

* fix(hooks): exit 0 on a project-gate miss and align the cmd wrapper path

hooks: the cmd gate `${gate} && (payload)` let a gate miss inherit findstr's
exit status — non-zero is surfaced as a hook error, and exit 2 blocks
UserPromptSubmit. Use `& if not errorlevel 1 (payload) else exit /b 0`: a miss
exits 0, a match still passes the payload's own status through.

builtin-hooks: the cmd wrapper searched `%USERPROFILE%\.teamai\bin` while the
writer uses getUserHome(), so HOME/USERPROFILE could diverge and `teamai` was
never found. Embed the resolved bin dir; drop TEAMAI_BIN_DIR_WIN.

tests: update the cmd-gate and codebuddy PATH assertions.

* fix: escape cmd project gate for CodeBuddy hooks on Windows

Cmd shell metacharacters in project roots break the team-hook
project gate; escape the root literal and match with findstr.
Sync tests and docs.

---------

Co-authored-by: 林志达 <linzhida@bosssoft.com.cn>
2026-09-21 23:09:45 +08:00
Hill Patel cc38871818 fix(env): detect the right shell profile file on Windows (#682) (#693) 2026-09-21 18:08:52 +08:00
zerolxy612 f4d09c2cc8 fix: redact prompt summaries before local persistence (#685) 2026-09-21 18:05:01 +08:00
Saul Moro a3366b9bf6 chore(git): ignore local copilot agent sync files (#694)
Ignore Copilot sync targets under .github/ (.github/hooks/teamai.json, .github/skills/, .github/instructions/*.instructions.md, .github/agents/, .github/copilot-instructions.md, and .github/mcp.json) as a best-effort ignore pattern while preserving other GitHub files.
2026-09-21 17:05:52 +08:00
yimi528 1f52967dbd fix(pid-monitor): resolve the process tree on Windows (#691)
`getParentPid()` and `getProcessComm()` read Linux /proc and fall back to
`ps -o`. Windows has neither: /proc does not exist, and the MSYS `ps` Git Bash
ships rejects `-o`. Both therefore returned undefined for every PID,
`resolveMonitorPid()` broke out of its walk on the first level, and the `best`
it returned was the hook's own parent — the launcher, not the AI tool.

`dashboard.ts` hands that value to `isProcessAlive()`, and a false answer
becomes a `process_exit` event, i.e. `status = 'stopped'`: on Windows the exit
signal was tied to the launcher's lifetime instead of the tool's.

Windows now reads the table from one Win32_Process query — the walk needs up to
five ancestors, so a per-pid PowerShell start would cost more than the whole
answer — and the walk itself moved into `walkToNonShell()`, leaving the POSIX
readers exactly as they were. `windowsPowerShell()` moves to
`src/utils/powershell.ts`, shared with the hook dispatcher that already needed
the same path resolution.
2026-09-21 11:28:44 +08:00
Carlos 71b3a5fab5 fix(env): warn when env.yaml has no top-level variables: key (#681)
* fix(env): warn when env.yaml has no top-level `variables:` key

`pullItem` parses env.yaml with a schema whose `variables` field carries
`.default([])`, and zod drops unknown keys without a word. A file that is a
bare `FOO: bar` mapping — or that misspells the key — therefore parsed as "no
variables", and `pullItem` returned at its length check. Every env variable
silently stopped being delivered and nothing in the output said why.

Report that one shape instead of tightening the schema: no `variables` key
plus at least one other top-level key produces a warning that names the keys
found and shows the expected `key`/`value` form.

`.strict()` was the obvious alternative and is deliberately not used. It turns
a valid `variables:` list that carries an extra top-level key into a hard parse
failure that stops env delivery for that team repo — an upgrade regression for
repos that never had a problem. Only the silent no-op is reported; every other
shape stays permissive.

Fixes #662

* fix(env): warn from pullForScope when env.yaml has no `variables:` key

pullForScope skips the env resource as soon as countEnvVars() reports 0
(src/pull.ts:828), and an env.yaml with no top-level `variables:` key answers
exactly that — so a warning raised inside pullItem never runs on a real pull,
which is what the review on #681 found.

The detection moves to a shape probe the orchestration layer can call before
it skips the resource: describeEnvYamlShapeProblem() names the top-level keys
that were found and states the shape `variables:` must have, and
describeEnvYamlShapeProblemAt() reads a file through it.

Only this one shape is reported, per #662: a valid `variables:` list that also
carries an extra top-level key keeps being delivered rather than starting to
fail to parse.

The handler-level tests cover the detection; pull-env-shape-warning.test.ts
drives pull() itself, so the wiring that makes the warning reachable is what
is under test.

* fix(env): also warn about a malformed env.yaml from the unchanged-rev fast path

The review pointed out that `pullForScope` raises the shape warning from the
Step 2 resource loop, and the "Already synced" branch returns before that loop
ever runs. A machine that recorded `lastPullRev` while the CLI still accepted a
bad shape therefore takes the fast path on every later pull and never sees the
warning — which is exactly the population the check exists for: the repo has
not moved, so the rev never changes and the warning can never fire.

The check now lives in `warnIfEnvYamlShapeIsWrong()` and is called from both
sites, so the two cannot drift apart. `resourceTypes` moves above the fast path
so that branch can tell whether this scope syncs env at all.

pull-env-shape-warning.test.ts gains a same-rev case: it pulls twice against the
same mocked rev and asserts both that the fast path was taken (the "Already
synced" success line) and that the warning still fired. Deleting the new call
fails that case alone — the other three stay green.
2026-09-21 10:45:01 +08:00