mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-02 02:07:25 +08:00
fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path > - Paperclip manages agents and the services they need for work. > - Agents use connection search to discover a setup path before they request access. > - Tool-only filtering hid channel and AI methods from this search. > - Requiring every query word to match rejected normal task descriptions. > - A small aggregator index also omitted supported apps such as Circleback. > - This pull request broadens retrieval and returns purpose-specific setup guidance. > - Agents can choose a relevant result while existing access and provider-choice checks still apply. ## Linked Issues or Issue Description **What happened?** A search such as “AgentMail create an email address and manage an agent mailbox” returned no usable result. “Help me find tools for circle back” also missed Composio's supported Circleback toolkit. Queries longer than 200 characters failed validation. **Expected behavior** Return useful native and verified aggregator matches from natural-language queries. Include channel/email methods when Chat connectors is enabled. Identify each method's purpose and the correct setup path. **Steps to reproduce** Use the queries above with `connections_search` from an active task. Enable Chat connectors for the AgentMail case. The regression suite reproduces these misses before the change. Related routing work: #13941. This change does not change the runner failure path or add channel setup to tool-only connection cards. ## What Changed - Rank name and capability matches. Accept extra words, split names, small spelling errors, and queries up to 4,000 characters. - Include tool, channel/email, and AI methods. Return company-prefix setup links for channel and AI flows. - Add a dated snapshot of 1,583 official Composio toolkit names and a refresh script. Merge duplicate MCP variants for search and link each support claim to official evidence. - Find authorized indexed aggregator namespaces within longer queries. Return multiple app matches when the agent needs to choose. - Prefer exact app names over fuzzy matches for other apps; retain existing AI readiness. - Preserve native preference, company and identity boundaries, administrative denials, and saved provider consent. - Add relevance and database regressions, extend native tool-authority coverage, and document search behavior. ## Verification - Red: 15 new assertions failed against the previous implementation; the existing baseline passed. Added red-green regressions for Motion versus fuzzy Notion and existing AI access during review. A further regression covers mixed ready/unconfigured AI results and their per-result setup guidance. - Green: all 103 tests in the eight focused shared, database, runtime-tool, fixture, and route suites pass on the latest commit. - `pnpm -r typecheck` and `pnpm build` passed. - Latest-commit CI passed: 54 successful checks and two skipped checks, including the full test matrix, browser E2E, typecheck, build, and canary dry run. - The long local `pnpm test:run` attempt began before the review fixes and retained transformed pre-fix search code; it also hit an unrelated timing failure. Fresh serial reruns of the affected search suites and three timeout cases passed all 124 tests. Parallel local route shards hit two additional database setup timeouts; both suites passed all 17 tests on a fresh serial rerun. The complete corresponding CI suites also passed. Duplicate broad local runs were stopped after CI completed. The local UI suite independently passed all 7,007 tests. - Greptile: 5/5 on `cc6a0180d`; all review findings resolved. - Browser verification passed in a disposable local instance through real process-agent search requests: AgentMail opened its setup flow with the requester selected; the saved Circleback choice produced the Composio setup card; a paragraph-length Notion query produced its setup card. No provider credentials or external accounts were created. - The browser test caught an invalid UUID-based setup URL. The fix uses the company prefix and has a regression assertion. - A 3,971-character catalog query found Circleback first in a local 10 ms spot check after sharing query preparation across the catalog scan. This is a single measurement, not a performance guarantee. ## Risks - Broader retrieval can return extra candidates. Named services rank first; agents must select the relevant method. - The public support snapshot can age. It proves catalog support, not account authorization or the availability of every requested action. - Channel and AI methods use existing setup links. The tool connection card still accepts tool methods only. - No schema, migration, credential, or runner lifecycle changes. ## Model Used OpenAI Codex (GPT-6). The session does not expose a more specific model identifier or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
# Connection search
|
||||
|
||||
`connections_search` accepts a service name or a natural-language query up to
|
||||
4,000 characters. It ranks names before capabilities, tolerates extra words,
|
||||
split names such as “Agent Mail,” and single-character spelling errors or
|
||||
transpositions. Short partial names work too. Capability searches return
|
||||
overlapping matches without requiring every query word to occur in the catalog.
|
||||
Results are bounded to 40 entries. Exact app names outrank typo matches for
|
||||
different apps (for example, Motion is not replaced by Notion). Query preparation
|
||||
is shared across the catalog scan.
|
||||
|
||||
Discovery includes tool, channel/email, and AI methods. Channel methods follow
|
||||
the instance's Chat connectors experimental setting. Each catalog method names
|
||||
its `purpose`. Channel and AI methods include a company-scoped `setupPath` to
|
||||
their existing setup flow. `connection_request` still creates tool setup cards;
|
||||
channel discovery does not turn an email inbox into an executable MCP tool or
|
||||
add channel setup to that card.
|
||||
|
||||
The agent chooses the relevant match using descriptions and method purposes.
|
||||
It should clarify only when the task remains ambiguous. A search match does not
|
||||
prove that an account is authorized, that a tool is installed, or that a provider
|
||||
can perform every requested action.
|
||||
|
||||
## Aggregator discovery
|
||||
|
||||
The local support index combines reviewed Arcade/Zapier claims with the public
|
||||
[Composio toolkit catalog](https://docs.composio.dev/toolkits). The Composio
|
||||
snapshot includes names and toolkit identifiers, a source URL, and a verification
|
||||
date. Each result links to its official toolkit page. This covers, for example,
|
||||
[Circleback MCP](https://docs.composio.dev/toolkits/circleback_mcp), including
|
||||
queries such as “help me find tools for circle back.”
|
||||
|
||||
Refresh the snapshot from the official public page with:
|
||||
|
||||
```sh
|
||||
node scripts/update-composio-search-catalog.mjs
|
||||
```
|
||||
|
||||
The script also accepts a saved HTML file as its first argument for reproducible
|
||||
extraction. Review the resulting diff before committing it. Searches use this
|
||||
local snapshot and authorized indexed tool namespaces; they do not send user
|
||||
queries to a provider. Broad provider descriptions are not evidence of support.
|
||||
Configured connection metadata remains scoped to the current company and
|
||||
identity. Archived and unauthorized catalogs cannot establish an aggregator route.
|
||||
|
||||
Built-in connections remain preferred for the same app. Existing AI access is
|
||||
checked through the agent's AI binding and credential selection. An aggregator route still requires a
|
||||
saved provider choice or a verified explicit user request. Fuzzy retrieval never
|
||||
grants consent. A provider connection is not proof that its underlying app is
|
||||
authorized. Multiple matching aggregator apps are returned as candidates; the
|
||||
agent searches the selected `aggregator.targetService` to get its provider question.
|
||||
|
||||
## Regression coverage
|
||||
|
||||
```sh
|
||||
pnpm exec vitest run packages/shared/src/connection-search.test.ts packages/shared/src/connection-routing.test.ts packages/shared/src/validators/connection-intent.test.ts server/src/__tests__/connection-aggregator-fallback.test.ts server/src/__tests__/connection-intents-service.test.ts
|
||||
```
|
||||
|
||||
The database suite covers the original AgentMail sentence, split names, typos,
|
||||
capability overlap, multiple services, AI discovery, experimental gating,
|
||||
Circleback/Attio/ClickUp, indexed Executor tools, and provider consent. The native
|
||||
tool authority test searches a paragraph longer than the old 200-character limit
|
||||
and then creates a real connection interaction. Existing tests retain checks for
|
||||
private metadata, administrative denials, stale identities, and saved declines.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -13,7 +13,7 @@ export const CONNECTION_INTENT_AGENT_GUIDANCE = [
|
||||
"- Do not use connection tools for arbitrary MCP URLs or unrelated work.",
|
||||
].join("\n");
|
||||
|
||||
export const CONNECTIONS_SEARCH_TOOL_DESCRIPTION = "Search Paperclip connections first and eligible external aggregator routes when no built-in service matches. Use first when the user asks to connect a service, or when usable access is uncertain; follow the returned instruction and exact providerQuestion, if any. Do not use for arbitrary MCP URLs. Search is read-only.";
|
||||
export const CONNECTIONS_SEARCH_TOOL_DESCRIPTION = "Search connections across tool, channel/email, and AI purposes with a service name or a natural-language description (up to 4000 characters). Results tolerate extra words, split names, and small typos. Choose the relevant match using its description, method purpose, and setupPath; tool methods use connection_request. Also discovers verified external aggregator apps. Use first when the user asks to connect a service, or when usable access is uncertain; follow the returned instruction and exact providerQuestion, if any. Do not use for arbitrary MCP URLs. Search is read-only.";
|
||||
|
||||
export const CONNECTION_REQUEST_TOOL_DESCRIPTION = "Request the service identifier returned by connections_search as available or needs_user_action. Follow the search instruction; aggregator routes require the saved provider-selection interaction ID. Only this tool creates the real setup card. If user action is needed, finish independent work, then yield without retrying or asking for credentials in comments.";
|
||||
|
||||
|
||||
@@ -6,6 +6,8 @@ import {
|
||||
isRemoteMcpConnectorId,
|
||||
type RemoteMcpConnectorId,
|
||||
} from "./remote-mcp-connectors.js";
|
||||
import composioCatalog from "./composio-search-catalog.json" with { type: "json" };
|
||||
import { prepareConnectionSearch, scoreConnectionSearch } from "./connection-search.js";
|
||||
|
||||
export const AGGREGATOR_PRIORITY = [
|
||||
"composio",
|
||||
@@ -24,7 +26,7 @@ export const AGGREGATOR_NAMES: Record<RemoteMcpConnectorId, string> = {
|
||||
* Add only services confirmed in the cited official catalog. No provider credentials.
|
||||
* Executor has no universal app catalog: use the authorized workspace's indexed tools.
|
||||
*/
|
||||
export const AGGREGATOR_SUPPORT_INDEX = [
|
||||
const REVIEWED_AGGREGATOR_SUPPORT = [
|
||||
{
|
||||
slug: "hubspot",
|
||||
name: "HubSpot",
|
||||
@@ -182,6 +184,43 @@ export const AGGREGATOR_SUPPORT_INDEX = [
|
||||
},
|
||||
},
|
||||
] as const;
|
||||
|
||||
export interface AggregatorServiceDefinition {
|
||||
slug: string;
|
||||
name: string;
|
||||
aliases: readonly string[];
|
||||
providers: Partial<Record<RemoteMcpConnectorId, string>>;
|
||||
evidenceUrls?: Partial<Record<RemoteMcpConnectorId, string>>;
|
||||
}
|
||||
|
||||
// Keep the other providers' reviewed support claims, and cover the complete
|
||||
// public Composio catalog instead of maintaining a tiny app-name allowlist.
|
||||
export const AGGREGATOR_SUPPORT_INDEX: AggregatorServiceDefinition[] = REVIEWED_AGGREGATOR_SUPPORT.map(entry => ({ ...entry }));
|
||||
for (const [toolkit, name] of composioCatalog.toolkits as Array<[string, string]>) {
|
||||
const slug = toolkit.replaceAll("_", "-").replace(/^-+/, "");
|
||||
const evidenceUrl = `https://docs.composio.dev/toolkits/${toolkit}`;
|
||||
const existing = AGGREGATOR_SUPPORT_INDEX.find(entry => entry.slug === slug
|
||||
|| (toolkit.endsWith("_mcp") && entry.slug === slug.slice(0, -4)));
|
||||
if (existing) {
|
||||
existing.aliases = [...new Set([...existing.aliases, toolkit, name])];
|
||||
existing.providers = { ...existing.providers, composio: composioCatalog.verifiedAt };
|
||||
existing.evidenceUrls ??= { composio: evidenceUrl };
|
||||
} else {
|
||||
AGGREGATOR_SUPPORT_INDEX.push({ slug, name,
|
||||
aliases: [toolkit, name.replace(/ MCP$/i, "")],
|
||||
providers: { composio: composioCatalog.verifiedAt },
|
||||
evidenceUrls: { composio: evidenceUrl },
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
export function searchAggregatorServices(query: string | ReturnType<typeof prepareConnectionSearch>) {
|
||||
const prepared = typeof query === "string" ? prepareConnectionSearch(query) : query;
|
||||
return AGGREGATOR_SUPPORT_INDEX.map(service => ({ service,
|
||||
...scoreConnectionSearch(prepared, [service.slug, service.name, ...service.aliases]),
|
||||
})).filter(match => match.nameScore > 0 && !isRemoteMcpConnectorId(match.service.slug))
|
||||
.sort((a, b) => b.score - a.score || a.service.slug.localeCompare(b.service.slug));
|
||||
}
|
||||
export const AGGREGATOR_CATALOG_SOURCES: Partial<
|
||||
Record<RemoteMcpConnectorId, string>
|
||||
> = {
|
||||
@@ -197,7 +236,7 @@ export function normalizeConnectionQuery(value: string): string {
|
||||
.trim();
|
||||
}
|
||||
|
||||
export function findAggregatorService(query: string) {
|
||||
export function findAggregatorService(query: string, rankedMatches?: ReturnType<typeof searchAggregatorServices>) {
|
||||
const normalized = normalizeConnectionQuery(query);
|
||||
const exact = AGGREGATOR_SUPPORT_INDEX.find((entry) =>
|
||||
[entry.slug, entry.name, ...entry.aliases].some(
|
||||
@@ -207,12 +246,14 @@ export function findAggregatorService(query: string) {
|
||||
if (exact) return exact;
|
||||
// Agents often include the desired capability ("HubSpot recent contacts").
|
||||
// Whole phrases avoid substring guesses; multiple named services need clarification.
|
||||
const matches = AGGREGATOR_SUPPORT_INDEX.filter((entry) =>
|
||||
[entry.slug, entry.name, ...entry.aliases].some((name) =>
|
||||
` ${normalized} `.includes(` ${normalizeConnectionQuery(name)} `),
|
||||
),
|
||||
);
|
||||
return matches.length === 1 ? matches[0] : undefined;
|
||||
const matches = (rankedMatches ?? searchAggregatorServices(query))
|
||||
.filter(match => match.nameScore >= 500).map(match => match.service);
|
||||
// "Atlassian Jira" identifies Jira, not both Jira and Atlassian's MCP.
|
||||
const phrases = (entry: AggregatorServiceDefinition) => [entry.slug, entry.name, ...entry.aliases]
|
||||
.map(normalizeConnectionQuery).filter(name => ` ${normalized} `.includes(` ${name} `));
|
||||
const specific = matches.filter(entry => !matches.some(other => other !== entry
|
||||
&& phrases(entry).every(name => phrases(other).some(longer => longer !== name && ` ${longer} `.includes(` ${name} `)))));
|
||||
return specific.length === 1 ? specific[0] : undefined;
|
||||
}
|
||||
|
||||
/** Preserve a provider explicitly named by the user instead of applying default ranking. */
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { scoreConnectionSearch } from "./connection-search.js";
|
||||
import { AGGREGATOR_SUPPORT_INDEX, searchAggregatorServices } from "./connection-routing.js";
|
||||
|
||||
describe("connection search relevance", () => {
|
||||
it.each(["agentmail", "agent mail", "AGENT-MAIL", "Agentmial", "Agentmai", "Agentmaill"])(
|
||||
"recognizes a name in a long query: %s", name => {
|
||||
expect(scoreConnectionSearch(`Please acquire an email address through ${name} and remember it`, ["AgentMail"]).nameScore).toBeGreaterThan(0);
|
||||
},
|
||||
);
|
||||
it("ranks a named service above generic capability overlap", () => {
|
||||
const query = "Help me connect Linear to read project documentation";
|
||||
expect(scoreConnectionSearch(query, ["Linear"]).score)
|
||||
.toBeGreaterThan(scoreConnectionSearch(query, ["Notion"], "Read project documentation").score);
|
||||
});
|
||||
it("ignores repeated filler and retains capability matches", () => {
|
||||
const base = scoreConnectionSearch("spreadsheets", ["Sheets"], "Read spreadsheets");
|
||||
expect(scoreConnectionSearch("please please help me find tools for spreadsheets", ["Sheets"], "Read spreadsheets").score).toBe(base.score);
|
||||
});
|
||||
it.each(["email", "contacts"])("retains the capability word %s without treating it as a product name", word => {
|
||||
expect(scoreConnectionSearch(`Find a connection for ${word}`, ["Workspace"], `Search ${word}`).score).toBeGreaterThan(0);
|
||||
expect(scoreConnectionSearch(`Find a connection for ${word}`, [word]).nameScore).toBe(0);
|
||||
});
|
||||
it("supports a partial name without turning substrings of prose into service claims", () => {
|
||||
expect(scoreConnectionSearch("agen", ["AgentMail"]).nameScore).toBeGreaterThan(0);
|
||||
expect(scoreConnectionSearch("notionally similar", ["Notion"]).nameScore).toBe(0);
|
||||
expect(scoreConnectionSearch("please find tools", ["Findymail"]).nameScore).toBe(0);
|
||||
expect(scoreConnectionSearch("our team's projects", ["Teams"]).nameScore).toBe(0);
|
||||
});
|
||||
it("indexes the public aggregator catalog with valid routes and evidence", () => {
|
||||
expect(AGGREGATOR_SUPPORT_INDEX.length).toBeGreaterThan(1000);
|
||||
expect(new Set(AGGREGATOR_SUPPORT_INDEX.map(app => app.slug)).size).toBe(AGGREGATOR_SUPPORT_INDEX.length);
|
||||
for (const app of AGGREGATOR_SUPPORT_INDEX) expect(app.slug).toMatch(/^[a-z0-9][a-z0-9-]{0,79}$/);
|
||||
expect(searchAggregatorServices("circle back meeting notes")[0]?.service).toMatchObject({
|
||||
slug: "circleback-mcp", evidenceUrls: { composio: "https://docs.composio.dev/toolkits/circleback_mcp" },
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,73 @@
|
||||
/** Retrieval only. A fuzzy match is never evidence of authorization or consent. */
|
||||
export function normalizeConnectionSearch(value: string): string {
|
||||
return value.normalize("NFKD").replace(/\p{M}/gu, "").toLowerCase().replace(/['’]s\b/g, "")
|
||||
.replace(/[^\p{L}\p{N}]+/gu, " ").trim();
|
||||
}
|
||||
|
||||
const STOP_WORDS = new Set("a an and are as at be by can connect connection connections do find for from get help i in is it me my need not of on or our please real service services so some that the their them there this to tool tools use want we with would you your".split(" "));
|
||||
const GENERIC_NAMES = new Set([...STOP_WORDS, "contacts", "email"]);
|
||||
|
||||
/** Reuse normalized words and deduplicated name windows across a catalog scan. */
|
||||
export function prepareConnectionSearch(query: string) {
|
||||
const normalized = normalizeConnectionSearch(query);
|
||||
const words = normalized ? normalized.split(" ") : [];
|
||||
const phrases = new Set<string>();
|
||||
const phrasesByLength = new Map<number, Set<string>>();
|
||||
for (let start = 0; start < words.length; start++) {
|
||||
let phrase = "";
|
||||
for (let size = 1; size <= 6 && start + size <= words.length; size++) {
|
||||
phrase += words[start + size - 1];
|
||||
phrases.add(phrase);
|
||||
const bucket = phrasesByLength.get(phrase.length) ?? new Set<string>();
|
||||
bucket.add(phrase);
|
||||
phrasesByLength.set(phrase.length, bucket);
|
||||
}
|
||||
}
|
||||
return { normalized, compact: words.join(""), phrases, phrasesByLength,
|
||||
terms: [...new Set(words.filter(word => word.length > 1 && !STOP_WORDS.has(word)))],
|
||||
};
|
||||
}
|
||||
|
||||
function oneEditApart(a: string, b: string): boolean {
|
||||
if (Math.abs(a.length - b.length) > 1) return false;
|
||||
let i = 0;
|
||||
while (i < Math.min(a.length, b.length) && a[i] === b[i]) i++;
|
||||
if (a.length === b.length) {
|
||||
return a.slice(i + 1) === b.slice(i + 1)
|
||||
|| (a[i] === b[i + 1] && a[i + 1] === b[i] && a.slice(i + 2) === b.slice(i + 2));
|
||||
}
|
||||
return a.length > b.length ? a.slice(i + 1) === b.slice(i) : a.slice(i) === b.slice(i + 1);
|
||||
}
|
||||
|
||||
export function scoreConnectionSearch(query: string | ReturnType<typeof prepareConnectionSearch>, names: readonly string[], description = "") {
|
||||
const prepared = typeof query === "string" ? prepareConnectionSearch(query) : query;
|
||||
const { normalized } = prepared;
|
||||
if (!normalized) return { score: 1, nameScore: 0 };
|
||||
let nameScore = 0;
|
||||
for (const name of names) {
|
||||
const normalizedName = normalizeConnectionSearch(name);
|
||||
const compact = normalizedName.replaceAll(" ", "");
|
||||
if (!compact) continue;
|
||||
if (compact === prepared.compact) {
|
||||
nameScore = 1000;
|
||||
break;
|
||||
}
|
||||
if (normalizedName.split(" ").every(word => GENERIC_NAMES.has(word))) continue;
|
||||
if (prepared.phrases.has(compact)) {
|
||||
nameScore = Math.max(nameScore, 500 + compact.length);
|
||||
continue;
|
||||
}
|
||||
if (compact.length >= 5 && nameScore < 500) {
|
||||
const fuzzy = [compact.length - 1, compact.length, compact.length + 1].some(length =>
|
||||
length >= 5 && [...(prepared.phrasesByLength.get(length) ?? [])].some(phrase => oneEditApart(phrase, compact)));
|
||||
if (fuzzy) nameScore = Math.max(nameScore, 200 + compact.length);
|
||||
}
|
||||
if (prepared.compact.length >= 4 && compact.startsWith(prepared.compact)) {
|
||||
nameScore = Math.max(nameScore, 100 + prepared.compact.length);
|
||||
}
|
||||
}
|
||||
const haystack = new Set(normalizeConnectionSearch(description).split(" "));
|
||||
const matched = prepared.terms.filter(term => haystack.has(term)
|
||||
|| (term.length >= 4 && [...haystack].some(word => word.startsWith(term) || (word.length >= 4 && term.startsWith(word)))));
|
||||
return { nameScore, score: nameScore + Math.min(80, matched.length * 5) };
|
||||
}
|
||||
@@ -2783,6 +2783,7 @@ export * from "./slack-tools.js";
|
||||
|
||||
export { MEMORY_CONNECTOR_IDS, isMemoryConnectorId, type MemoryConnectorId } from "./memory-connectors.js";
|
||||
export * from "./connection-routing.js";
|
||||
export * from "./connection-search.js";
|
||||
|
||||
export { WORKSPACE_RESTORE_FAILURE_CODES, hasWorkspaceRestoreFailure, safeWorkspaceRestorePath, isNativeWorkspaceExportRepairCause } from "./workspace-restore.js";
|
||||
|
||||
|
||||
@@ -28,6 +28,10 @@ export interface ConnectionSearchResultItem {
|
||||
key: string;
|
||||
label: string;
|
||||
auth: "oauth" | "api_key" | "none";
|
||||
/** Discovery is open across purposes; connection_request creates tool cards. */
|
||||
purpose?: "tool" | "channel" | "ai";
|
||||
/** Company-scoped setup destination for methods with a separate setup flow. */
|
||||
setupPath?: string;
|
||||
}>;
|
||||
state: ConnectionAvailabilityState;
|
||||
connectionId: string | null;
|
||||
|
||||
@@ -63,4 +63,9 @@ describe("connection intent contracts", () => {
|
||||
expect(connectionsSearchInputSchema.parse({ query: " notion " })).toEqual({ query: "notion" });
|
||||
expect(connectionRequestInputSchema.parse({ service: " notion " })).toEqual({ service: "notion" });
|
||||
});
|
||||
it("accepts long natural-language searches while bounding input size", () => {
|
||||
const query = "Please help me find a connection. ".repeat(30) + "AgentMail inbox";
|
||||
expect(connectionsSearchInputSchema.parse({ query }).query).toBe(query);
|
||||
expect(connectionsSearchInputSchema.safeParse({ query: "x".repeat(4001) }).success).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
import { z } from "zod";
|
||||
|
||||
export const connectionsSearchInputSchema = z.object({
|
||||
query: z.string().trim().max(200).default(""),
|
||||
query: z.string().trim().max(4000).default(""),
|
||||
retryProviderChoice: z.boolean().optional().describe("Only when the user explicitly asks to reconsider a previous provider choice or decline"),
|
||||
}).strict();
|
||||
|
||||
|
||||
@@ -0,0 +1,28 @@
|
||||
import { readFile, writeFile } from "node:fs/promises";
|
||||
|
||||
// Public names and toolkit identifiers only. Search never sends user queries
|
||||
// to this website or requires a Composio account.
|
||||
const source = "https://docs.composio.dev/toolkits";
|
||||
const response = process.argv[2] ? null : await fetch(source);
|
||||
if (response && !response.ok) throw new Error(`Catalog returned ${response.status}`);
|
||||
const html = process.argv[2] ? await readFile(process.argv[2], "utf8") : await response.text();
|
||||
const decode = (value) => value.replace(/&#x([0-9a-f]+);|&#(\d+);|&(amp|quot|apos|lt|gt);/gi,
|
||||
(_, hex, decimal, named) => hex || decimal ? String.fromCodePoint(parseInt(hex ?? decimal, hex ? 16 : 10))
|
||||
: ({ amp: "&", quot: '"', apos: "'", lt: "<", gt: ">" })[named.toLowerCase()]);
|
||||
const entries = new Map();
|
||||
for (const match of html.matchAll(/<a\b[^>]*href="\/toolkits\/([a-z0-9_]+)"[^>]*>([\s\S]*?)<\/a>/g)) {
|
||||
const name = match[2].match(/<span\b[^>]*class="truncate text-sm font-medium[^"\n]*"[^>]*>([^<]+)<\/span>/)?.[1];
|
||||
if (name) entries.set(match[1], decode(name));
|
||||
}
|
||||
if (entries.size < 1000) throw new Error(`Catalog extraction returned only ${entries.size} entries; inspect the source before updating`);
|
||||
const rows = [...entries].sort(([a], [b]) => a.localeCompare(b));
|
||||
const result = `{
|
||||
"source": ${JSON.stringify(source)},
|
||||
"verifiedAt": ${JSON.stringify(new Date().toISOString().slice(0, 10))},
|
||||
"toolkits": [
|
||||
${rows.map(row => ` ${JSON.stringify(row)}`).join(",\n")}
|
||||
]
|
||||
}
|
||||
`;
|
||||
await writeFile(new URL("../packages/shared/src/composio-search-catalog.json", import.meta.url), result);
|
||||
console.log(`Saved ${rows.length} public toolkit names from ${source}`);
|
||||
@@ -21,6 +21,7 @@ import {
|
||||
} from "@paperclipai/db";
|
||||
import type { RuntimeToolsTokenClaims } from "../runtime-tools-token.js";
|
||||
import { connectionIntentService } from "../services/connection-intents.js";
|
||||
import { instanceSettingsService } from "../services/instance-settings.js";
|
||||
import { issueThreadInteractionService } from "../services/issue-thread-interactions.js";
|
||||
import {
|
||||
getEmbeddedPostgresTestSupport,
|
||||
@@ -256,6 +257,100 @@ const support = await getEmbeddedPostgresTestSupport();
|
||||
}
|
||||
return connection!;
|
||||
}
|
||||
it.each([
|
||||
["AgentMail create an email address and manage an agent mailbox", "agentmail"],
|
||||
["Please acquire an email address through agent mail and remember it", "agentmail"],
|
||||
["I need an Agentmial inbox for receiving mail from customers", "agentmail"],
|
||||
["Please connect my Linear workspace so I can triage the team's backlog", "linear"],
|
||||
["Find a Notion connection to search our engineering documentation", "notion"],
|
||||
["Can you find the OpenRouter connection for the model I want to use?", "openrouter"],
|
||||
["Please help me connect to Git Hub to look at my pull requests", "github"],
|
||||
])("finds native connections in natural-language queries: %s", async (query, service) => {
|
||||
await resetQuestions();
|
||||
await instanceSettingsService(db).updateExperimental({ enableChatConnectors: true });
|
||||
try {
|
||||
const result = await connectionIntentService(db).search(claims, query);
|
||||
expect(result.results[0]).toMatchObject({ service, state: "available" });
|
||||
expect(result.providerQuestion).toBeUndefined();
|
||||
} finally {
|
||||
await instanceSettingsService(db).updateExperimental({ enableChatConnectors: false });
|
||||
}
|
||||
});
|
||||
it("discovers all purposes with truthful setup actions and the chat feature gate", async () => {
|
||||
await resetQuestions();
|
||||
const service = connectionIntentService(db);
|
||||
expect((await service.search(claims, "agentmail")).results.some(item => item.service === "agentmail")).toBe(false);
|
||||
await instanceSettingsService(db).updateExperimental({ enableChatConnectors: true });
|
||||
try {
|
||||
const result = await service.search(claims, "agentmail");
|
||||
expect(result.results[0]?.methods).toEqual([expect.objectContaining({
|
||||
key: "email-agent", purpose: "channel", setupPath: `/AGG/apps/chat/connect?provider=agentmail&purpose=chat&agentId=${claims.sub}`,
|
||||
})]);
|
||||
expect(result.instruction).toContain("setupPath");
|
||||
const browse = await service.search(claims, "");
|
||||
expect(browse.results.some(item => item.service === "agentmail")).toBe(true);
|
||||
expect(browse.results.some(item => item.service === "anthropic")).toBe(true);
|
||||
} finally {
|
||||
await instanceSettingsService(db).updateExperimental({ enableChatConnectors: false });
|
||||
}
|
||||
});
|
||||
it("keeps useful capability matches even when other query words do not occur in the catalog", async () => {
|
||||
await resetQuestions();
|
||||
const result = await connectionIntentService(db).search(claims,
|
||||
"I need something that can search documents and spreadsheets for our quarterly planning discussion");
|
||||
expect(result.results.some(item => item.service === "google-drive")).toBe(true);
|
||||
expect(result.results.some(item => item.service === "google-sheets")).toBe(true);
|
||||
expect(result.instruction).not.toContain("search its exact name");
|
||||
});
|
||||
it.each([
|
||||
["help me find tools for circle back", "circleback-mcp"],
|
||||
["Find Circlebak meeting transcripts and action items from yesterday", "circleback-mcp"],
|
||||
["Can you connect Circleback MCP to get all our meeting notes?", "circleback-mcp"],
|
||||
["Find Attio tools to review all our customer contacts before next week's meeting", "attio"],
|
||||
["Help me find a ClickUp connection to organize our team's projects", "clickup"],
|
||||
])("finds verified aggregator apps in natural-language queries: %s", async (query, target) => {
|
||||
await resetQuestions();
|
||||
const result = await connectionIntentService(db).search(claims, query);
|
||||
expect(result.results).toEqual(expect.arrayContaining([expect.objectContaining({
|
||||
service: `via:composio:${target}`, source: "aggregator",
|
||||
aggregator: expect.objectContaining({ evidenceUrl: expect.stringContaining("composio.dev/"), targetService: target }),
|
||||
})]));
|
||||
expect(result.providerQuestion?.options.map(option => option.id)).toContain(`via:composio:${target}`);
|
||||
await expect(connectionIntentService(db).request(claims, `via:composio:${target}`)).rejects.toThrow();
|
||||
});
|
||||
it.each([false, true])("prefers an exact Motion match over fuzzy Notion (Notion denied: %s)", async denied => {
|
||||
await resetQuestions();
|
||||
if (denied) await seedProvider("notion", "notion_search", "responsible-user", false);
|
||||
const result = await connectionIntentService(db).search(claims, "Help me find Motion tools");
|
||||
expect(result.results[0]?.service).toBe("via:composio:motion");
|
||||
expect(result.providerQuestion?.id).toBe("connection-provider:motion");
|
||||
expect(result.instruction).not.toContain("administratively restricted");
|
||||
});
|
||||
it("finds namespaced installed aggregator tools with extra query words", async () => {
|
||||
await resetQuestions();
|
||||
await seedProvider("executor", "heliotrope:list_records");
|
||||
const result = await connectionIntentService(db).search(claims, "Please find heliotrope tools for our weekly report");
|
||||
expect(result.results[0]?.service).toBe("via:executor:heliotrope");
|
||||
});
|
||||
it("returns multiple named services so the agent can choose the relevant result", async () => {
|
||||
await resetQuestions();
|
||||
const result = await connectionIntentService(db).search(claims, "Find Linear or Notion for our project planning");
|
||||
expect(result.results.map(item => item.service)).toEqual(expect.arrayContaining(["linear", "notion"]));
|
||||
});
|
||||
it("returns multiple aggregator apps without inventing a combined route or choosing consent", async () => {
|
||||
await resetQuestions();
|
||||
const result = await connectionIntentService(db).search(claims, "Find Circleback and Attio tools");
|
||||
expect(result.results.map(item => item.service)).toEqual(expect.arrayContaining(["via:composio:circleback-mcp", "via:composio:attio"]));
|
||||
expect(result.providerQuestion).toBeUndefined();
|
||||
expect(result.instruction).toContain("aggregator.targetService");
|
||||
});
|
||||
it("keeps native and aggregator app matches from the same query", async () => {
|
||||
await resetQuestions();
|
||||
const result = await connectionIntentService(db).search(claims, "Find Notion and Circleback tools");
|
||||
expect(result.results.map(item => item.service)).toEqual(expect.arrayContaining(["notion", "via:composio:circleback-mcp"]));
|
||||
expect(result.providerQuestion).toBeUndefined();
|
||||
});
|
||||
|
||||
it("prefers the built-in Jira connection and returns a direct instruction", async () => {
|
||||
await resetQuestions();
|
||||
const result = await connectionIntentService(db).search(
|
||||
|
||||
@@ -671,7 +671,8 @@ describeEmbeddedPostgres("connectionIntentService", () => {
|
||||
await db.update(issues).set({ executionRunId: runId }).where(eq(issues.id, issueId));
|
||||
await db.update(heartbeatRuns).set({ runtimeMode: "native", nativeIssueId: issueId }).where(eq(heartbeatRuns.id, runId));
|
||||
const authority = new PaperclipRunnerToolAuthority(db, { companyId: claims.company_id, issueId, agentId: claims.sub, runId });
|
||||
const result = await authority.execute({ tool: "connections_search", callId: "discover", arguments: { query: "github" } });
|
||||
const query = "Please help me find tools to continue this task. ".repeat(8) + "Look at my Git Hub pull requests";
|
||||
const result = await authority.execute({ tool: "connections_search", callId: "discover", arguments: { query } });
|
||||
expect(result).toMatchObject({ results: expect.arrayContaining([expect.objectContaining({ service: "github", state: "available" })]) });
|
||||
const request = await authority.execute({ tool: "connection_request", callId: "request", arguments: { service: "github" } });
|
||||
expect(request).toMatchObject({ state: "needs_user_action", interactionId: expect.any(String) });
|
||||
@@ -840,8 +841,20 @@ describeEmbeddedPostgres("connectionIntentService", () => {
|
||||
await expect(service.complete(toolRequest.id, connection!.id, claims.responsible_user_id!)).rejects.toThrow("cannot satisfy");
|
||||
await expect(service.complete(aiRequest.interactionId!, connection!.id, claims.responsible_user_id!)).resolves.toMatchObject({ status: "accepted" });
|
||||
expect((await service.request(aiClaims, "anthropic", { purpose: "ai" })).state).toBe("ready");
|
||||
const search = await service.search(aiClaims, "Find Anthropic authentication for this agent");
|
||||
expect(search.results[0]).toMatchObject({ service: "anthropic", state: "ready", connectionId: connection!.id });
|
||||
expect(search.instruction).not.toContain("Share its setupPath");
|
||||
const multipleAi = await service.search(aiClaims, "Anthropic and xAI");
|
||||
expect(multipleAi.results).toEqual(expect.arrayContaining([
|
||||
expect.objectContaining({ service: "anthropic", state: "ready" }),
|
||||
expect.objectContaining({ service: "xai", state: "available" }),
|
||||
]));
|
||||
expect(multipleAi.instruction).toContain("setupPath");
|
||||
expect(multipleAi.instruction).toContain("ready AI");
|
||||
await expect(service.request(aiClaims, "anthropic")).rejects.toMatchObject({ status: 422 });
|
||||
expect((await service.search(aiClaims, "openrouter")).results.some(result => result.service === "openrouter")).toBe(false);
|
||||
expect((await service.search(aiClaims, "openrouter")).results[0]).toMatchObject({
|
||||
service: "openrouter", methods: [expect.objectContaining({ purpose: "ai", setupPath: expect.any(String) })],
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ describe("runtime connection MCP contract", () => {
|
||||
inputSchema: {
|
||||
type: "object",
|
||||
properties: {
|
||||
query: { type: "string" },
|
||||
query: { type: "string", maxLength: 4000 },
|
||||
retryProviderChoice: {
|
||||
type: "boolean",
|
||||
description: "Only when the user explicitly asks to reconsider a previous provider choice or decline",
|
||||
|
||||
@@ -7,6 +7,7 @@ import { and, eq, desc, isNull, sql } from "drizzle-orm";
|
||||
import type { Db } from "@paperclipai/db";
|
||||
import {
|
||||
agents,
|
||||
companies,
|
||||
toolConnections,
|
||||
toolCatalogEntries,
|
||||
companyMemberships,
|
||||
@@ -18,13 +19,14 @@ import {
|
||||
import {
|
||||
APP_STORE_DEFINITIONS,
|
||||
AGGREGATOR_PRIORITY, AGGREGATOR_NAMES, AGGREGATOR_CATALOG_SOURCES,
|
||||
findAggregatorService, explicitAggregatorQuery, normalizeConnectionQuery, parseAggregatorRoute, aggregatorProviderQuestion,
|
||||
findAggregatorService, searchAggregatorServices, prepareConnectionSearch, scoreConnectionSearch, explicitAggregatorQuery, normalizeConnectionQuery, parseAggregatorRoute, aggregatorProviderQuestion,
|
||||
aggregatorContinuationInstruction, isRemoteMcpConnectorId, askUserQuestionsPayloadSchema, askUserQuestionsResultSchema,
|
||||
CONNECTABLE_APP_DEFINITIONS,
|
||||
connectionIntentPayloadSchema,
|
||||
getAvailableConnectionMethods,
|
||||
isToolConnectionAttentionHealth,
|
||||
type ConnectionSearchResultItem,
|
||||
type AggregatorServiceDefinition,
|
||||
getAppStoreDefinition,
|
||||
type ConnectionIntentInteraction,
|
||||
type ConnectionIntentSetupOptions,
|
||||
@@ -35,6 +37,7 @@ import {
|
||||
type AiConnectionAttribution,
|
||||
type AiConnectionBinding,
|
||||
} from "@paperclipai/shared";
|
||||
import { instanceSettingsService } from "./instance-settings.js";
|
||||
import { conflict, forbidden, notFound, unprocessable } from "../errors.js";
|
||||
import type { RuntimeToolsTokenClaims } from "../runtime-tools-token.js";
|
||||
import { issueThreadInteractionService } from "./issue-thread-interactions.js";
|
||||
@@ -351,11 +354,18 @@ export function connectionIntentService(db: Db) {
|
||||
|
||||
async function search(claims: ConnectionRunClaims, query: string, options: { retryProviderChoice?: boolean } = {}): Promise<ConnectionsSearchResult> {
|
||||
const { run, agent, issue } = await loadRunContext(claims);
|
||||
const normalized = query.trim().toLocaleLowerCase();
|
||||
const tokens = normalized.split(/[^\p{L}\p{N}]+/u).filter(Boolean);
|
||||
const settings = await instanceSettingsService(db).getExperimental();
|
||||
const [company] = await db.select({ prefix: companies.issuePrefix }).from(companies).where(eq(companies.id, run.companyId));
|
||||
if (!company) throw notFound("Company was not found");
|
||||
const explicit = explicitAggregatorQuery(query);
|
||||
const serviceQuery = explicit?.serviceQuery ?? query;
|
||||
const preparedQuery = prepareConnectionSearch(serviceQuery);
|
||||
const inventory = await connectionInventory(run.companyId);
|
||||
const candidates: Array<{ item: ConnectionSearchResultItem; score: number }> = [];
|
||||
const services = [...APP_STORE_DEFINITIONS.filter(app => getAvailableConnectionMethods(app).some(method => method.transport !== "runtime_auth")).map((app) => app.slug),
|
||||
const candidates: Array<{ item: ConnectionSearchResultItem; score: number; nameScore: number }> = [];
|
||||
const authorizedCatalogs = new Map<string, Awaited<ReturnType<typeof indexedCatalog>>>();
|
||||
const discoveryMethods = (app: (typeof APP_STORE_DEFINITIONS)[number]) => getAvailableConnectionMethods(app)
|
||||
.filter(method => method.purpose !== "channel" || settings.enableChatConnectors);
|
||||
const services = [...APP_STORE_DEFINITIONS.filter(app => discoveryMethods(app).length).map((app) => app.slug),
|
||||
...inventory.connections.filter((connection) =>
|
||||
sourceSlugForConnection(connection, inventory.applicationsById)?.startsWith("connection:")
|
||||
&& connection.status !== "archived").map((connection) => `connection:${connection.id}`)];
|
||||
@@ -363,8 +373,22 @@ export function connectionIntentService(db: Db) {
|
||||
let app;
|
||||
try { app = await resolveService(service, run.companyId, run.responsibleUserId!, agent.id); }
|
||||
catch (error) { if (service.startsWith("connection:") && (error as { status?: number }).status === 404) continue; throw error; }
|
||||
const definition = getAppStoreDefinition(service);
|
||||
if (definition) {
|
||||
const methods = discoveryMethods(definition);
|
||||
app = { ...app, searchCapabilities: methods.map(method => `${method.whenToUse} ${method.label ?? ""} ${method.capabilityProfile?.description ?? ""}`).join(" "),
|
||||
methods: methods.map(method => ({ key: method.key, label: method.label ?? method.key, auth: method.auth,
|
||||
purpose: method.purpose ?? "tool",
|
||||
...(method.purpose === "channel" && method.provider ? {
|
||||
setupPath: `/${company.prefix}/apps/chat/connect?${new URLSearchParams({ provider: method.provider, purpose: "chat", agentId: agent.id })}`,
|
||||
} : method.purpose === "ai" ? { setupPath: `/${company.prefix}/apps/connect?${new URLSearchParams({ source: service })}` } : {}),
|
||||
})),
|
||||
};
|
||||
}
|
||||
const aiOnly = definition && discoveryMethods(definition).every(method => method.purpose === "ai");
|
||||
const matching = inventory.connections.filter((connection) =>
|
||||
sourceSlugForConnection(connection, inventory.applicationsById) === service && connection.status !== "archived" && connection.connectionPurpose !== "ai");
|
||||
sourceSlugForConnection(connection, inventory.applicationsById) === service && connection.status !== "archived"
|
||||
&& (aiOnly ? connection.connectionPurpose === "ai" : connection.connectionPurpose !== "ai"));
|
||||
// Indexed descriptions can contain private workspace metadata, including
|
||||
// for catalog providers. Check each configured connection's audience first.
|
||||
const catalogs = await Promise.all(matching.map(async (connection) => {
|
||||
@@ -373,17 +397,18 @@ export function connectionIntentService(db: Db) {
|
||||
grant.kind === "organization" || (grant.kind === "user" && grant.subjectUserId === run.responsibleUserId)
|
||||
|| (grant.kind === "agent" && grant.subjectAgentId === agent.id)
|
||||
));
|
||||
return authorized ? indexedCatalog(connection.id, run.companyId) : [];
|
||||
const entries = authorized ? await indexedCatalog(connection.id, run.companyId) : [];
|
||||
authorizedCatalogs.set(connection.id, entries);
|
||||
return entries;
|
||||
}));
|
||||
const catalog = catalogs.flat().filter((entry) => entry.status === "active");
|
||||
const haystack = `${app.slug} ${app.name} ${app.description ?? ""} ${app.searchCapabilities} ${catalog.map((tool) => `${tool.toolName} ${tool.description ?? ""}`).join(" ")}`.toLocaleLowerCase();
|
||||
const score = !normalized ? 1 : app.slug === normalized || app.name.toLocaleLowerCase() === normalized
|
||||
? 1000 : tokens.reduce((sum, token) => sum + (haystack.includes(token) ? 1 : 0), 0);
|
||||
const { score, nameScore } = scoreConnectionSearch(preparedQuery, [app.slug, app.name],
|
||||
`${app.description ?? ""} ${app.searchCapabilities} ${catalog.map(tool => `${tool.toolName} ${tool.description ?? ""}`).join(" ")}`);
|
||||
if (!score) continue;
|
||||
const ready = await usableConnectionForAgent({ companyId: run.companyId, agentId: agent.id,
|
||||
responsibleUserId: run.responsibleUserId!, serviceSlug: service, inventory });
|
||||
const denied = !ready && matching.length > 0 && await administrativeDenial(run.companyId, agent.id, service, inventory);
|
||||
candidates.push({ score, item: {
|
||||
responsibleUserId: run.responsibleUserId!, serviceSlug: service, purpose: aiOnly ? "ai" : undefined, inventory });
|
||||
const denied = !aiOnly && !ready && matching.length > 0 && await administrativeDenial(run.companyId, agent.id, service, inventory);
|
||||
candidates.push({ score, nameScore, item: {
|
||||
service, name: app.name, description: app.description ?? null, logoUrl: app.branding.logoUrl ?? null,
|
||||
methods: app.methods, source: app.source,
|
||||
state: ready ? "ready" : denied ? "unavailable" : !app.available || !app.methods.length ? "unavailable"
|
||||
@@ -394,67 +419,107 @@ export function connectionIntentService(db: Db) {
|
||||
connectionId: ready?.id ?? null,
|
||||
}});
|
||||
}
|
||||
const explicit = explicitAggregatorQuery(query);
|
||||
const serviceQuery = explicit?.serviceQuery ?? query;
|
||||
const publicService = findAggregatorService(serviceQuery);
|
||||
const exact = candidates.filter(({ item }) => normalizeConnectionQuery(item.service) === normalizeConnectionQuery(query)
|
||||
|| normalizeConnectionQuery(item.name) === normalizeConnectionQuery(query)
|
||||
|| item.service === publicService?.slug);
|
||||
const targetService = publicService?.slug ?? normalizeConnectionQuery(serviceQuery).replaceAll(" ", "-");
|
||||
const publicMatches = searchAggregatorServices(preparedQuery);
|
||||
const bestMatches = publicMatches.filter(match => Math.floor(match.nameScore / 100) === Math.floor((publicMatches[0]?.nameScore ?? 0) / 100));
|
||||
const publicService = findAggregatorService(serviceQuery, publicMatches) ?? (bestMatches.length === 1 ? bestMatches[0]!.service : undefined);
|
||||
// Extract service namespaces only from the current identity's indexed tools.
|
||||
// Generic provider search/execute descriptions do not prove app support.
|
||||
const indexedServices = new Set<string>();
|
||||
for (const connection of inventory.connections) {
|
||||
const provider = sourceSlugForConnection(connection, inventory.applicationsById);
|
||||
if (!isRemoteMcpConnectorId(provider) || !connection.enabled || connection.status === "archived") continue;
|
||||
for (const entry of authorizedCatalogs.get(connection.id) ?? []) {
|
||||
const namespace = entry.toolName.toLowerCase().match(/^([a-z0-9-]+)[_.:]/)?.[1];
|
||||
if (namespace && !isRemoteMcpConnectorId(namespace)) indexedServices.add(namespace);
|
||||
if (provider === "executor" && entry.toolName === "execute") {
|
||||
for (const line of (entry.description ?? "").split("\n")) {
|
||||
const slug = line.trim().match(/^- `([a-z0-9][a-z0-9-]{0,79})`$/)?.[1];
|
||||
if (slug) indexedServices.add(slug);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
const indexedMatches = [...indexedServices].map(slug => ({ slug, ...scoreConnectionSearch(preparedQuery, [slug]) }))
|
||||
.filter(match => match.nameScore > 0).sort((a, b) => b.score - a.score);
|
||||
const bestExternalTier = Math.floor(Math.max(publicMatches[0]?.nameScore ?? 0, indexedMatches[0]?.nameScore ?? 0) / 100);
|
||||
// Native preference applies to the same app or equally strong name matches.
|
||||
// A typo match such as Notion must not hide the distinctly named app Motion.
|
||||
const exact = candidates.filter(({ item, nameScore }) => ((nameScore >= 200 && Math.floor(nameScore / 100) >= bestExternalTier) || item.service === publicService?.slug)
|
||||
&& (!isRemoteMcpConnectorId(item.service) || (!publicService && !indexedMatches.length)));
|
||||
const ranked = () => candidates.sort((a, b) => b.score - a.score || a.item.name.localeCompare(b.item.name))
|
||||
.slice(0, 40).map(({ item }) => item);
|
||||
const targetService = publicService?.slug ?? (indexedMatches.length === 1 ? indexedMatches[0]!.slug : normalizeConnectionQuery(serviceQuery).replaceAll(" ", "-"));
|
||||
// Indexed-only app labels must be identical when requests re-search the slug.
|
||||
const targetName = publicService?.name ?? targetService.split("-").map(word => word.charAt(0).toUpperCase() + word.slice(1)).join(" ");
|
||||
const previous = await providerSelections(run.companyId, issue.id, agent.id, run.responsibleUserId!, `connection-provider:${targetService}`);
|
||||
const explicitConsent = explicit && await hasExplicitProviderRequest(run.companyId, issue.id, run.responsibleUserId!, targetService, explicit.provider, previous[0]);
|
||||
if (exact.length && (!explicitConsent || exact.some(({ item }) => item.state === "unavailable"))) return directSearchResult(query, exact.map(({ item }) => item));
|
||||
const alternatives: ConnectionSearchResultItem[] = [];
|
||||
if (/^[a-z0-9][a-z0-9-]{0,79}$/.test(targetService)) {
|
||||
for (const provider of AGGREGATOR_PRIORITY) {
|
||||
// Broad execute/search descriptions are not evidence of app support. Only a
|
||||
// namespaced action (or an explicitly listed Executor integration) qualifies.
|
||||
const connections = inventory.connections.filter(connection => connection.status !== "archived" && connection.enabled && sourceSlugForConnection(connection, inventory.applicationsById) === provider);
|
||||
let indexedAt: Date | undefined;
|
||||
for (const connection of connections) {
|
||||
const { grants } = await access.listConnectionGrants(connection.id, run.companyId);
|
||||
if (!grants.some(grant => grant.status === "active" && (grant.kind === "organization"
|
||||
|| grant.kind === "user" && grant.subjectUserId === run.responsibleUserId
|
||||
|| grant.kind === "agent" && grant.subjectAgentId === agent.id))) continue;
|
||||
const entries = await indexedCatalog(connection.id, run.companyId);
|
||||
const matching = entries.filter(entry => {
|
||||
const name = entry.toolName.toLowerCase();
|
||||
return name.startsWith(targetService + "_") || name.startsWith(targetService + ".") || name.startsWith(targetService + ":")
|
||||
|| provider === "executor" && entry.toolName === "execute"
|
||||
&& (entry.description ?? "").split("\n").some(line => line.trim() === "- `" + targetService + "`");
|
||||
});
|
||||
for (const entry of matching) if (!indexedAt || entry.lastSeenAt > indexedAt) indexedAt = entry.lastSeenAt;
|
||||
}
|
||||
const publishedAt = publicService ? (publicService.providers as Partial<Record<string, string>>)[provider] : undefined;
|
||||
const published = Boolean(publishedAt);
|
||||
if (!published && !indexedAt) continue;
|
||||
if (await administrativeDenial(run.companyId, agent.id, provider, inventory)) continue;
|
||||
const app = await resolveService(provider, run.companyId, run.responsibleUserId!, agent.id);
|
||||
if (!app.available || !app.methods.length) continue;
|
||||
const ready = await usableConnectionForAgent({ companyId: run.companyId, agentId: agent.id, responsibleUserId: run.responsibleUserId!, serviceSlug: provider, inventory });
|
||||
alternatives.push({
|
||||
service: `via:${provider}:${targetService}`, name: `${targetName} through ${app.name}`,
|
||||
source: "aggregator", state: "available", description: `Connect ${targetName} through ${app.name}, an external service.`,
|
||||
logoUrl: app.branding.logoUrl ?? null, methods: app.methods, connectionId: ready?.id ?? null,
|
||||
reason: ready ? `${app.name} is connected; verify ${targetName} authorization and the requested action.`
|
||||
: provider === "arcade" ? "Set up an Arcade gateway with this app's tools, then authorize the app."
|
||||
: `Connect ${app.name}, then verify and authorize ${targetName}.`,
|
||||
aggregator: { provider, targetService, targetName, evidenceUrl: published ? AGGREGATOR_CATALOG_SOURCES[provider] ?? null : null,
|
||||
verifiedAt: publishedAt ?? indexedAt!.toISOString(),
|
||||
readiness: ready ? "requires_app_verification" : "requires_provider_setup" },
|
||||
});
|
||||
if (exact.length && (!explicitConsent || exact.some(({ item }) => item.state === "unavailable"))) {
|
||||
const otherApps = bestMatches.filter(({ service, nameScore }) => nameScore >= 200 && !exact.some(({ item }) =>
|
||||
item.service === service.slug || scoreConnectionSearch(service.name, [item.name]).nameScore === 1000));
|
||||
if (otherApps.length && !exact.some(({ item }) => item.state === "unavailable")) {
|
||||
const otherRoutes = (await Promise.all(otherApps.slice(0, 10).map(({ service }) =>
|
||||
aggregatorAlternatives(service.slug, service.name, service)))).flat();
|
||||
if (otherRoutes.length) return discoverySuggestions([...ranked(), ...otherRoutes]);
|
||||
}
|
||||
return directSearchResult(query, ranked());
|
||||
}
|
||||
async function aggregatorAlternatives(targetService: string, targetName: string, publicService?: AggregatorServiceDefinition) {
|
||||
const alternatives: ConnectionSearchResultItem[] = [];
|
||||
if (/^[a-z0-9][a-z0-9-]{0,79}$/.test(targetService)) {
|
||||
for (const provider of AGGREGATOR_PRIORITY) {
|
||||
// Broad execute/search descriptions are not evidence of app support. Only a
|
||||
// namespaced action (or an explicitly listed Executor integration) qualifies.
|
||||
const connections = inventory.connections.filter(connection => connection.status !== "archived" && connection.enabled && sourceSlugForConnection(connection, inventory.applicationsById) === provider);
|
||||
let indexedAt: Date | undefined;
|
||||
for (const connection of connections) {
|
||||
const { grants } = await access.listConnectionGrants(connection.id, run.companyId);
|
||||
if (!grants.some(grant => grant.status === "active" && (grant.kind === "organization"
|
||||
|| grant.kind === "user" && grant.subjectUserId === run.responsibleUserId
|
||||
|| grant.kind === "agent" && grant.subjectAgentId === agent.id))) continue;
|
||||
const entries = authorizedCatalogs.get(connection.id) ?? [];
|
||||
const matching = entries.filter(entry => {
|
||||
const name = entry.toolName.toLowerCase();
|
||||
return name.startsWith(targetService + "_") || name.startsWith(targetService + ".") || name.startsWith(targetService + ":")
|
||||
|| provider === "executor" && entry.toolName === "execute"
|
||||
&& (entry.description ?? "").split("\n").some(line => line.trim() === "- `" + targetService + "`");
|
||||
});
|
||||
for (const entry of matching) if (!indexedAt || entry.lastSeenAt > indexedAt) indexedAt = entry.lastSeenAt;
|
||||
}
|
||||
const publishedAt = publicService ? (publicService.providers as Partial<Record<string, string>>)[provider] : undefined;
|
||||
const published = Boolean(publishedAt);
|
||||
if (!published && !indexedAt) continue;
|
||||
if (await administrativeDenial(run.companyId, agent.id, provider, inventory)) continue;
|
||||
const app = await resolveService(provider, run.companyId, run.responsibleUserId!, agent.id);
|
||||
if (!app.available || !app.methods.length) continue;
|
||||
const ready = await usableConnectionForAgent({ companyId: run.companyId, agentId: agent.id, responsibleUserId: run.responsibleUserId!, serviceSlug: provider, inventory });
|
||||
alternatives.push({
|
||||
service: `via:${provider}:${targetService}`, name: `${targetName} through ${app.name}`,
|
||||
source: "aggregator", state: "available", description: `Connect ${targetName} through ${app.name}, an external service.`,
|
||||
logoUrl: app.branding.logoUrl ?? null, methods: app.methods, connectionId: ready?.id ?? null,
|
||||
reason: ready ? `${app.name} is connected; verify ${targetName} authorization and the requested action.`
|
||||
: provider === "arcade" ? "Set up an Arcade gateway with this app's tools, then authorize the app."
|
||||
: `Connect ${app.name}, then verify and authorize ${targetName}.`,
|
||||
aggregator: { provider, targetService, targetName, evidenceUrl: published ? publicService?.evidenceUrls?.[provider] ?? AGGREGATOR_CATALOG_SOURCES[provider] ?? null : null,
|
||||
verifiedAt: publishedAt ?? indexedAt!.toISOString(),
|
||||
readiness: ready ? "requires_app_verification" : "requires_provider_setup" },
|
||||
});
|
||||
}
|
||||
}
|
||||
return alternatives;
|
||||
}
|
||||
const alternatives = await aggregatorAlternatives(targetService, targetName, publicService);
|
||||
if (explicit && explicitConsent) {
|
||||
const selected = alternatives.find(item => item.aggregator?.provider === explicit.provider);
|
||||
if (!selected) return { version: 1, query, results: [], instruction: "The explicitly requested external provider is unavailable or its app support could not be verified. Explain the limitation. Do not switch providers automatically." };
|
||||
return { version: 1, query, results: [{ ...selected, service: explicit.provider }],
|
||||
instruction: `The user explicitly requested ${AGGREGATOR_NAMES[explicit.provider]}. Disclose that this external service handles the connection and requests to ${targetName}. Do not ask another provider-choice question or substitute another provider. Call connection_request with service ${explicit.provider} and targetService ${targetService}. Follow its returned instruction; app authorization is not yet verified.` };
|
||||
}
|
||||
if (!alternatives.length) return directSearchResult(query, candidates.filter(({ item, score }) => !isRemoteMcpConnectorId(item.service) && score >= tokens.length)
|
||||
.sort((a, b) => b.score - a.score || a.item.name.localeCompare(b.item.name)).slice(0, 20).map(({ item }) => item), true);
|
||||
if (!alternatives.length && bestMatches.length > 1) {
|
||||
const matches = (await Promise.all(bestMatches.slice(0, 10).map(({ service }) =>
|
||||
aggregatorAlternatives(service.slug, service.name, service)))).flat();
|
||||
if (matches.length) return discoverySuggestions([...ranked(), ...matches]);
|
||||
}
|
||||
if (!alternatives.length) return directSearchResult(query, ranked(), true);
|
||||
const question = aggregatorProviderQuestion(targetService, targetName, alternatives);
|
||||
const latest = previous[0];
|
||||
if (latest && (latest.status === "pending" || !options.retryProviderChoice)) {
|
||||
@@ -472,6 +537,11 @@ export function connectionIntentService(db: Db) {
|
||||
if (!offered.length) return { version: 1, query, results: [], instruction: "The requested external provider is unavailable. Do not switch providers automatically." };
|
||||
return { version: 1, query, results: offered, providerQuestion: aggregatorProviderQuestion(targetService, targetName, offered),
|
||||
instruction: "No matching built-in Paperclip connection was found. Ask the responsible user with providerQuestion exactly as returned (including its id, full prompt, and options). With native request_human_input, use interactionKind questions, continuationPolicy wake_assignee, and payload {version:1, questions:[providerQuestion]}; do not add questionSet. Otherwise use ask_user_questions with the same questions payload. These are external services. Wait for the saved answer; then call connection_request with the selected service identifier and selectionInteractionId set to the answered question interaction ID. None for now means do not connect. Do not claim app access yet." };
|
||||
|
||||
function discoverySuggestions(results: ConnectionSearchResultItem[]): ConnectionsSearchResult {
|
||||
return { version: 1, query, results: results.slice(0, 40),
|
||||
instruction: "Multiple apps match this query. Choose the relevant service and method using descriptions and purposes. For available or needs_user_action results, tool methods use connection_request and channel or AI methods use setupPath. Use ready tools as installed; ready AI authentication applies to the next execution without reconnection. For an aggregator result, search its aggregator.targetService to obtain the provider-choice question and follow that instruction before requesting a connection. Respect unavailable states. Do not treat a search match as provider consent or app authorization." };
|
||||
}
|
||||
}
|
||||
|
||||
function directSearchResult(query: string, results: ConnectionSearchResultItem[], suggestions = false): ConnectionsSearchResult {
|
||||
@@ -479,7 +549,12 @@ export function connectionIntentService(db: Db) {
|
||||
? "No verified connection route was found. Explain that support could not be verified; do not invent a provider route or request unsupported services."
|
||||
: !suggestions && isRemoteMcpConnectorId(results[0]!.service) && results[0]!.state !== "unavailable"
|
||||
? `${results[0]!.name} is an external service. When the user explicitly names this provider, disclose that it handles the connection and requests to the requested app; no additional provider-choice question is necessary. ${results[0]!.state === "ready" ? aggregatorContinuationInstruction(results[0]!.service, "The requested app") : "Call connection_request with the returned service identifier and follow its instruction. Provider setup does not yet verify underlying app access."}`
|
||||
: suggestions ? "These are possible Paperclip matches, not an exact service match. Clarify the service if ambiguous, then search its exact name. Do not treat unrelated matches as support."
|
||||
: !suggestions && results[0]!.state === "unavailable" ? "This connection is unavailable or administratively restricted. Explain the reason. Do not bypass it using another provider."
|
||||
: suggestions || results.length > 1 ? "These are ranked connection matches. Choose the relevant service and method using their descriptions and purposes; extra query words need not match. For available or needs_user_action tool methods, call connection_request with the service identifier to present its setup card. For available or needs_user_action channel or AI methods, share the method's setupPath. Use ready tools as installed; ready AI authentication is available for the agent's next execution and does not need reconnection. Respect unavailable states and recorded user choices. Clarify only if the intended service is still ambiguous; unrelated matches are not evidence of support."
|
||||
: results[0]!.state === "ready" && results[0]!.methods.every(method => method.purpose === "ai")
|
||||
? "AI authentication is available for this agent's next execution. Do not create another connection request or ask the user to reconnect."
|
||||
: results[0]!.methods.every(method => method.purpose && method.purpose !== "tool")
|
||||
? "Choose the method relevant to the task using its purpose and label. Share its setupPath with the user to open the existing channel or AI setup flow. These methods do not use the tool connection_request card. Do not claim tools or an inbox are ready before setup finishes."
|
||||
: results[0]!.state === "ready" ? "Use the installed connection. Do not create another connection request."
|
||||
: results[0]!.state === "unavailable" ? "This connection is unavailable or administratively restricted. Explain the reason. Do not bypass it using another provider."
|
||||
: "Call connection_request with the returned service identifier. The user already asked to connect: do not ask a generic confirmation or imitate the setup card. Follow the returned instruction." };
|
||||
|
||||
@@ -6,7 +6,7 @@ export const RUNTIME_CONNECTION_TOOL_DEFINITIONS = [
|
||||
description: CONNECTIONS_SEARCH_TOOL_DESCRIPTION,
|
||||
inputSchema: {
|
||||
type: "object",
|
||||
properties: { query: { type: "string" }, retryProviderChoice: { type: "boolean", description: "Only when the user explicitly asks to reconsider a previous provider choice or decline" } },
|
||||
properties: { query: { type: "string", maxLength: 4000 }, retryProviderChoice: { type: "boolean", description: "Only when the user explicitly asks to reconsider a previous provider choice or decline" } },
|
||||
additionalProperties: false,
|
||||
},
|
||||
},
|
||||
@@ -21,4 +21,3 @@ export const RUNTIME_CONNECTION_TOOL_DEFINITIONS = [
|
||||
},
|
||||
},
|
||||
] as const;
|
||||
|
||||
|
||||
Reference in New Issue
Block a user