mirror of
https://github.com/mksglu/context-mode.git
synced 2026-10-02 04:14:38 +08:00
feat(prose): retire prose-style enforcement entirely (#482)
Issue #482 (makoMakoGo) reported that context-mode's caveman/terse injection pressures the model toward brevity on its FINAL ANSWER, not just on tool-output reporting. Cited evidence: Moonshot AI on kimi-k2.5 (anomalyco/opencode#20258, PR #20259) — aggressive brevity prompts measurably degrade coding/reasoning benchmarks because the model drops assumptions, caveats, verification evidence, failure modes, and security warnings the user actually needs. Considered: A — config switch ("injectCommunicationStyle: false"). Rejected: switches default-on become dead code. B — close FR with rationale. Rejected: ignores valid evidence. C — refine wording ("compress when reporting raw tool output, be complete for technical answers"). Rejected: still text injection, model-dependent, half-measure. D — full strip everywhere. Adopted. The decision after grilling: context-mode's value is data routing (sandbox, FTS5, session continuity), not prose styling. The brevity injection conflated three goals — keeping raw data out of context (real, hard-enforced), summarizing tool output compactly (LLMs auto- calibrate), and final-answer prose style (the wrong target). Strip all 22 sites where prose-style language landed. Sites stripped (A-Z): hooks/routing-block.mjs - <communication_style> block (Terse like caveman, fragments OK, auto-expand for security warnings) - <response_format> block (Concise summary, 2-3 bullets) src/server.ts (5 MCP tool descriptions + 2 cosmetic comments) - ctx_execute "When reporting results — terse..." - ctx_execute_file same - ctx_search same - ctx_fetch_and_index same (URL + commands shapes, both) - cosmetic comment "Caveman style — terse status line" - rewrote concurrency note: "Indexing is serial regardless of concurrency" → "Fetches parallelize up to your concurrency setting; FTS5 indexing serializes the writes after (SQLite single-writer rule)." — same fact, less jargon. configs/ (15 adapter MD files — every shipped system prompt) antigravity/GEMINI.md, claude-code/CLAUDE.md, codex/AGENTS.md, cursor/context-mode.mdc, gemini-cli/GEMINI.md, jetbrains-copilot/ copilot-instructions.md, kilo/AGENTS.md, kiro/KIRO.md, omp/SYSTEM.md, openclaw/AGENTS.md, opencode/AGENTS.md, pi/AGENTS.md, qwen-code/ QWEN.md, vscode-copilot/copilot-instructions.md, zed/AGENTS.md All had identical "## Output" block: 3 caveman lines stripped, workflow lines ("Write artifacts to FILES", "Descriptive source labels") kept. CLAUDE.md (repo root — internal dev instructions) Same caveman block stripped. We don't ship this file but we do eat our own dog food. README.md Pillar 4 ("Output Compression — Terse like caveman...") rewritten to "No prose-style enforcement" — explicitly cites the kimi-k2.5 benchmark evidence as the rationale. web/index.html Removed Ch 4b entirely (the "Output compression" chapter with before/after example pushing terse style on the model). Tests: - tests/session/continuity.test.ts: SessionStart routing-block assertion flipped from "must include 'Terse like caveman'" to "must NOT include caveman/terse-style directive". - tests/core/server.test.ts: Task hook injection assertion same flip. Two cosmetic comment renames ("Caveman style — terse status line" → "Status line: counts + sections + size"), test name rename ("caveman style" → "compact format"). Added new "prose-style policy (#482)" describe block at end of file with 3 negative-pin tests covering server.ts MCP descriptions, routing-block, and README. Full suite: 82/82 files passed, 2645 passed, 20 skipped, 0 failed. Net +3 new tests (the policy describe block). CONTRIBUTING.md New "Prose-style policy (#482)" section documents the decision so future contributors don't re-add the injection. Cites the Moonshot benchmark evidence + the regression test that pins the deletion. This addresses #482 in full. Closing the issue with a comment that walks the requester through the decision and links the policy section.
This commit is contained in:
@@ -57,9 +57,6 @@ Routing block auto-injected into subagent prompts. Bash-type subagents upgraded
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -339,6 +339,23 @@ To test against a running OpenClaw gateway:
|
||||
|
||||
See [`docs/adapters/openclaw.md`](docs/adapters/openclaw.md) for hook registration details and known upstream issues.
|
||||
|
||||
## Prose-style policy (issue [#482](https://github.com/mksglu/context-mode/issues/482))
|
||||
|
||||
context-mode does not dictate how the model writes its final answer. The four pillars (sandbox routing, session continuity, think-in-code, no prose-style enforcement) keep raw data out of context but leave editorial style — brevity vs. completeness, formatting, tone — entirely to the model and the user's own `CLAUDE.md` / `AGENTS.md`.
|
||||
|
||||
**Why:** aggressive brevity instructions have been shown to degrade coding/reasoning benchmarks. Moonshot AI's report on `kimi-k2.5` (cited in [#482](https://github.com/mksglu/context-mode/issues/482), with the OpenCode fix at [anomalyco/opencode#20259](https://github.com/anomalyco/opencode/pull/20259)) showed that prompts like "minimize output tokens", "MUST answer concisely with fewer than 4 lines", and "one-word answers are best" pushed coding models to drop assumptions, caveats, verification evidence, failure modes, and security warnings the user actually needed.
|
||||
|
||||
**What this means for contributors:**
|
||||
|
||||
- Do **not** add brevity directives to MCP tool descriptions in `src/server.ts`.
|
||||
- Do **not** add `<communication_style>` or `<response_format>` blocks to `hooks/routing-block.mjs`.
|
||||
- Do **not** put "Terse like caveman" / "Only fluff die" / "Drop articles, filler" / "fewer than N lines" wording in any shipped adapter config under `configs/*/`.
|
||||
- Workflow-discipline rules — "write artifacts to FILES", "use descriptive `ctx_search` source labels", `<artifact_policy>` — are fine. They describe *what to do* (file vs. inline), not *how to write*.
|
||||
|
||||
The regression test at `tests/core/server.test.ts > prose-style policy (#482)` pins the deletion: any caveman-style language landing in `src/server.ts`, `hooks/routing-block.mjs`, or `README.md` will fail CI.
|
||||
|
||||
If you genuinely need to nudge the model on style for a specific use case, do it in your own project's `CLAUDE.md` / `AGENTS.md`. Don't ship it inside the framework.
|
||||
|
||||
## Submitting a Bug Report
|
||||
|
||||
When filing a bug, **always include your prompt**. The exact message you sent to the agent is critical for reproduction. Without it, we can't debug the issue.
|
||||
|
||||
@@ -48,7 +48,7 @@ Context Mode is an MCP server that solves all four sides of this problem:
|
||||
files.forEach(f => console.log(f + ': ' + fs.readFileSync('src/'+f,'utf8').split('\\n').length + ' lines'));
|
||||
`);
|
||||
```
|
||||
4. **Output Compression** — Terse like caveman. Technical substance exact. Only fluff die. Drop articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged. Pattern: [thing] [action] [reason]. [next step]. Auto-expand for security warnings, irreversible actions, and user confusion. ~65-75% output token reduction with full technical accuracy.
|
||||
4. **No prose-style enforcement** — context-mode keeps raw data out of context but never dictates how the model writes its final answer. Brevity, completeness, formatting — your model's call (or yours via your own `CLAUDE.md` / `AGENTS.md`). Aggressive brevity prompts have been shown to degrade coding/reasoning benchmarks ([Moonshot AI on `kimi-k2.5`](https://github.com/anomalyco/opencode/issues/20258)) — the routing block stays focused on *where data goes*, not on *how the model talks*.
|
||||
|
||||
<a href="https://www.youtube.com/watch?v=QUHrntlfPo4">
|
||||
<picture>
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -57,9 +57,6 @@ Routing block auto-injected into subagent prompts. Bash-type subagents upgraded
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -54,9 +54,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -46,9 +46,6 @@ ALWAYS use native file editing tools to create/modify files. NEVER use `ctx_exec
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
|
||||
## Session Continuity
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -45,9 +45,6 @@ Pass `concurrency: 4-8` to `ctx_batch_execute` and `ctx_fetch_and_index` for net
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `context-mode_ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -54,9 +54,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `context-mode__ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -54,9 +54,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -57,9 +57,6 @@ Routing block auto-injected into subagent prompts. Bash-type subagents upgraded
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `mcp__context-mode__ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -45,9 +45,6 @@ Pass `concurrency: 4-8` to `ctx_batch_execute` and `ctx_fetch_and_index` for net
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `ctx_search(source: "label")`.
|
||||
|
||||
|
||||
@@ -53,9 +53,6 @@ GitHub API rate-limit: cap at 4 for `gh` calls.
|
||||
|
||||
## Output
|
||||
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Drop: articles, filler (just/really/basically), pleasantries, hedging. Fragments OK. Short synonyms. Code unchanged.
|
||||
Pattern: [thing] [action] [reason]. [next step]. Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
Write artifacts to FILES — never inline. Return: file path + 1-line description.
|
||||
Descriptive source labels for `search(source: "label")`.
|
||||
|
||||
|
||||
@@ -50,22 +50,10 @@ export function createRoutingBlock(t, options = {}) {
|
||||
</file_writing_policy>
|
||||
|
||||
<output_constraints>
|
||||
<communication_style>
|
||||
Terse like caveman. Technical substance exact. Only fluff die.
|
||||
Use fragments when clear. Short synonyms (fix not "implement a solution for").
|
||||
Technical terms exact. Code blocks unchanged.
|
||||
Auto-expand for: security warnings, irreversible actions, user confusion.
|
||||
</communication_style>
|
||||
<artifact_policy>
|
||||
Write artifacts (code, configs, PRDs) to FILES. NEVER inline.
|
||||
Return only: file path + 1-line description.
|
||||
</artifact_policy>
|
||||
<response_format>
|
||||
Concise summary:
|
||||
- Actions taken (2-3 bullets)
|
||||
- File paths created/modified
|
||||
- Key findings
|
||||
</response_format>
|
||||
</output_constraints>
|
||||
<session_continuity>
|
||||
Skills, roles, and decisions set during this session remain active until the user revokes them.
|
||||
|
||||
+7
-10
@@ -1007,7 +1007,7 @@ server.registerTool(
|
||||
"ctx_execute",
|
||||
{
|
||||
title: "Execute Code",
|
||||
description: `MANDATORY: Use for any command where output exceeds 20 lines. Execute code in a sandboxed subprocess. Only stdout enters context — raw data stays in the subprocess.${bunNote} Available: ${langList}.\n\nPREFER THIS OVER BASH for: API calls (gh, curl, aws), test runners (npm test, pytest), git queries (git log, git diff), data processing, and ANY CLI command that may produce large output. Bash should only be used for file mutations, git writes, and navigation.\n\nTHINK IN CODE: When you need to analyze, count, filter, compare, or process data — write code that does the work and console.log() only the answer. Do NOT read raw data into context to process mentally. Program the analysis, don't compute it in your reasoning. Write robust, pure JavaScript (no npm dependencies). Use only Node.js built-ins (fs, path, child_process). Always wrap in try/catch. Handle null/undefined. Works on both Node.js and Bun.\n\nWhen reporting results — terse like caveman. Technical substance exact. Only fluff die. Pattern: [thing] [action] [reason]. [next step].`,
|
||||
description: `MANDATORY: Use for any command where output exceeds 20 lines. Execute code in a sandboxed subprocess. Only stdout enters context — raw data stays in the subprocess.${bunNote} Available: ${langList}.\n\nPREFER THIS OVER BASH for: API calls (gh, curl, aws), test runners (npm test, pytest), git queries (git log, git diff), data processing, and ANY CLI command that may produce large output. Bash should only be used for file mutations, git writes, and navigation.\n\nTHINK IN CODE: When you need to analyze, count, filter, compare, or process data — write code that does the work and console.log() only the answer. Do NOT read raw data into context to process mentally. Program the analysis, don't compute it in your reasoning. Write robust, pure JavaScript (no npm dependencies). Use only Node.js built-ins (fs, path, child_process). Always wrap in try/catch. Handle null/undefined. Works on both Node.js and Bun.`,
|
||||
inputSchema: z.object({
|
||||
language: z
|
||||
.enum([
|
||||
@@ -1337,7 +1337,7 @@ server.registerTool(
|
||||
{
|
||||
title: "Execute File Processing",
|
||||
description:
|
||||
"Read a file and process it without loading contents into context. The file is read into a FILE_CONTENT variable inside the sandbox. Only your printed summary enters context.\n\nPREFER THIS OVER Read/cat for: log files, data files (CSV, JSON, XML), large source files for analysis, and any file where you need to extract specific information rather than read the entire content.\n\nTHINK IN CODE: Write code that processes FILE_CONTENT and console.log() only the answer. Don't read files into context to analyze mentally. Write robust, pure JavaScript — no npm deps, try/catch, null-safe. Node.js + Bun compatible.\n\nWhen reporting results — terse like caveman. Technical substance exact. Only fluff die. Pattern: [thing] [action] [reason]. [next step].",
|
||||
"Read a file and process it without loading contents into context. The file is read into a FILE_CONTENT variable inside the sandbox. Only your printed summary enters context.\n\nPREFER THIS OVER Read/cat for: log files, data files (CSV, JSON, XML), large source files for analysis, and any file where you need to extract specific information rather than read the entire content.\n\nTHINK IN CODE: Write code that processes FILE_CONTENT and console.log() only the answer. Don't read files into context to analyze mentally. Write robust, pure JavaScript — no npm deps, try/catch, null-safe. Node.js + Bun compatible.",
|
||||
inputSchema: z.object({
|
||||
path: z
|
||||
.string()
|
||||
@@ -1621,8 +1621,7 @@ server.registerTool(
|
||||
"Pass ALL search questions as queries array in ONE call. " +
|
||||
"File-backed sources are auto-refreshed when the source file changes.\n\n" +
|
||||
"TIPS: 2-4 specific terms per query. Use 'source' to scope results.\n\n" +
|
||||
"SESSION STATE: If skills, roles, or decisions were set earlier in this conversation, they are still active. Do not discard or contradict them.\n\n" +
|
||||
"When reporting results — terse like caveman. Technical substance exact. Only fluff die. Pattern: [thing] [action] [reason]. [next step].",
|
||||
"SESSION STATE: If skills, roles, or decisions were set earlier in this conversation, they are still active. Do not discard or contradict them.",
|
||||
inputSchema: z.object({
|
||||
queries: z.preprocess(coerceJsonArray, z
|
||||
.array(z.string())
|
||||
@@ -2356,8 +2355,7 @@ server.registerTool(
|
||||
" ✅ Use concurrency: 4-8 for: library docs sweep, multi-changelog scan, competitive pricing pages, multi-region docs, GitHub raw file pulls.\n" +
|
||||
" ❌ Single URL → use the legacy {url, source} shape (concurrency irrelevant).\n" +
|
||||
" Example: requests: [{url: 'https://react.dev/...', source: 'react'}, {url: 'https://vuejs.org/...', source: 'vue'}], concurrency: 5.\n" +
|
||||
" Indexing is serial regardless of concurrency — fetches race, FTS5 writes don't (avoids SQLite WAL contention).\n\n" +
|
||||
"When reporting results — terse like caveman. Technical substance exact. Only fluff die. Pattern: [thing] [action] [reason]. [next step].",
|
||||
" Fetches parallelize up to your concurrency setting; FTS5 indexing serializes the writes after (SQLite single-writer rule).",
|
||||
inputSchema: z.object({
|
||||
url: z.string().optional().describe("Single URL to fetch and index (legacy single-shape)"),
|
||||
source: z
|
||||
@@ -2537,8 +2535,8 @@ server.registerTool(
|
||||
const cappedNote = capped
|
||||
? ` cap=${effectiveConcurrency}/${cpus().length}cpu`
|
||||
: "";
|
||||
// Caveman style — terse status line: counts + sections + size.
|
||||
// Singular forms used at count=1 to avoid grammar drift ("1 errors" → "1 error").
|
||||
// Status line: counts + sections + size, with singular/plural agreement
|
||||
// (count=1 → "1 error" not "1 errors") so the line stays grammatical.
|
||||
const fmt = (n: number, sing: string, plur: string) => `${n} ${n === 1 ? sing : plur}`;
|
||||
const headerLine =
|
||||
`fetched ${batch.length} c=${effectiveConcurrency}${cappedNote}. ` +
|
||||
@@ -2580,8 +2578,7 @@ server.registerTool(
|
||||
" ❌ Keep concurrency: 1 for: npm test, build, lint, image processing (CPU-bound), or commands sharing state (ports, lock files, same-repo writes).\n" +
|
||||
" Example: [gh issue view 1, gh issue view 2, gh issue view 3] → concurrency: 3.\n" +
|
||||
" Speedup depends on workload — applies to I/O wait, not CPU work.\n\n" +
|
||||
"THINK IN CODE — NON-NEGOTIABLE: When commands produce data you need to analyze, count, filter, compare, or transform — add a processing command that runs JavaScript and console.log() ONLY the answer. NEVER pull raw output into context to reason over. Concurrency parallelizes the FETCH; THINK IN CODE owns the PROCESSING. One programmed analysis replaces ten read-and-reason rounds. Pure JavaScript, Node.js built-ins (fs, path, child_process), try/catch, null-safe.\n\n" +
|
||||
"When reporting results — terse like caveman. Technical substance exact. Only fluff die. Pattern: [thing] [action] [reason]. [next step].",
|
||||
"THINK IN CODE — NON-NEGOTIABLE: When commands produce data you need to analyze, count, filter, compare, or transform — add a processing command that runs JavaScript and console.log() ONLY the answer. NEVER pull raw output into context to reason over. Concurrency parallelizes the FETCH; THINK IN CODE owns the PROCESSING. One programmed analysis replaces ten read-and-reason rounds. Pure JavaScript, Node.js built-ins (fs, path, child_process), try/catch, null-safe.",
|
||||
inputSchema: z.object({
|
||||
commands: z.preprocess(coerceCommandsArray, z
|
||||
.array(
|
||||
|
||||
@@ -1842,7 +1842,16 @@ describe("Hook Injection", () => {
|
||||
const parsed = JSON.parse(output);
|
||||
const prompt = parsed.hookSpecificOutput.updatedInput.prompt;
|
||||
assert.ok(prompt.includes("<output_constraints>"), "Should inject output_constraints");
|
||||
assert.ok(prompt.includes("Terse like caveman"), "Should mention concise communication style");
|
||||
// Pillar 4 (caveman/Output Compression) retired in #482. Routing block
|
||||
// must NOT push a prose-style directive — assert the negative.
|
||||
assert.ok(
|
||||
!prompt.toLowerCase().includes("terse like caveman"),
|
||||
"Routing block must not contain caveman/terse style directive",
|
||||
);
|
||||
assert.ok(
|
||||
!prompt.includes("<communication_style>"),
|
||||
"Routing block must not contain communication_style block",
|
||||
);
|
||||
assert.ok(
|
||||
prompt.includes("<tool_selection_hierarchy>"),
|
||||
"Should inject tool_selection_hierarchy",
|
||||
@@ -3135,7 +3144,7 @@ describe("ctx_fetch_and_index batch refactor", () => {
|
||||
|
||||
test("capped-concurrency note appears only when capped", () => {
|
||||
expect(fetchHandlerSrc).toMatch(/cappedNote\s*=\s*capped\s*\?/);
|
||||
// Caveman style — `cap=N/Mcpu` instead of "capped from N to M; M cores available".
|
||||
// Compact form `cap=N/Mcpu` (replaces verbose "capped from N to M; M cores available").
|
||||
expect(fetchHandlerSrc).toContain("cap=${effectiveConcurrency}/${cpus().length}cpu");
|
||||
});
|
||||
|
||||
@@ -3156,13 +3165,13 @@ describe("ctx_fetch_and_index batch refactor", () => {
|
||||
});
|
||||
|
||||
test("batch header uses singular form for count=1 (review F5 plural fix)", () => {
|
||||
// Per CLAUDE.md "Terse like caveman" + grammar correctness:
|
||||
// "1 errors" → "1 error" via the fmt() helper.
|
||||
// Grammar correctness in the compact status line: "1 errors" → "1 error"
|
||||
// via the fmt() helper, so the line stays grammatical at any count.
|
||||
expect(fetchHandlerSrc).toContain('const fmt = (n: number, sing: string, plur: string)');
|
||||
expect(fetchHandlerSrc).toContain('n === 1 ? sing : plur');
|
||||
});
|
||||
|
||||
test("batch header uses caveman style (review F5 terse format)", () => {
|
||||
test("batch header uses compact format (review F5)", () => {
|
||||
// Old: "Batch fetched N URLs at concurrency=X (capped from Y to X; Z cores available): a fetched, b cached, c errors. d new sections (eKB total)."
|
||||
// New: "fetched N c=X cap=X/Zcpu. ok=a cache=b err=c. d sections eKB."
|
||||
expect(fetchHandlerSrc).toContain("`fetched ${batch.length} c=${effectiveConcurrency}");
|
||||
@@ -4241,3 +4250,42 @@ describe("killProcessOnPort — Windows non-English locale (#441 follow-up)", ()
|
||||
expect(calls).toHaveLength(1); // netstat only — no taskkill issued
|
||||
});
|
||||
});
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════════════════
|
||||
// Prose-style policy (issue #482)
|
||||
// ═══════════════════════════════════════════════════════════════════════════
|
||||
// Decision: context-mode keeps raw data out of context but does not dictate
|
||||
// how the model writes its final answer. Aggressive brevity prompts have been
|
||||
// shown to degrade coding/reasoning benchmarks (Moonshot AI on kimi-k2.5).
|
||||
// Any caveman/terse-style language in shipped artifacts is a regression.
|
||||
|
||||
describe("prose-style policy (#482)", () => {
|
||||
const serverSrc = readFileSync(
|
||||
resolve(__dirname, "../../src/server.ts"),
|
||||
"utf-8",
|
||||
);
|
||||
const routingBlock = readFileSync(
|
||||
resolve(__dirname, "../../hooks/routing-block.mjs"),
|
||||
"utf-8",
|
||||
);
|
||||
|
||||
test("no caveman/terse directive lands in any MCP tool description", () => {
|
||||
expect(serverSrc).not.toMatch(/terse like caveman/i);
|
||||
expect(serverSrc).not.toMatch(/only fluff die/i);
|
||||
});
|
||||
|
||||
test("routing-block has no <communication_style> or <response_format> blocks", () => {
|
||||
expect(routingBlock).not.toMatch(/<communication_style>/);
|
||||
expect(routingBlock).not.toMatch(/<response_format>/);
|
||||
expect(routingBlock).not.toMatch(/terse like caveman/i);
|
||||
});
|
||||
|
||||
test("README does not advertise an Output Compression pillar", () => {
|
||||
const readme = readFileSync(
|
||||
resolve(__dirname, "../../README.md"),
|
||||
"utf-8",
|
||||
);
|
||||
expect(readme).not.toMatch(/\*\*Output Compression\*\*/);
|
||||
expect(readme).not.toMatch(/terse like caveman/i);
|
||||
});
|
||||
});
|
||||
|
||||
@@ -102,11 +102,21 @@ describe("SessionStart Hook", () => {
|
||||
const result = runHook({});
|
||||
const parsed = JSON.parse(result.stdout);
|
||||
const ctx = parsed.hookSpecificOutput.additionalContext;
|
||||
assert.ok(ctx.includes("Terse like caveman"), "Expected communication style directive");
|
||||
// Pillar 4 ("Output Compression"/caveman) was retired in #482 — the
|
||||
// routing block must keep the artifact-policy block but MUST NOT push a
|
||||
// prose-style directive on the model.
|
||||
assert.ok(
|
||||
ctx.includes("Write artifacts"),
|
||||
"Expected artifact policy",
|
||||
);
|
||||
assert.ok(
|
||||
!ctx.toLowerCase().includes("terse like caveman"),
|
||||
"Routing block must not contain caveman/terse style directive",
|
||||
);
|
||||
assert.ok(
|
||||
!ctx.includes("<communication_style>"),
|
||||
"Routing block must not contain communication_style block",
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
+3
-28
@@ -297,34 +297,9 @@ code{font-family:var(--mono);font-size:.88em;background:var(--bg-code);color:var
|
||||
|
||||
<hr class="sep">
|
||||
|
||||
<!-- Ch 4b -->
|
||||
<section class="chapter rev">
|
||||
<p class="section-label">Output</p>
|
||||
<h2 class="section-title">Output compression</h2>
|
||||
<div class="rule"></div>
|
||||
|
||||
<p class="body-text">Context saving handles the input side. But the agent wastes output tokens too — filler words, pleasantries, verbose explanations, hedging. "Sure! I'd be happy to help you with that. Let me take a look at the file and see what we can do." That's 30 tokens of nothing.</p>
|
||||
|
||||
<p class="body-text">context-mode enforces output compression across all 15 platforms. Technical substance stays exact. Only fluff dies.</p>
|
||||
|
||||
<div style="background:var(--bg-alt);border:1px solid var(--border-light);border-radius:6px;padding:22px 24px;margin:1.5rem 0">
|
||||
<p style="font-size:15px;color:var(--text);font-weight:600;margin-bottom:10px">The rules</p>
|
||||
<p style="font-size:14px;color:var(--text-2);line-height:1.8;margin:0">Drop: articles, filler (just/really/basically), pleasantries, hedging.<br>Fragments OK. Short synonyms. Code unchanged.<br>Pattern: <code>[thing] [action] [reason]. [next step].</code><br>Auto-expand for: security warnings, irreversible actions, user confusion.</p>
|
||||
</div>
|
||||
|
||||
<div style="display:grid;grid-template-columns:1fr 1fr;gap:14px;margin:1.5rem 0">
|
||||
<div style="padding:16px 18px;border:1px solid var(--border-light);border-radius:6px">
|
||||
<p style="font-size:11px;font-weight:600;color:var(--text-4);letter-spacing:.06em;text-transform:uppercase;margin-bottom:8px">Before</p>
|
||||
<p style="font-size:13.5px;color:var(--text-3);line-height:1.6;margin:0;font-style:italic">"Sure, I'd be happy to help! I've looked at the database connection pooling configuration and it appears that the issue is related to the maximum idle timeout setting being set too low for your workload. Let me walk you through the fix."</p>
|
||||
</div>
|
||||
<div style="padding:16px 18px;border:1px solid var(--accent);border-radius:6px;background:var(--accent-light)">
|
||||
<p style="font-size:11px;font-weight:600;color:var(--accent-dark);letter-spacing:.06em;text-transform:uppercase;margin-bottom:8px">After</p>
|
||||
<p style="font-size:13.5px;color:var(--text);line-height:1.6;margin:0">"DB pool idle timeout too low for workload. Set <code>max_idle_ms: 30000</code> in config. Fix:"</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<p class="body-text">~65-75% output token reduction. Full technical accuracy. The agent says the same thing in fewer words. Your context window lasts longer from both sides.</p>
|
||||
</section>
|
||||
<!-- Ch 4b removed: prose-style enforcement was retired (issue #482).
|
||||
context-mode now keeps raw data out of context but does not dictate
|
||||
how the model writes its final answer. -->
|
||||
|
||||
<hr class="sep">
|
||||
|
||||
|
||||
Reference in New Issue
Block a user