mirror of
https://github.com/Imbad0202/academic-research-skills.git
synced 2026-10-02 07:14:34 +08:00
Whether to run a systematic review is the author's decision. ARS now shows one fixed note listing five review forms (systematic, scoping, narrative or integrative, rapid, no formal review), each with one line on what it serves and roughly what it costs, with no default, ranking, or recommendation. - shared/references/review_form_note.md holds the rules and the note text in English and Traditional Chinese; every other language gets the English text. deep-research/SKILL.md and academic-paper/SKILL.md carry the block verbatim. - The note appears at the first of two author actions: selecting lit-review mode, or confirming the research question (full mode's RQ Brief before Phase 2, or the Socratic closing RQ Brief or RQ Summary). It never depends on the question's content, is skipped when the author already named a form, and appears at most once per project or run. The session waits for the reply; skipping counts as a decision. With a passport file the note and the reply go to the run ledger as a review-form-note checkpoint, and an unanswered checkpoint is resumed rather than reopened. - Paths that could move into systematic-review mode on the model's reading now require the author's choice (three-way-scan escalation, the lit-review -> systematic-review transition, lit-review trigger examples), and requests to conduct a systematic review no longer route to academic-paper lit-review. - scripts/check_review_form_note_sync.py keeps the copies byte-identical and checks each note's five labels, order, neutrality sentence, quoting, ranking words, and selection marks; 34 mutation tests; wired into spec-consistency and the pytest manifest. - Routing fixtures 13-16 add the negative case, its content-independence pair, a named-form case, and a broader-synthesis case; 02 and 04 now expect the note. The fixtures have not been run; session behaviour is prompt-level and unmeasured. Cost lines are sourced in the canonical file (Borah et al. 2017; Tricco et al. 2018; Garritty et al. 2021; Whittemore & Knafl 2005). Cross-model review (gpt-6-astra, xhigh) converged in three rounds: 9 -> 2 -> 0 P1/P2. Closes #921 Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
8120a34368
commit
c5c1b454cc
@@ -407,6 +407,13 @@ jobs:
|
||||
PYTHONPATH: .
|
||||
run: python scripts/check_method_weaknesses_sync.py
|
||||
|
||||
- name: Review-form note sync lint (#921)
|
||||
# Every surface listed in shared/references/review_form_note.md must
|
||||
# carry its block byte for byte, and the note's options stay unranked.
|
||||
env:
|
||||
PYTHONPATH: .
|
||||
run: python scripts/check_review_form_note_sync.py
|
||||
|
||||
- name: Judge-prompt-version drift guard (#361)
|
||||
# Recomputes the SHA-256 of the canonical judge-prompt section (between
|
||||
# the JUDGE-PROMPT-CANONICAL markers in claim_ref_alignment_audit_agent.md)
|
||||
|
||||
@@ -10,6 +10,8 @@ All notable changes to this project will be documented in this file.
|
||||
|
||||
- **Reading outputs state how one source's method can fail, and say where each weakness came from (#916).** `shared/references/per_source_method_weaknesses.md` holds five rules, copied verbatim into the evidence assessment card (§4), the three-way scan WHAT field, and the annotated-bibliography formats of `bibliography_agent` and `literature_strategist_agent`: name the design choice and the condition under which it distorts the result; label each item `author-acknowledged` (with a v3.7.3 anchor) or `reader-inferred`; name the sections checked before saying the authors do not address something (#548 pattern); walk the paper-type table in `review_criteria_framework.md` §2 once instead of expanding open-endedly; and write `not assessed (read scope: <scope>)` when only the abstract or table of contents was read (#513 scopes). The format borrows from ScientistTwo (Nam et al., 2026, arXiv:2609.19644, Appendix C) and leaves out its unlabelled inferences and its research-direction field. Nothing blocks, gates, or asks the scholar a question. `scripts/check_method_weaknesses_sync.py` keeps the four copies byte-identical to the canonical block. `evals/heldout/per_source_method_weaknesses/` adds two synthetic items (a section-addressed issue and an abstract-only source); they are a smoke check, `NOT_RUN` as a measured suite, and no claim of improved review quality is made.
|
||||
|
||||
- **A fixed-point reminder that the author decides the form of a literature review (#921).** `shared/references/review_form_note.md` holds one note, in English and Traditional Chinese, that lists five forms (systematic, scoping, narrative or integrative, rapid, no formal review) with one line each on what the form serves and roughly what it costs, and states that ARS does not choose. The note has no default, ranking, or recommendation. It appears at the first of two actions the author takes: selecting `lit-review` mode, or confirming the research question (the RQ Brief in `deep-research` `full` mode before Phase 2, or the closing RQ Brief or RQ Summary of `socratic` mode; an ending the author has not confirmed does not count). It never appears because of what the question asks, and not at all when the author has already named a review form. The session then waits for the author's reply; skipping counts as a decision, and the note appears at most once per project or run. A systematic review is entered only on the author's explicit choice. In a run with a passport file the note and the reply go to the run ledger as a `review-form-note` checkpoint. `deep-research/SKILL.md` and `academic-paper/SKILL.md` carry the block verbatim; `scripts/check_review_form_note_sync.py` keeps the copies identical and checks that each note lists the five forms under their fixed labels in order and keeps its neutrality sentence, with no ranking word or selection mark elsewhere in the note. The two existing paths that could move into `systematic-review` mode on the model's own reading (a three-way scan escalating for "PRISMA-like coverage", and a `lit-review` → `systematic-review` transition "when the topic warrants" PRISMA) now require the author's choice, and requests for a systematic review no longer route to `academic-paper` `lit-review`. Routing fixtures 13 to 16 add a negative case ("help me write the literature review" reaches neither `systematic-review` mode nor screening), its content-independence pair, a named-form case, and a broader-synthesis case; fixtures 02 and 04 now expect the note. The fixtures have not been run; whether sessions show the note at the right points is prompt-level and unmeasured. Other languages get the English text, so no user sees a paraphrase written on the spot.
|
||||
|
||||
- **Held-out cases: an abstract that claims more than its own result tables (#915).** `evals/heldout/experiment_alignment_overclaim/` holds two synthetic Material Passports built from the DynaSpec-RAG showcase manuscript in ScientistTwo (Nam et al., 2026, arXiv:2609.19644 v1, Appendix D), with the paper's own tables as scholar-declared `experiment_provenance[]` and its claims as ClaimIntent entries. Case st2-01 pairs the abstract's "strictly eliminating negative transfer" and the conclusion's "strictly preventing negative transfer" with test rows where the method is worse than its frozen base (Electricity, ETTh1 MAE, GIFT-Eval bizitobs l2c MAE), the calibrated guard's own 2-of-7 regressing datasets (Table 7), and a calibration result scoped to validation splits; expected `OVERSTATED` or `NOT_SUPPORTED_BY_PROVENANCE`, both of which block. Case st2-02 pairs the contributions list's credit to neighbour attention (cut from a seven-component sentence) with an ablation where uniform averaging ties it on MSE and beats it on MAE; expected `OVERSTATED` or `NOT_SUPPORTED_BY_PROVENANCE`. Each passport also carries one supported control claim, expected `ALIGNED`. Only short quotes, numbers, and page references are stored. The suite is registered as `mechanical_match` and is `NOT_RUN`; two cases from one paper are a smoke check, not a detection rate. `scripts/test_experiment_alignment_overclaim_fixtures.py` keeps the passports valid against the #260 contracts and free of answers.
|
||||
|
||||
### Changed
|
||||
|
||||
+63
-1
@@ -362,7 +362,7 @@ See `references/mode_selection_guide.md` for details.
|
||||
| Just need an abstract | `abstract-only` | fidelity |
|
||||
| Need to check/fix citations | `citation-check` | fidelity |
|
||||
| Need to convert format (LaTeX, DOCX) or citation style | `format-convert` | fidelity |
|
||||
| Want a systematic literature review paper | `lit-review` | fidelity |
|
||||
| Want a literature review section or paper (to conduct a systematic review, use `deep-research` `systematic-review` mode) | `lit-review` | fidelity |
|
||||
| Need a venue-specific AI-usage disclosure bundle for submission | `disclosure` | fidelity |
|
||||
| Have a written rebuttal draft to QA against reviewer comments | `rebuttal-audit` | fidelity |
|
||||
|
||||
@@ -378,6 +378,68 @@ committees are peer review, not a committee for this variant, even when the user
|
||||
names the venue or the venue calls the role a committee (#854). The separate
|
||||
artifact is a source-accounted drafting aid and never enters peer-review Schema 11.
|
||||
|
||||
The canonical copy of the block below is `shared/references/review_form_note.md`; `scripts/check_review_form_note_sync.py` keeps this copy identical to it.
|
||||
|
||||
<!-- review-form-note:begin -->
|
||||
### Review-form note (#921)
|
||||
|
||||
The author decides whether to run a systematic review. ARS reminds the author that the choice exists; it does not judge whether a question fits a systematic review, and no review form is ever a default step.
|
||||
|
||||
**When to show it.** Show the note at the first of these two points. Both are actions the author takes:
|
||||
|
||||
1. The author selects `lit-review` mode (in `deep-research` or `academic-paper`, by slash command or by request).
|
||||
2. The author confirms the research question: in `deep-research` `full` mode, the author confirms the RQ Brief before Phase 2; in `socratic` mode, the author confirms the Mentor's closing RQ Brief or RQ Summary as their research question. Show the note right after that confirmation. A Socratic ending the author has not confirmed (a turn-cap ending, an ending the author calls unfinished, the stagnation suggestion to switch to `full` mode, or a switch to `full` mode) is not this point; a later confirmation is.
|
||||
|
||||
Whether the note appears must not depend on the topic, the wording, or the kind of research question. Do not show it at any other point, and do not show it, or hold it back, because a question looks like an effect question.
|
||||
|
||||
**When not to show it.**
|
||||
|
||||
- The note was already answered or skipped in this project or run. In a run with a passport file, look for a `checkpoint_closed` entry with `checkpoint_id: review-form-note` in the run ledger; without one, look in this conversation. Across separate sessions without a passport file the note can appear again; this is accepted. If the ledger holds a `checkpoint_opened` entry for `review-form-note` and no closing entry, the note is still awaiting its answer: show it again, append no second opening entry, and append the closing entry after the reply.
|
||||
- The author already named a review form in their own words or actions: entered `systematic-review` mode, asked for a systematic, scoping, rapid, narrative, or integrative review, or said they want no formal review. This test reads what the author said, not the content of the research question.
|
||||
|
||||
**How to show it.**
|
||||
|
||||
- Show the note text below verbatim: the English text in English conversations, the Traditional Chinese text in Traditional Chinese conversations, and the English text in every other language. Do not shorten, reorder, paraphrase, or add to it. Add no recommendation, default, or comment on which form fits the question, before or after it.
|
||||
- Then stop and wait for the author's reply. Do not start the literature search, the review, or Phase 2 before the author replies.
|
||||
- The note does not reopen the choice of workflow. If the author skips it, the mode the author asked for continues unchanged.
|
||||
- If the author asks which form fits their question, say that the choice is theirs. On request, describe any form in more detail, without a comparison that favours one form for their question.
|
||||
|
||||
**After the reply.**
|
||||
|
||||
- Skip, or a reply that keeps the current work: continue in the current mode. Skipping is a decision.
|
||||
- Systematic review: offer `deep-research` `systematic-review` mode, and enter it only when the author confirms.
|
||||
- Scoping review or rapid review: continue in the current mode, and say once that ARS has no separate mode for this form, so its protocol and reporting checklist (PRISMA-ScR for a scoping review) stay with the author.
|
||||
- Narrative or integrative review, or no formal review: continue in the current mode.
|
||||
- In a run with a passport file, record the note through `scripts/run_ledger.py append`: before waiting, unless the ledger already holds one, a `checkpoint_opened` entry (`checkpoint_id: review-form-note`, `stage`: the current stage or mode, `checkpoint_type: SLIM`, `question`: the note as shown, `options`: the five forms and `skip`); after the reply, a `checkpoint_closed` entry with `answer`: the form chosen or `skip`, and the author's exact words in `user_words`.
|
||||
- No path enters `systematic-review` mode on ARS's initiative. Only the author's explicit choice does.
|
||||
|
||||
**Note text (English):**
|
||||
|
||||
> **Before the review starts: which form of literature review?**
|
||||
> ARS does not choose this for you. There is no default and no recommendation, and the order below is not a ranking.
|
||||
>
|
||||
> - **Systematic review**: answers a focused question with a search and screening plan fixed in advance, with two people screening independently where possible. Months to more than a year for a team of several people, often with a registered protocol. In ARS: `systematic-review` mode.
|
||||
> - **Scoping review**: maps what has been studied on a topic, the main concepts, and the gaps, using a systematic approach. The work grows with the breadth of the topic. ARS has no separate mode for it.
|
||||
> - **Narrative or integrative review**: builds an argument or a framework from the literature. In a narrative review the author chooses the sources; an integrative review documents its search and evaluation and can combine different study designs. The work depends on the scope the author sets. In ARS: `lit-review` mode.
|
||||
> - **Rapid review**: a systematic review with some steps shortened or left out to deliver sooner, typically within weeks to a few months. ARS has no separate mode for it.
|
||||
> - **No formal review**: background from the sources at hand, for example for an introduction. It does not claim to cover the literature.
|
||||
>
|
||||
> Reply with the form you want, or reply "skip" to continue as you are. Either reply is your decision, and this note will not appear again in this project.
|
||||
|
||||
**Note text (Traditional Chinese):**
|
||||
|
||||
> **開始回顧之前:要做哪一種文獻回顧?**
|
||||
> 這件事由你決定,ARS 不替你選。下列選項沒有預設、沒有推薦,排列順序也不代表高下。
|
||||
>
|
||||
> - **系統性回顧(systematic review)**:回答一個聚焦的問題,檢索與篩選方式事先訂好,盡可能由兩人各自獨立篩選。一個數人團隊需要數個月到一年以上,通常有已登錄的研究計畫書。ARS 對應:`systematic-review` 模式。
|
||||
> - **範疇回顧(scoping review)**:用系統化的做法盤點一個主題已經研究了什麼、有哪些主要概念、缺口在哪裡。工作量隨主題的廣度增加。ARS 沒有專屬模式。
|
||||
> - **敘事或整合性回顧(narrative / integrative review)**:從文獻建立論證或架構。敘事回顧由作者選擇文獻;整合性回顧會記錄檢索與評估過程,並可合併不同研究設計。工作量取決於作者設定的範圍。ARS 對應:`lit-review` 模式。
|
||||
> - **快速回顧(rapid review)**:為了早點交出結果而縮短或省略部分步驟的系統性回顧,通常在數週到數個月內完成。ARS 沒有專屬模式。
|
||||
> - **不做正式回顧**:用手邊的文獻寫背景,例如論文的緒論。不宣稱涵蓋整體文獻。
|
||||
>
|
||||
> 請回覆你要的形式,或回覆「跳過」照目前的做法繼續。兩種回覆都算你的決定,這個專案裡不會再出現這則提醒。
|
||||
<!-- review-form-note:end -->
|
||||
|
||||
### Mode Selection Logic
|
||||
|
||||
> See `references/mode_selection_guide.md` for trigger-to-mode mappings and the full selection flowchart.
|
||||
|
||||
@@ -11,6 +11,8 @@ Trigger the `academic-paper` skill in `lit-review` mode. Produces an annotated b
|
||||
|
||||
Stay in `academic-paper` `lit-review` mode and do not reopen the choice of workflow: if the papers or sources the request refers to are missing, ask the user for them or offer to search for them within this mode.
|
||||
|
||||
The skill's review-form note (#921) belongs to this mode and does not reopen that choice: show it and wait when the skill says to, and continue in `lit-review` mode if the author skips it.
|
||||
|
||||
Mode reference: `${CLAUDE_PLUGIN_ROOT}/MODE_REGISTRY.md` § academic-paper.
|
||||
Skill entry: `${CLAUDE_PLUGIN_ROOT}/academic-paper/SKILL.md`.
|
||||
|
||||
|
||||
+67
-2
@@ -184,6 +184,68 @@ User Input
|
||||
+-- Only need fact-checking? --> fact-check mode
|
||||
```
|
||||
|
||||
The canonical copy of the block below is `shared/references/review_form_note.md`; `scripts/check_review_form_note_sync.py` keeps this copy identical to it.
|
||||
|
||||
<!-- review-form-note:begin -->
|
||||
### Review-form note (#921)
|
||||
|
||||
The author decides whether to run a systematic review. ARS reminds the author that the choice exists; it does not judge whether a question fits a systematic review, and no review form is ever a default step.
|
||||
|
||||
**When to show it.** Show the note at the first of these two points. Both are actions the author takes:
|
||||
|
||||
1. The author selects `lit-review` mode (in `deep-research` or `academic-paper`, by slash command or by request).
|
||||
2. The author confirms the research question: in `deep-research` `full` mode, the author confirms the RQ Brief before Phase 2; in `socratic` mode, the author confirms the Mentor's closing RQ Brief or RQ Summary as their research question. Show the note right after that confirmation. A Socratic ending the author has not confirmed (a turn-cap ending, an ending the author calls unfinished, the stagnation suggestion to switch to `full` mode, or a switch to `full` mode) is not this point; a later confirmation is.
|
||||
|
||||
Whether the note appears must not depend on the topic, the wording, or the kind of research question. Do not show it at any other point, and do not show it, or hold it back, because a question looks like an effect question.
|
||||
|
||||
**When not to show it.**
|
||||
|
||||
- The note was already answered or skipped in this project or run. In a run with a passport file, look for a `checkpoint_closed` entry with `checkpoint_id: review-form-note` in the run ledger; without one, look in this conversation. Across separate sessions without a passport file the note can appear again; this is accepted. If the ledger holds a `checkpoint_opened` entry for `review-form-note` and no closing entry, the note is still awaiting its answer: show it again, append no second opening entry, and append the closing entry after the reply.
|
||||
- The author already named a review form in their own words or actions: entered `systematic-review` mode, asked for a systematic, scoping, rapid, narrative, or integrative review, or said they want no formal review. This test reads what the author said, not the content of the research question.
|
||||
|
||||
**How to show it.**
|
||||
|
||||
- Show the note text below verbatim: the English text in English conversations, the Traditional Chinese text in Traditional Chinese conversations, and the English text in every other language. Do not shorten, reorder, paraphrase, or add to it. Add no recommendation, default, or comment on which form fits the question, before or after it.
|
||||
- Then stop and wait for the author's reply. Do not start the literature search, the review, or Phase 2 before the author replies.
|
||||
- The note does not reopen the choice of workflow. If the author skips it, the mode the author asked for continues unchanged.
|
||||
- If the author asks which form fits their question, say that the choice is theirs. On request, describe any form in more detail, without a comparison that favours one form for their question.
|
||||
|
||||
**After the reply.**
|
||||
|
||||
- Skip, or a reply that keeps the current work: continue in the current mode. Skipping is a decision.
|
||||
- Systematic review: offer `deep-research` `systematic-review` mode, and enter it only when the author confirms.
|
||||
- Scoping review or rapid review: continue in the current mode, and say once that ARS has no separate mode for this form, so its protocol and reporting checklist (PRISMA-ScR for a scoping review) stay with the author.
|
||||
- Narrative or integrative review, or no formal review: continue in the current mode.
|
||||
- In a run with a passport file, record the note through `scripts/run_ledger.py append`: before waiting, unless the ledger already holds one, a `checkpoint_opened` entry (`checkpoint_id: review-form-note`, `stage`: the current stage or mode, `checkpoint_type: SLIM`, `question`: the note as shown, `options`: the five forms and `skip`); after the reply, a `checkpoint_closed` entry with `answer`: the form chosen or `skip`, and the author's exact words in `user_words`.
|
||||
- No path enters `systematic-review` mode on ARS's initiative. Only the author's explicit choice does.
|
||||
|
||||
**Note text (English):**
|
||||
|
||||
> **Before the review starts: which form of literature review?**
|
||||
> ARS does not choose this for you. There is no default and no recommendation, and the order below is not a ranking.
|
||||
>
|
||||
> - **Systematic review**: answers a focused question with a search and screening plan fixed in advance, with two people screening independently where possible. Months to more than a year for a team of several people, often with a registered protocol. In ARS: `systematic-review` mode.
|
||||
> - **Scoping review**: maps what has been studied on a topic, the main concepts, and the gaps, using a systematic approach. The work grows with the breadth of the topic. ARS has no separate mode for it.
|
||||
> - **Narrative or integrative review**: builds an argument or a framework from the literature. In a narrative review the author chooses the sources; an integrative review documents its search and evaluation and can combine different study designs. The work depends on the scope the author sets. In ARS: `lit-review` mode.
|
||||
> - **Rapid review**: a systematic review with some steps shortened or left out to deliver sooner, typically within weeks to a few months. ARS has no separate mode for it.
|
||||
> - **No formal review**: background from the sources at hand, for example for an introduction. It does not claim to cover the literature.
|
||||
>
|
||||
> Reply with the form you want, or reply "skip" to continue as you are. Either reply is your decision, and this note will not appear again in this project.
|
||||
|
||||
**Note text (Traditional Chinese):**
|
||||
|
||||
> **開始回顧之前:要做哪一種文獻回顧?**
|
||||
> 這件事由你決定,ARS 不替你選。下列選項沒有預設、沒有推薦,排列順序也不代表高下。
|
||||
>
|
||||
> - **系統性回顧(systematic review)**:回答一個聚焦的問題,檢索與篩選方式事先訂好,盡可能由兩人各自獨立篩選。一個數人團隊需要數個月到一年以上,通常有已登錄的研究計畫書。ARS 對應:`systematic-review` 模式。
|
||||
> - **範疇回顧(scoping review)**:用系統化的做法盤點一個主題已經研究了什麼、有哪些主要概念、缺口在哪裡。工作量隨主題的廣度增加。ARS 沒有專屬模式。
|
||||
> - **敘事或整合性回顧(narrative / integrative review)**:從文獻建立論證或架構。敘事回顧由作者選擇文獻;整合性回顧會記錄檢索與評估過程,並可合併不同研究設計。工作量取決於作者設定的範圍。ARS 對應:`lit-review` 模式。
|
||||
> - **快速回顧(rapid review)**:為了早點交出結果而縮短或省略部分步驟的系統性回顧,通常在數週到數個月內完成。ARS 沒有專屬模式。
|
||||
> - **不做正式回顧**:用手邊的文獻寫背景,例如論文的緒論。不宣稱涵蓋整體文獻。
|
||||
>
|
||||
> 請回覆你要的形式,或回覆「跳過」照目前的做法繼續。兩種回覆都算你的決定,這個專案裡不會再出現這則提醒。
|
||||
<!-- review-form-note:end -->
|
||||
|
||||
---
|
||||
|
||||
## Orchestration Workflow (6 Phases)
|
||||
@@ -212,6 +274,7 @@ User: "Research [topic]"
|
||||
- Verdict: PASS / REVISE (with specific feedback)
|
||||
|
|
||||
** User confirmation before Phase 2 **
|
||||
(review-form note at this confirmation, per § Review-form note)
|
||||
|
|
||||
=== Phase 2: INVESTIGATION ===
|
||||
|
|
||||
@@ -301,7 +364,7 @@ User: "Research [topic]"
|
||||
1. ⚠️ **IRON RULE**: **Devil's Advocate** has 3 mandatory checkpoints; **Critical-severity** issues block progression
|
||||
2. Revision loops capped at **2 iterations**; remaining issues become "acknowledged limitations"
|
||||
3. ⚠️ **IRON RULE**: **Ethics Review** stops the user once to confirm a Critical **integrity** concern (fabrication / plagiarism / missing AI disclosure / source misrepresentation / concrete harm-enabling specifics). Overridable with recorded reasoning — it confirms, it does not veto. Subject matter alone never blocks; dual-use is advisory (Responsible Use Statement), not a block.
|
||||
4. User confirmation required after Phase 1 before proceeding
|
||||
4. User confirmation required after Phase 1 before proceeding; the review-form note (§ Review-form note, #921) is shown at this confirmation unless its skip conditions apply
|
||||
|
||||
---
|
||||
|
||||
@@ -337,6 +400,8 @@ line before any clearly labeled AI-generated candidate. Never switch silently.
|
||||
|
||||
> See `references/socratic_mode_protocol.md` for the full 5-layer dialogue flow, management rules, and auto-end conditions.
|
||||
|
||||
When the author confirms the closing RQ Brief or RQ Summary as their research question, show the review-form note (§ Review-form note, #921) unless its skip conditions apply.
|
||||
|
||||
### Opt-in Reading Probe (v3.5.1)
|
||||
|
||||
Setting `ARS_SOCRATIC_READING_PROBE=1` enables a one-time honesty probe during **goal-oriented** Socratic sessions. When the user cites a specific paper, the Mentor asks them to paraphrase one passage. Decline is logged without penalty. Default OFF. See `agents/socratic_mentor_agent.md` §"Optional Reading Probe Layer".
|
||||
@@ -424,7 +489,7 @@ Then add:
|
||||
- strongest `WHAT`
|
||||
- unresolved global gap
|
||||
|
||||
If the user later wants a broader evidence matrix, thematic synthesis, or PRISMA-like coverage, escalate from `three-way-scan` to `lit-review` or `systematic-review`.
|
||||
If the user later wants a broader evidence matrix or thematic synthesis, escalate from `three-way-scan` to `lit-review`. Escalate to `systematic-review` only when the author chooses a systematic review (§ Review-form note, #921).
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -108,7 +108,7 @@ User Input
|
||||
|------|------|
|
||||
| **Applicable Scenario** | Need systematic literature search and synthesis analysis, but not a complete research report |
|
||||
| **Not Applicable** | Need a complete report with original analysis; only need to verify a few facts; need methodology design |
|
||||
| **Typical Users** | Graduate students writing the literature review chapter of their thesis, research teams conducting systematic reviews, coursework assignments |
|
||||
| **Typical Users** | Graduate students writing the literature review chapter of their thesis, coursework assignments (teams conducting a systematic review use `systematic-review` mode) |
|
||||
| **Expected Output** | Annotated bibliography + synthesis analysis (1,500-4,000 words), including thematic classification, evidence matrix, research gaps |
|
||||
| **Expected Dialogue Rounds** | 1-2 rounds (confirm search scope) |
|
||||
| **Agents Activated** | 3 (Biblio + Verification + Synthesis) |
|
||||
@@ -118,7 +118,6 @@ User Input
|
||||
```
|
||||
"Literature review on SDGs in higher education"
|
||||
"Literature review: the evolution of quality assurance in Taiwan's higher education"
|
||||
"Systematic review of AI-assisted assessment"
|
||||
```
|
||||
|
||||
---
|
||||
@@ -128,7 +127,7 @@ User Input
|
||||
| Item | Description |
|
||||
|------|------|
|
||||
| **Applicable Scenario** | Need a disciplined shortlist of papers compared in a stable WHY/HOW/WHAT frame, lighter than a full literature review |
|
||||
| **Not Applicable** | Need a complete evidence matrix, thematic synthesis, or PRISMA-like coverage (escalate to `lit-review` / `systematic-review`); need to verify specific facts (`fact-check`) |
|
||||
| **Not Applicable** | Need a complete evidence matrix, thematic synthesis, or PRISMA-like coverage (escalate to `lit-review`; `systematic-review` only when the author chooses it); need to verify specific facts (`fact-check`) |
|
||||
| **Typical Users** | Researchers scoping a new area, students triaging a reading list, anyone deciding which papers to read in depth |
|
||||
| **Expected Output** | Per-paper WHY/HOW/WHAT shortlist + cross-paper synthesis (common WHY, divergent HOW, strongest WHAT, unresolved gap) (800-2,000 words) |
|
||||
| **Expected Dialogue Rounds** | 1-2 rounds (confirm candidate set) |
|
||||
@@ -238,7 +237,7 @@ User Input
|
||||
socratic → full Continue with complete research after Socratic completion
|
||||
socratic → academic-paper Write paper directly after Socratic completion
|
||||
lit-review → full Want complete analysis after literature review
|
||||
lit-review → systematic-review Need formal PRISMA compliance after initial lit survey
|
||||
lit-review → systematic-review Author chooses a systematic review after initial lit survey
|
||||
fact-check → full Need deeper research after fact-checking
|
||||
quick → full Worth going deeper after quick research
|
||||
review → full Need to re-research after review
|
||||
@@ -314,7 +313,7 @@ Rules for switching between modes mid-research. Not all transitions are safe.
|
||||
- **Quality Delta**: Fact-check is binary (true/false/mixed); full mode produces nuanced analysis
|
||||
|
||||
### Transition: lit-review → systematic-review
|
||||
- **When**: Literature review reveals the topic warrants formal PRISMA compliance (e.g., for publication in a journal that requires it)
|
||||
- **When**: The author chooses a systematic review, for example after the review-form note (`deep-research/SKILL.md` § Review-form note, #921) or because their target journal requires PRISMA compliance. ARS does not start this transition because the topic seems to warrant it
|
||||
- **Reusable**: Initial keyword strategy, some identified sources (need re-screening)
|
||||
- **Must Redo**: Protocol registration, formal inclusion/exclusion criteria, dual screening, risk of bias assessment, meta-analysis feasibility assessment
|
||||
- **Quality Delta**: systematic-review requires protocol, RoB assessment, GRADE; lit-review has none of these
|
||||
|
||||
@@ -68,6 +68,10 @@ path = "scripts/test_check_routing_core_sync.py"
|
||||
id = "916-method-weaknesses-sync"
|
||||
path = "scripts/test_check_method_weaknesses_sync.py"
|
||||
|
||||
[[pytest]]
|
||||
id = "921-review-form-note-sync"
|
||||
path = "scripts/test_check_review_form_note_sync.py"
|
||||
|
||||
[[pytest]]
|
||||
id = "915-experiment-alignment-overclaim-fixtures"
|
||||
path = "scripts/test_experiment_alignment_overclaim_fixtures.py"
|
||||
|
||||
@@ -0,0 +1,173 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Review-form note sync lint (#921).
|
||||
|
||||
`shared/references/review_form_note.md` holds the canonical review-form note:
|
||||
the rules for when ARS reminds the author that a literature review can take
|
||||
several forms, and the fixed note text in English and Traditional Chinese.
|
||||
This lint keeps each surface's verbatim copy byte-identical to it and checks
|
||||
that the note offers its options without a default or a ranking.
|
||||
|
||||
Checks:
|
||||
RF-1 The canonical file holds exactly one begin marker and one end marker,
|
||||
each alone on its line, begin before end, around a non-empty block.
|
||||
RF-2 Every surface holds exactly one such marker pair, and its block is
|
||||
byte-identical to the canonical block, line endings included.
|
||||
RF-3 The canonical block holds an English and a Traditional Chinese note
|
||||
text. Every non-blank line of each opens with `>`. Each has exactly
|
||||
five list items, opening with the five review forms' labels in the
|
||||
same order, and its neutrality sentence verbatim;
|
||||
no other note line carries a word that ranks, recommends, or
|
||||
preselects, or a selection mark.
|
||||
|
||||
Usage:
|
||||
python scripts/check_review_form_note_sync.py [--root PATH]
|
||||
|
||||
Exit codes: 0 all checks pass; 1 a check failed; 2 a required file is missing.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
from _skill_lint import read_or_exit2 # noqa: E402
|
||||
from check_routing_core_sync import first_difference # noqa: E402
|
||||
|
||||
CANONICAL = Path("shared/references/review_form_note.md")
|
||||
SURFACES = (
|
||||
Path("deep-research/SKILL.md"),
|
||||
Path("academic-paper/SKILL.md"),
|
||||
)
|
||||
BEGIN = "<!-- review-form-note:begin -->"
|
||||
END = "<!-- review-form-note:end -->"
|
||||
|
||||
NOTE_HEADINGS = ("**Note text (English):**", "**Note text (Traditional Chinese):**")
|
||||
# The exact bold label that opens each option line, per note language, in order.
|
||||
OPTIONS = (
|
||||
("**Systematic review**:", "**系統性回顧(systematic review)**:"),
|
||||
("**Scoping review**:", "**範疇回顧(scoping review)**:"),
|
||||
("**Narrative or integrative review**:", "**敘事或整合性回顧(narrative / integrative review)**:"),
|
||||
("**Rapid review**:", "**快速回顧(rapid review)**:"),
|
||||
("**No formal review**:", "**不做正式回顧**:"),
|
||||
)
|
||||
# The one line per note that says there is no default, recommendation, or
|
||||
# ranking. It must be present verbatim, and it is the only note line allowed to
|
||||
# use those words.
|
||||
NEUTRALITY = (
|
||||
"> ARS does not choose this for you. There is no default and no recommendation, "
|
||||
"and the order below is not a ranking.",
|
||||
"> 這件事由你決定,ARS 不替你選。下列選項沒有預設、沒有推薦,排列順序也不代表高下。",
|
||||
)
|
||||
LIST_ITEM = re.compile(r"^>\s*(?:[-*+]|\d+[.)])\s")
|
||||
RANKING_WORDS = re.compile(
|
||||
r"\b(?:recommend\w*|default\w*|best|prefer\w*|suggest\w*|ideal\w*|should|most|least|"
|
||||
r"better|optimal\w*|advis\w*)\b|推薦|建議|預設|最|首選|應該|較適合|更好|優先",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
SELECTION_MARK = re.compile(r"\[[ xX✓✔]\]|[☑☒✅✔✓★⭐←]|\(\*\)")
|
||||
|
||||
|
||||
def extract_block(text: str, label: str) -> tuple[str | None, list[str]]:
|
||||
"""Return the text between the one marker pair, or None with the errors.
|
||||
A marker line may end in CR; the block keeps its CRs for the comparison."""
|
||||
lines = text.split("\n")
|
||||
begins = [i for i, line in enumerate(lines) if line.rstrip("\r") == BEGIN]
|
||||
ends = [i for i, line in enumerate(lines) if line.rstrip("\r") == END]
|
||||
errors: list[str] = []
|
||||
for marker, whole in ((BEGIN, begins), (END, ends)):
|
||||
total = text.count(marker)
|
||||
if total != 1 or len(whole) != 1:
|
||||
errors.append(f"{label}: expected one {marker} alone on its line, "
|
||||
f"found {total} occurrence(s), {len(whole)} on their own line")
|
||||
if errors:
|
||||
return None, errors
|
||||
if begins[0] > ends[0]:
|
||||
return None, [f"{label}: {END} comes before {BEGIN}"]
|
||||
block = "\n".join(lines[begins[0] + 1:ends[0]])
|
||||
if not block.strip():
|
||||
return None, [f"{label}: the review-form note block is empty"]
|
||||
return block, []
|
||||
|
||||
|
||||
def note_lines(block: str, heading: str) -> list[str] | None:
|
||||
"""Return the non-blank lines of the note text under `heading`, up to the
|
||||
next note-text heading, or None when the heading is missing. Line
|
||||
terminators are dropped here; RF-2 compares the raw bytes."""
|
||||
lines = [line.rstrip("\r") for line in block.split("\n")]
|
||||
if heading not in lines:
|
||||
return None
|
||||
note = []
|
||||
for line in lines[lines.index(heading) + 1:]:
|
||||
if line.startswith("**Note text"):
|
||||
break
|
||||
if line.strip():
|
||||
note.append(line)
|
||||
return note
|
||||
|
||||
|
||||
def check_neutrality(block: str) -> list[str]:
|
||||
"""RF-3: each note lists the five forms under their exact labels, in order,
|
||||
keeps its neutrality sentence, and has no ranking word or selection mark."""
|
||||
errors: list[str] = []
|
||||
for index, heading in enumerate(NOTE_HEADINGS):
|
||||
where = f"RF-3 {CANONICAL}: {heading}"
|
||||
lines = note_lines(block, heading)
|
||||
if lines is None:
|
||||
errors.append(f"RF-3 {CANONICAL}: note text heading {heading} is missing")
|
||||
continue
|
||||
for line in lines:
|
||||
if not line.startswith(">"):
|
||||
errors.append(f"{where} has a line outside the quoted note text: {line[:60]!r}")
|
||||
if NEUTRALITY[index] not in lines:
|
||||
errors.append(f"{where} lacks its neutrality sentence verbatim")
|
||||
items = [line for line in lines if LIST_ITEM.match(line)]
|
||||
labels = [names[index] for names in OPTIONS]
|
||||
if len(items) != len(labels):
|
||||
errors.append(f"{where} lists {len(items)} options, expected {len(labels)}")
|
||||
else:
|
||||
for line, label in zip(items, labels):
|
||||
if not line.startswith(f"> - {label}"):
|
||||
errors.append(f"{where} option out of order or relabelled: "
|
||||
f"expected it to open with {label!r}, found {line[:60]!r}")
|
||||
for line in lines:
|
||||
if line == NEUTRALITY[index]:
|
||||
continue
|
||||
for pattern, kind in ((RANKING_WORDS, "ranking word"), (SELECTION_MARK, "selection mark")):
|
||||
match = pattern.search(line)
|
||||
if match:
|
||||
errors.append(f"{where} carries the {kind} {match.group(0)!r} in {line[:60]!r}")
|
||||
return errors
|
||||
|
||||
|
||||
def check(root: Path) -> list[str]:
|
||||
"""Run RF-1 to RF-3 under `root`; a missing file exits 2."""
|
||||
canonical, errors = extract_block(read_or_exit2(root, str(CANONICAL), exact=True),
|
||||
f"RF-1 {CANONICAL}")
|
||||
if canonical is not None:
|
||||
errors += check_neutrality(canonical)
|
||||
for rel in SURFACES:
|
||||
block, copy_errors = extract_block(read_or_exit2(root, str(rel), exact=True), f"RF-2 {rel}")
|
||||
errors += copy_errors
|
||||
if block is not None and canonical is not None and block != canonical:
|
||||
errors.append(f"RF-2 {rel}: review-form note block differs from {CANONICAL} "
|
||||
f"({first_difference(block, canonical)})")
|
||||
return errors
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
|
||||
parser.add_argument("--root", type=Path, default=Path(__file__).resolve().parent.parent)
|
||||
args = parser.parse_args(argv)
|
||||
errors = check(args.root)
|
||||
if errors:
|
||||
for error in errors:
|
||||
print(error, file=sys.stderr)
|
||||
return 1
|
||||
print(f"check_review_form_note_sync: OK ({len(SURFACES)} surfaces match {CANONICAL})")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,232 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Tests for check_review_form_note_sync.py (#921).
|
||||
|
||||
Mutation tests confirm the lint is not accept-all: every break in the marker
|
||||
grammar, in a surface's bytes, or in the note's neutral option list must fail
|
||||
it, and the clean repository must pass.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import shutil
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from scripts.check_review_form_note_sync import (
|
||||
BEGIN,
|
||||
CANONICAL,
|
||||
END,
|
||||
NEUTRALITY,
|
||||
SURFACES,
|
||||
check,
|
||||
extract_block,
|
||||
)
|
||||
from tests.test_helpers import run_script
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
LINT = REPO_ROOT / "scripts" / "check_review_form_note_sync.py"
|
||||
SKILL = Path("academic-paper/SKILL.md")
|
||||
PHRASE = "The author decides whether to run a systematic review."
|
||||
EN_SCOPING = "> - **Scoping review**: maps what has been studied"
|
||||
ZH_SCOPING = "> - **範疇回顧(scoping review)**:用系統化的做法"
|
||||
EN_NEUTRALITY = NEUTRALITY[0].removeprefix("> ")
|
||||
EN_NO_FORMAL = "It does not claim to cover the literature."
|
||||
ZH_NO_FORMAL = "不宣稱涵蓋整體文獻。"
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def tree(tmp_path: Path) -> Path:
|
||||
"""A copy of just the files the lint reads, under a temp root."""
|
||||
for rel in (CANONICAL, *SURFACES):
|
||||
dest = tmp_path / rel
|
||||
dest.parent.mkdir(parents=True, exist_ok=True)
|
||||
shutil.copyfile(REPO_ROOT / rel, dest)
|
||||
return tmp_path
|
||||
|
||||
|
||||
def _edit(root: Path, rel: Path, old: str, new: str) -> None:
|
||||
path = root / rel
|
||||
text = path.read_text(encoding="utf-8")
|
||||
assert old in text, f"{old!r} not in {rel}"
|
||||
path.write_text(text.replace(old, new, 1), encoding="utf-8")
|
||||
|
||||
|
||||
def _edit_everywhere(root: Path, old: str, new: str) -> None:
|
||||
"""Change the canonical block and every copy, so only RF-3 can fail."""
|
||||
for rel in (CANONICAL, *SURFACES):
|
||||
_edit(root, rel, old, new)
|
||||
|
||||
|
||||
def _line(root: Path, prefix: str) -> str:
|
||||
"""The first canonical line that starts with `prefix`."""
|
||||
text = (root / CANONICAL).read_text(encoding="utf-8")
|
||||
return next(line for line in text.split("\n") if line.startswith(prefix))
|
||||
|
||||
|
||||
def _errors(root: Path) -> str:
|
||||
return "\n".join(check(root))
|
||||
|
||||
|
||||
def test_repository_passes() -> None:
|
||||
result = run_script(LINT, "--root", str(REPO_ROOT))
|
||||
assert result.returncode == 0, result.stderr
|
||||
assert "2 surfaces match" in result.stdout
|
||||
|
||||
|
||||
def test_canonical_lists_every_surface() -> None:
|
||||
text = (REPO_ROOT / CANONICAL).read_text(encoding="utf-8")
|
||||
for rel in SURFACES:
|
||||
assert f"`{rel}`" in text
|
||||
|
||||
|
||||
def test_copied_tree_passes(tree: Path) -> None:
|
||||
assert check(tree) == []
|
||||
|
||||
|
||||
@pytest.mark.parametrize("rel", SURFACES)
|
||||
def test_one_changed_byte_in_a_surface_fails(tree: Path, rel: Path) -> None:
|
||||
_edit(tree, rel, PHRASE, PHRASE.replace("decides", "Decides"))
|
||||
assert f"RF-2 {rel}: review-form note block differs" in _errors(tree)
|
||||
|
||||
|
||||
def test_line_ending_drift_in_a_surface_fails(tree: Path) -> None:
|
||||
path = tree / SKILL
|
||||
path.write_bytes(path.read_bytes().replace(b"\n", b"\r\n"))
|
||||
assert "differs only in its line ending" in _errors(tree)
|
||||
|
||||
|
||||
def test_surface_without_markers_fails(tree: Path) -> None:
|
||||
_edit(tree, SKILL, BEGIN + "\n", "")
|
||||
_edit(tree, SKILL, "\n" + END, "")
|
||||
assert f"RF-2 {SKILL}: expected one {BEGIN}" in _errors(tree)
|
||||
|
||||
|
||||
def test_surface_with_block_twice_fails(tree: Path) -> None:
|
||||
block, errors = extract_block((REPO_ROOT / CANONICAL).read_text(encoding="utf-8"), "t")
|
||||
assert block is not None, errors
|
||||
text = (tree / SKILL).read_text(encoding="utf-8")
|
||||
(tree / SKILL).write_text(f"{text}\n{BEGIN}\n{block}\n{END}\n", encoding="utf-8")
|
||||
assert "found 2 occurrence(s)" in _errors(tree)
|
||||
|
||||
|
||||
def test_marker_not_alone_on_its_line_fails(tree: Path) -> None:
|
||||
_edit(tree, SKILL, BEGIN + "\n", "Text " + BEGIN + "\n")
|
||||
assert "0 on their own line" in _errors(tree)
|
||||
|
||||
|
||||
def test_canonical_markers_reversed_fails(tree: Path) -> None:
|
||||
_edit(tree, CANONICAL, BEGIN, "@@BEGIN@@")
|
||||
_edit(tree, CANONICAL, END, BEGIN)
|
||||
_edit(tree, CANONICAL, "@@BEGIN@@", END)
|
||||
assert f"RF-1 {CANONICAL}: {END} comes before {BEGIN}" in _errors(tree)
|
||||
|
||||
|
||||
def test_empty_canonical_block_fails(tree: Path) -> None:
|
||||
text = (tree / CANONICAL).read_text(encoding="utf-8")
|
||||
head, rest = text.split(BEGIN + "\n", 1)
|
||||
_, tail = rest.split(END, 1)
|
||||
(tree / CANONICAL).write_text(head + BEGIN + "\n\n" + END + tail, encoding="utf-8")
|
||||
assert "block is empty" in _errors(tree)
|
||||
|
||||
|
||||
def test_changed_canonical_fails_every_surface(tree: Path) -> None:
|
||||
_edit(tree, CANONICAL, PHRASE, PHRASE.replace("decides", "Decides"))
|
||||
errors = _errors(tree)
|
||||
for rel in SURFACES:
|
||||
assert f"RF-2 {rel}: review-form note block differs" in errors
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("old", "new", "word"),
|
||||
[
|
||||
(EN_SCOPING, EN_SCOPING + " Recommended for most theses.", "Recommended"),
|
||||
(EN_SCOPING, EN_SCOPING + " (default)", "default"),
|
||||
(EN_SCOPING, EN_SCOPING + " Often the best fit.", "best"),
|
||||
(EN_NO_FORMAL, "Least effort; " + EN_NO_FORMAL, "Least"),
|
||||
(ZH_SCOPING, ZH_SCOPING + "(推薦)", "推薦"),
|
||||
(ZH_SCOPING, ZH_SCOPING + "多數論文較適合", "較適合"),
|
||||
(ZH_NO_FORMAL, "最省力,但" + ZH_NO_FORMAL, "最"),
|
||||
("> Reply with the form you want", "> We suggest a systematic review. Reply with the form you want",
|
||||
"suggest"),
|
||||
],
|
||||
)
|
||||
def test_ranking_word_in_the_note_fails(tree: Path, old: str, new: str, word: str) -> None:
|
||||
_edit_everywhere(tree, old, new)
|
||||
errors = _errors(tree)
|
||||
assert f"carries the ranking word {word!r}" in errors
|
||||
assert "RF-2" not in errors
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mark", ["[x] ", "✅ ", "(*) "])
|
||||
def test_selection_mark_on_an_option_fails(tree: Path, mark: str) -> None:
|
||||
_edit_everywhere(tree, EN_NO_FORMAL, EN_NO_FORMAL + " " + mark)
|
||||
errors = _errors(tree)
|
||||
assert f"carries the selection mark {mark.strip()!r}" in errors
|
||||
|
||||
|
||||
def test_replaced_neutrality_sentence_fails(tree: Path) -> None:
|
||||
_edit_everywhere(tree, EN_NEUTRALITY, "We recommend a systematic review.")
|
||||
errors = _errors(tree)
|
||||
assert "(English):** lacks its neutrality sentence verbatim" in errors
|
||||
assert "carries the ranking word 'recommend'" in errors
|
||||
|
||||
|
||||
def test_ranking_word_inside_another_word_passes(tree: Path) -> None:
|
||||
_edit_everywhere(tree, EN_NO_FORMAL, EN_NO_FORMAL + " Asbestos studies included.")
|
||||
assert check(tree) == []
|
||||
|
||||
|
||||
def test_crlf_in_every_file_passes(tree: Path) -> None:
|
||||
for rel in (CANONICAL, *SURFACES):
|
||||
path = tree / rel
|
||||
path.write_bytes(path.read_bytes().replace(b"\n", b"\r\n"))
|
||||
assert check(tree) == []
|
||||
|
||||
|
||||
def test_dropped_option_fails(tree: Path) -> None:
|
||||
line = _line(tree, "> - **Rapid review**")
|
||||
_edit_everywhere(tree, line + "\n", "")
|
||||
assert "(English):** lists 4 options, expected 5" in _errors(tree)
|
||||
|
||||
|
||||
def test_sixth_unbolded_option_fails(tree: Path) -> None:
|
||||
_edit_everywhere(tree, "> - **No formal review**", "> - Umbrella review: a review of reviews.\n> - **No formal review**")
|
||||
assert "(English):** lists 6 options, expected 5" in _errors(tree)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"inserted",
|
||||
[
|
||||
"We recommend a systematic review.",
|
||||
" > - Umbrella review: a review of reviews.",
|
||||
],
|
||||
)
|
||||
def test_line_outside_the_quote_fails(tree: Path, inserted: str) -> None:
|
||||
_edit_everywhere(tree, EN_NEUTRALITY + "\n", EN_NEUTRALITY + "\n" + inserted + "\n")
|
||||
assert f"(English):** has a line outside the quoted note text: {inserted!r}" in _errors(tree)
|
||||
|
||||
|
||||
def test_relabelled_option_fails(tree: Path) -> None:
|
||||
_edit_everywhere(tree, "> - **Scoping review**: maps", "> - **Mapping review**: a Scoping review maps")
|
||||
assert "(English):** option out of order or relabelled" in _errors(tree)
|
||||
|
||||
|
||||
def test_reordered_options_fail(tree: Path) -> None:
|
||||
first = _line(tree, "> - **系統性回顧")
|
||||
second = _line(tree, "> - **範疇回顧")
|
||||
_edit_everywhere(tree, first + "\n" + second, second + "\n" + first)
|
||||
assert "(Traditional Chinese):** option out of order or relabelled" in _errors(tree)
|
||||
|
||||
|
||||
def test_missing_language_fails(tree: Path) -> None:
|
||||
_edit_everywhere(tree, "**Note text (Traditional Chinese):**", "**Note text (zh):**")
|
||||
assert "heading **Note text (Traditional Chinese):** is missing" in _errors(tree)
|
||||
|
||||
|
||||
def test_cli_exit_codes(tree: Path) -> None:
|
||||
_edit(tree, SKILL, PHRASE, PHRASE.replace("decides", "Decides"))
|
||||
assert run_script(LINT, "--root", str(tree)).returncode == 1
|
||||
(tree / SKILL).unlink()
|
||||
result = run_script(LINT, "--root", str(tree))
|
||||
assert result.returncode == 2
|
||||
assert f"required file missing: {SKILL}" in result.stderr
|
||||
@@ -0,0 +1,82 @@
|
||||
# Review-Form Note (#921)
|
||||
|
||||
This file holds the canonical review-form note: a fixed-point reminder that lists the forms a literature review can take and leaves the choice to the author. The block between the markers is copied verbatim into each surface below; `scripts/check_review_form_note_sync.py` keeps every copy byte-identical to this one and checks the note text for a default, a ranking, or a recommendation.
|
||||
|
||||
| Surface | Where the block sits |
|
||||
|---|---|
|
||||
| `deep-research/SKILL.md` | After the Mode Selection Guide (covers `lit-review` mode and the RQ Brief confirmation in `full` and `socratic` modes) |
|
||||
| `academic-paper/SKILL.md` | After the Mode Selection Guide (covers `lit-review` mode) |
|
||||
|
||||
Why the note exists: whether to run a systematic review is a research-design decision that costs the author months of work and a team. ARS enters `systematic-review` mode only when the user asks for it, so users who do not know the option exists never hear about it. The note closes that gap without letting the model decide: deciding *when* to remind is itself a judgment, so the note is tied to two actions the author takes (selecting `lit-review` mode, confirming the RQ Brief) and never to the content of the question.
|
||||
|
||||
Sources for the one-line descriptions:
|
||||
|
||||
- Systematic review time and team size: Borah, R., Brown, A. W., Capers, P. L., & Kaiser, K. A. (2017). Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. *BMJ Open, 7*(2), e012545. https://doi.org/10.1136/bmjopen-2016-012545 (195 registered, completed reviews: mean 67.3 weeks from registration to publication, mean 5 authors; the same abstract reports means of 42 weeks for funded and 26 weeks for unfunded reviews). The sample is medical-intervention reviews and the means vary by measure, so the note says "months to more than a year", not a figure.
|
||||
- Scoping review purpose: Tricco, A. C., Lillie, E., Zarin, W., et al. (2018). PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and explanation. *Annals of Internal Medicine, 169*(7), 467–473. https://doi.org/10.7326/M18-0850 ("follow a systematic approach to map evidence on a topic and identify main concepts, theories, sources, and knowledge gaps").
|
||||
- Rapid review, and two independent screeners for a systematic review: Garritty, C., Gartlehner, G., Nussbaumer-Streit, B., et al. (2021). Cochrane Rapid Reviews Methods Group offers evidence-informed guidance to conduct rapid reviews. *Journal of Clinical Epidemiology, 130*, 13–22. https://doi.org/10.1016/j.jclinepi.2020.10.007 (§3.3.4: "MECIR states that it is desirable to use two screeners working independently"; the guidance applies "when decisions need to be made in a period of weeks to a few months").
|
||||
- Integrative review: Whittemore, R., & Knafl, K. (2005). The integrative review: Updated methodology. *Journal of Advanced Nursing, 52*(5), 546–553. https://doi.org/10.1111/j.1365-2648.2005.03621.x (the only review approach that "allows for the combination of diverse methodologies"; its stages include searching the literature and evaluating data from primary sources).
|
||||
- The scoping and narrative lines give no time figure because none of these sources supports one; they say what the work depends on instead.
|
||||
|
||||
Other languages: the note has an English and a Traditional Chinese text. Every other language gets the English text, so that no user sees a paraphrase the model wrote on the spot. A community locale pack may propose a reviewed translation as a third fixed text.
|
||||
|
||||
Out of scope: any model assessment of whether a question suits a systematic review; the screening skill proposed in #919. Whether sessions show the note at the right points is prompt-level and unmeasured; the routing fixtures in `tests/fixtures/issue_133_routing/` (13–15) are smoke tests, not a rate.
|
||||
|
||||
<!-- review-form-note:begin -->
|
||||
### Review-form note (#921)
|
||||
|
||||
The author decides whether to run a systematic review. ARS reminds the author that the choice exists; it does not judge whether a question fits a systematic review, and no review form is ever a default step.
|
||||
|
||||
**When to show it.** Show the note at the first of these two points. Both are actions the author takes:
|
||||
|
||||
1. The author selects `lit-review` mode (in `deep-research` or `academic-paper`, by slash command or by request).
|
||||
2. The author confirms the research question: in `deep-research` `full` mode, the author confirms the RQ Brief before Phase 2; in `socratic` mode, the author confirms the Mentor's closing RQ Brief or RQ Summary as their research question. Show the note right after that confirmation. A Socratic ending the author has not confirmed (a turn-cap ending, an ending the author calls unfinished, the stagnation suggestion to switch to `full` mode, or a switch to `full` mode) is not this point; a later confirmation is.
|
||||
|
||||
Whether the note appears must not depend on the topic, the wording, or the kind of research question. Do not show it at any other point, and do not show it, or hold it back, because a question looks like an effect question.
|
||||
|
||||
**When not to show it.**
|
||||
|
||||
- The note was already answered or skipped in this project or run. In a run with a passport file, look for a `checkpoint_closed` entry with `checkpoint_id: review-form-note` in the run ledger; without one, look in this conversation. Across separate sessions without a passport file the note can appear again; this is accepted. If the ledger holds a `checkpoint_opened` entry for `review-form-note` and no closing entry, the note is still awaiting its answer: show it again, append no second opening entry, and append the closing entry after the reply.
|
||||
- The author already named a review form in their own words or actions: entered `systematic-review` mode, asked for a systematic, scoping, rapid, narrative, or integrative review, or said they want no formal review. This test reads what the author said, not the content of the research question.
|
||||
|
||||
**How to show it.**
|
||||
|
||||
- Show the note text below verbatim: the English text in English conversations, the Traditional Chinese text in Traditional Chinese conversations, and the English text in every other language. Do not shorten, reorder, paraphrase, or add to it. Add no recommendation, default, or comment on which form fits the question, before or after it.
|
||||
- Then stop and wait for the author's reply. Do not start the literature search, the review, or Phase 2 before the author replies.
|
||||
- The note does not reopen the choice of workflow. If the author skips it, the mode the author asked for continues unchanged.
|
||||
- If the author asks which form fits their question, say that the choice is theirs. On request, describe any form in more detail, without a comparison that favours one form for their question.
|
||||
|
||||
**After the reply.**
|
||||
|
||||
- Skip, or a reply that keeps the current work: continue in the current mode. Skipping is a decision.
|
||||
- Systematic review: offer `deep-research` `systematic-review` mode, and enter it only when the author confirms.
|
||||
- Scoping review or rapid review: continue in the current mode, and say once that ARS has no separate mode for this form, so its protocol and reporting checklist (PRISMA-ScR for a scoping review) stay with the author.
|
||||
- Narrative or integrative review, or no formal review: continue in the current mode.
|
||||
- In a run with a passport file, record the note through `scripts/run_ledger.py append`: before waiting, unless the ledger already holds one, a `checkpoint_opened` entry (`checkpoint_id: review-form-note`, `stage`: the current stage or mode, `checkpoint_type: SLIM`, `question`: the note as shown, `options`: the five forms and `skip`); after the reply, a `checkpoint_closed` entry with `answer`: the form chosen or `skip`, and the author's exact words in `user_words`.
|
||||
- No path enters `systematic-review` mode on ARS's initiative. Only the author's explicit choice does.
|
||||
|
||||
**Note text (English):**
|
||||
|
||||
> **Before the review starts: which form of literature review?**
|
||||
> ARS does not choose this for you. There is no default and no recommendation, and the order below is not a ranking.
|
||||
>
|
||||
> - **Systematic review**: answers a focused question with a search and screening plan fixed in advance, with two people screening independently where possible. Months to more than a year for a team of several people, often with a registered protocol. In ARS: `systematic-review` mode.
|
||||
> - **Scoping review**: maps what has been studied on a topic, the main concepts, and the gaps, using a systematic approach. The work grows with the breadth of the topic. ARS has no separate mode for it.
|
||||
> - **Narrative or integrative review**: builds an argument or a framework from the literature. In a narrative review the author chooses the sources; an integrative review documents its search and evaluation and can combine different study designs. The work depends on the scope the author sets. In ARS: `lit-review` mode.
|
||||
> - **Rapid review**: a systematic review with some steps shortened or left out to deliver sooner, typically within weeks to a few months. ARS has no separate mode for it.
|
||||
> - **No formal review**: background from the sources at hand, for example for an introduction. It does not claim to cover the literature.
|
||||
>
|
||||
> Reply with the form you want, or reply "skip" to continue as you are. Either reply is your decision, and this note will not appear again in this project.
|
||||
|
||||
**Note text (Traditional Chinese):**
|
||||
|
||||
> **開始回顧之前:要做哪一種文獻回顧?**
|
||||
> 這件事由你決定,ARS 不替你選。下列選項沒有預設、沒有推薦,排列順序也不代表高下。
|
||||
>
|
||||
> - **系統性回顧(systematic review)**:回答一個聚焦的問題,檢索與篩選方式事先訂好,盡可能由兩人各自獨立篩選。一個數人團隊需要數個月到一年以上,通常有已登錄的研究計畫書。ARS 對應:`systematic-review` 模式。
|
||||
> - **範疇回顧(scoping review)**:用系統化的做法盤點一個主題已經研究了什麼、有哪些主要概念、缺口在哪裡。工作量隨主題的廣度增加。ARS 沒有專屬模式。
|
||||
> - **敘事或整合性回顧(narrative / integrative review)**:從文獻建立論證或架構。敘事回顧由作者選擇文獻;整合性回顧會記錄檢索與評估過程,並可合併不同研究設計。工作量取決於作者設定的範圍。ARS 對應:`lit-review` 模式。
|
||||
> - **快速回顧(rapid review)**:為了早點交出結果而縮短或省略部分步驟的系統性回顧,通常在數週到數個月內完成。ARS 沒有專屬模式。
|
||||
> - **不做正式回顧**:用手邊的文獻寫背景,例如論文的緒論。不宣稱涵蓋整體文獻。
|
||||
>
|
||||
> 請回覆你要的形式,或回覆「跳過」照目前的做法繼續。兩種回覆都算你的決定,這個專案裡不會再出現這則提醒。
|
||||
<!-- review-form-note:end -->
|
||||
+7
-1
@@ -2,6 +2,7 @@ fixture_id: 02_single_phase_literature_only
|
||||
expected_routing_class: proceed
|
||||
expected_destination: academic-paper:lit-review
|
||||
escape_hatch_applied: false
|
||||
review_form_note: shown
|
||||
notes: |
|
||||
Materials: literature only (Phase 2). Intent: explicit ("Run lit-review on these papers")
|
||||
matches `lit-review` trigger keyword unambiguously.
|
||||
@@ -10,6 +11,11 @@ notes: |
|
||||
lit-review mode. No clarification needed — explicit intent overrides cross-phase check
|
||||
(and materials are single-phase anyway).
|
||||
|
||||
Since #921, selecting `lit-review` mode is a fixed point for the review-form note,
|
||||
so the response shows the note and waits before the lit-review work starts.
|
||||
Results recorded in CALIBRATION_LOG.md before #921 predate this criterion.
|
||||
|
||||
Pass criteria:
|
||||
- Response begins lit-review work (search strategy, screening, annotated bibliography)
|
||||
- Response enters lit-review mode and shows the English review-form note text
|
||||
verbatim, then waits for the author's reply
|
||||
- Response does NOT trigger clarification a-d options
|
||||
|
||||
@@ -2,10 +2,16 @@ fixture_id: 04_explicit_slash_command
|
||||
expected_routing_class: proceed
|
||||
expected_destination: academic-paper:lit-review
|
||||
escape_hatch_applied: false
|
||||
review_form_note: shown
|
||||
notes: |
|
||||
Slash command `/ars-lit-review` is an explicit invocation. Per Routing Discipline
|
||||
Step 1, route directly with no clarification, no orchestrator detour.
|
||||
|
||||
Since #921, selecting `lit-review` mode is a fixed point for the review-form note,
|
||||
so the response shows the note and waits before the lit-review work starts.
|
||||
Results recorded in CALIBRATION_LOG.md before #921 predate this criterion.
|
||||
|
||||
Pass criteria:
|
||||
- Response begins lit-review work
|
||||
- Response enters lit-review mode and shows the English review-form note text
|
||||
verbatim, then waits for the author's reply
|
||||
- Response does NOT clarify
|
||||
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
fixture_id: 13_lit_review_request_descriptive_question
|
||||
expected_routing_class: proceed
|
||||
expected_destination: academic-paper:lit-review
|
||||
accepted_destinations:
|
||||
- academic-paper:lit-review
|
||||
- deep-research:lit-review
|
||||
excluded_destinations:
|
||||
- deep-research:systematic-review
|
||||
- sr-screener
|
||||
escape_hatch_applied: false
|
||||
review_form_note: shown
|
||||
notes: |
|
||||
Negative case from issue #921: a request to write a literature review names no
|
||||
review form, so it routes to a `lit-review` mode and to neither
|
||||
`systematic-review` mode nor a screening skill. Selecting `lit-review` mode is
|
||||
fixed point 1, so the review-form note appears verbatim and the response waits
|
||||
for the author's reply before starting the review.
|
||||
|
||||
Fixture 14 is the same request with an effect question. Both must show the same
|
||||
note: whether it appears must not depend on the content of the question.
|
||||
|
||||
Pass criteria:
|
||||
- Response routes to a `lit-review` mode (either skill), not to
|
||||
`systematic-review` mode or screening
|
||||
- Response shows the English review-form note text verbatim, with nothing added
|
||||
that favours one form, and waits
|
||||
- Response does NOT trigger workflow clarification a-d options
|
||||
+1
@@ -0,0 +1 @@
|
||||
Help me write the literature review for my thesis on how teachers in rural schools describe burnout.
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
fixture_id: 14_lit_review_request_effect_question
|
||||
expected_routing_class: proceed
|
||||
expected_destination: academic-paper:lit-review
|
||||
accepted_destinations:
|
||||
- academic-paper:lit-review
|
||||
- deep-research:lit-review
|
||||
excluded_destinations:
|
||||
- deep-research:systematic-review
|
||||
- sr-screener
|
||||
escape_hatch_applied: false
|
||||
review_form_note: shown
|
||||
notes: |
|
||||
Content-independence pair of fixture 13 (issue #921). The research question is an
|
||||
effect question, the kind a model might think "fits" a systematic review. The
|
||||
route and the note must be the same as in fixture 13: the note appears because
|
||||
the author selected `lit-review` mode, not because of what the question asks.
|
||||
|
||||
Pass criteria:
|
||||
- Same as fixture 13
|
||||
- Response does NOT suggest, recommend, or lean toward a systematic review, and
|
||||
does NOT enter `systematic-review` mode on its own initiative
|
||||
+1
@@ -0,0 +1 @@
|
||||
Help me write the literature review for my thesis on whether mindfulness training reduces burnout among teachers in rural schools.
|
||||
@@ -0,0 +1,14 @@
|
||||
fixture_id: 15_systematic_review_named
|
||||
expected_routing_class: proceed
|
||||
expected_destination: deep-research:systematic-review
|
||||
escape_hatch_applied: false
|
||||
review_form_note: not_shown
|
||||
notes: |
|
||||
Issue #921: the author named the review form ("Run a systematic review"), so the
|
||||
review-form note does not appear. This test reads the author's words, not the
|
||||
content of the question; the question is the same as fixture 14's.
|
||||
|
||||
Pass criteria:
|
||||
- Response routes to `deep-research` `systematic-review` mode
|
||||
- Response does NOT show the review-form note
|
||||
- Response does NOT trigger workflow clarification a-d options
|
||||
@@ -0,0 +1 @@
|
||||
Run a systematic review on whether mindfulness training reduces burnout among teachers in rural schools.
|
||||
+22
@@ -0,0 +1,22 @@
|
||||
fixture_id: 16_broader_synthesis_not_systematic
|
||||
expected_routing_class: proceed
|
||||
expected_destination: deep-research:lit-review
|
||||
accepted_destinations:
|
||||
- deep-research:lit-review
|
||||
- academic-paper:lit-review
|
||||
excluded_destinations:
|
||||
- deep-research:systematic-review
|
||||
- sr-screener
|
||||
escape_hatch_applied: false
|
||||
review_form_note: shown
|
||||
notes: |
|
||||
Issue #921, from codex review round 1: a wish for a broader evidence matrix or a
|
||||
thematic synthesis after a three-way scan escalates to `lit-review`, never to
|
||||
`systematic-review` (deep-research/SKILL.md, Three-Way Scan Mode, last paragraph).
|
||||
Entering `lit-review` mode is fixed point 1, so the review-form note appears and
|
||||
the response waits; a systematic review starts only if the author then chooses it.
|
||||
|
||||
Pass criteria:
|
||||
- Response routes to a `lit-review` mode, not to `systematic-review` mode
|
||||
- Response shows the English review-form note text verbatim and waits
|
||||
- Response does NOT trigger workflow clarification a-d options
|
||||
@@ -0,0 +1,3 @@
|
||||
I ran a three-way scan on these 12 papers about peer feedback in writing courses. Now I want a broader evidence matrix and a thematic synthesis across them.
|
||||
|
||||
Papers attached in `~/papers/peer-feedback/`.
|
||||
+13
@@ -40,6 +40,10 @@ If you cannot reach 100% on the current primary model, the routing prose in CLAU
|
||||
| 10 | `10_korean_review_not_revision/` | Korean 심사 (referee) request + manuscript (#452) | **Proceed** → `academic-paper-reviewer:full` (not paper) |
|
||||
| 11 | `11_spanish_revision_not_review/` | Spanish enmendar (revise) request + draft (#856 es-ES) | **Proceed** → `academic-paper:revision` (not reviewer) |
|
||||
| 12 | `12_spanish_review_not_revision/` | Spanish revisar (referee) request + manuscript (#856 es-ES) | **Proceed** → `academic-paper-reviewer:full` (not paper) |
|
||||
| 13 | `13_lit_review_request_descriptive_question/` | "Help me write the literature review" on a descriptive question (#921) | **Proceed** → a `lit-review` mode, not `systematic-review` or screening; review-form note shown |
|
||||
| 14 | `14_lit_review_request_effect_question/` | Same request on an effect question (#921, content-independence pair of 13) | **Proceed** → same as 13; same note |
|
||||
| 15 | `15_systematic_review_named/` | "Run a systematic review" on the question of 14 (#921) | **Proceed** → `deep-research:systematic-review`; review-form note not shown |
|
||||
| 16 | `16_broader_synthesis_not_systematic/` | After a three-way scan, "a broader evidence matrix and a thematic synthesis" (#921) | **Proceed** → a `lit-review` mode, not `systematic-review`; review-form note shown |
|
||||
|
||||
## Fixture file format
|
||||
|
||||
@@ -61,8 +65,17 @@ expected_destination: <skill_name> | <agent_name> | clarification_only
|
||||
escape_hatch_applied: true | false
|
||||
direct_mode_stripped_message: <string, only when escape_hatch_applied: true>
|
||||
notes: <optional free-text>
|
||||
accepted_destinations: [<skill>:<mode>, ...] # optional; any listed destination passes
|
||||
excluded_destinations: [<skill>:<mode>, ...] # optional; entering any listed destination fails
|
||||
review_form_note: shown | not_shown # optional (#921)
|
||||
```
|
||||
|
||||
Scoring the #921 fields (added after the 2026-09-24 pass; the rules in `CALIBRATION_LOG.md` stay as they were fixed for that pass):
|
||||
|
||||
- **`review_form_note: shown`** passes when the response shows the review-form note text from `shared/references/review_form_note.md` verbatim in the conversation's language (English for every language other than Traditional Chinese), adds nothing that favours one form, and does not start the review before the author replies. The note is not a workflow clarification: showing it does not turn a `proceed` into a `clarify`.
|
||||
- **`review_form_note: not_shown`** passes when the note does not appear.
|
||||
- A fixture pair that differs only in the content of the question (13 and 14) passes only when both members pass; a note shown for one and not the other fails both.
|
||||
|
||||
### Running the smoke test (manual)
|
||||
|
||||
Until v3.10 conductor brings deterministic dispatch, these fixtures are run **manually against a live ARS session**:
|
||||
|
||||
Reference in New Issue
Block a user