docs(readme): correct front-page and showcase claims to match the evidence (#908)

* docs(readme): correct front-page and showcase claims to match the evidence

- Tagline: say that ARS can draft a whole paper in full mode, that the
  pipeline stops for confirmation at every stage, and that the user
  remains the author and answers for every submitted claim.
- Showcase: add a version note (academic-pipeline v2.3, March 2026; the
  current gates were not measured on this paper; why the byline names
  Claude); drop the 10-stage label; restate the Stage 2.5 counts as the
  report gives them (8 bibliographic errors, 6-8 likely fabricated);
  describe the post-publication audit as a Claude Code + WebSearch run
  that found problems in 21 of the 68 final references, 19 of which were
  removed or corrected; call the Round 1 panel role-separated reviews
  generated by Claude, not independent reviews.
- Calibration (README, ARCHITECTURE): the gold set has 25 synthetic
  tuples, the shipped test uses a stub judge that returns the gold
  labels, and no live-judge result is recorded.
- Pipeline section: "guarantees" becomes rules that are protocol, not
  runtime guarantees; Schema 11 maps each reviewer concern to the
  author's revision claim and records whether re-review verified it.
- Cost and time (QUICKSTART, README cost pointer): use the 2026-09
  list-price estimate from docs/PERFORMANCE.md (about US$3-7 per run,
  before cache discounts); collaborative work spans hours to days.

All six README locales carry the same changes, per the README sync rule
in CONTRIBUTING.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(readme): keep the audit row's translations as neutral as the English

The English row says 21 of the 68 final references still had problems
after three rounds of integrity checks. The zh-TW, zh-CN, ja-JP, and
ko-KR rows said the checks missed those problems, which is stronger:
several of the affected references had been flagged at Stage 2.5 and
marked FIXED before the audit found other errors in them. The four rows
now say the problems remained after the three rounds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): record the front-page claim corrections (#908)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Edward Cheng-I Wu
2026-09-25 17:08:39 +08:00
committed by GitHub
co-authored by Claude Opus 5.5
parent a3f65690a9
commit 1063d78f13
10 changed files with 67 additions and 51 deletions
+2
View File
@@ -32,6 +32,8 @@ All notable changes to this project will be documented in this file.
- The routing core now reaches plugin and skills-copy installs (#892). Claude Code loads `.claude/CLAUDE.md` only for sessions started inside the checkout and does not load a plugin's `CLAUDE.md`, so a plugin install ran without Routing Discipline Steps 0-3 and the #133 anti-pattern; in the #133 fixtures, a plugin-install session on Claude Opus 5.5 failed 6 of 12. The block now has one home, `shared/references/routing_core.md`: `.claude/CLAUDE.md` and every skill's `SKILL.md` carry byte-identical copies, the SessionStart announce reads the file at every session start, and `scripts/check_routing_core_sync.py` fails CI on any difference or on a skill without the block. After compaction, resume, or a fork, the announce's lead-in limits the core to a new request, so a run under way is not routed again; that wording is untested, because every fixture is a session's first message. After the fix, the plugin install matched the repo clone on all 12 fixtures in the Opus 5.5 sessions and on 11 of 12 in the Claude Fable 5.1 sessions. In both, fixture 04 runs on Sonnet 5 through its command's model pin, and fixture 05 passes in the plugin install only under two scoring readings the log records (a read the permission check refused counts as reading the agent file, and a paraphrased Skill argument counts as carrying the stripped message); scored strictly, it fails there on both models. On fixture 06, a `[direct-mode]` token placed mid-message, Fable 5.1 honored the token in 5 of 6 plugin-install sessions and in none of 6 repo-clone sessions, so #892 stays open for it. A skills-copy install gets the core at session start only when the project's `.claude/CLAUDE.md` carries the merged ARS text, as SETUP Method 1 instructs; otherwise it, like the other channels, gets the core only once a skill loads (`docs/CONTROL_AVAILABILITY.md` note 9). The announce reads the file with shell builtins, so the core survives a minimal `PATH` and a CRLF checkout, and the lint compares line endings too. `tests/fixtures/issue_133_routing/CALIBRATION_LOG.md` records the condition, the acceptance rule fixed before the run, and every result; one session per fixture is a smoke test, not a rate.
- Front-page and showcase statements now match their sources (#908). The README tagline says that full mode can draft a whole paper, that the pipeline stops for confirmation at every stage, and that the user remains the author. The showcase gains a version note (the run used academic-pipeline v2.3 in March 2026, the current citation gates have not been measured on that paper, and the byline names Claude because the researcher asked for it), restates the Stage 2.5 counts as the report gives them (8 bibliographic errors, 6–8 likely fabricated), describes the post-publication audit as a Claude Code + WebSearch run that found problems in 21 of the 68 final references (19 removed or corrected), and calls the Round 1 panel role-separated reviews generated by Claude rather than independent reviews. The claim-audit calibration text says the gold set has 25 synthetic tuples and the shipped test uses a stub judge, with no live-judge result recorded. "Pipeline guarantees" become rules that are protocol, not runtime guarantees, and Schema 11 is described by its documented purpose. `QUICKSTART.md` and the README cost pointer use the 2026-09 list-price estimate from `docs/PERFORMANCE.md`. All six README locales carry the changes.
## [3.22.1] - 2026-09-23 — Model currency for Claude Opus 5.5, citation-check loading and Chinese APA 7 repairs, and a Pi wrapper fix
> **Model currency and repairs; the new prompt-level guard is unmeasured:** v3.22.1 names Claude Opus 5.5 beside Claude Fable 5.1 as a supported session model, after a two-reader audit of the Opus 5.5 system card that retires no guardrail (#883). The docs add effort guidance (Claude Code starts Opus 5.5 at `medium`; heavy runs should use `high` or above), one list-price re-derivation for both models, and tiering guidance that reads the ladder as the vendor's lineup order, not a capability order. Because the card reports that Opus 5.5 acts more often than earlier models on instructions inside pasted text, the revision coach now treats pasted reviewer and committee text as data, pinned by a lint; the effect of that prompt-level guard is unmeasured. The release also repairs mode loading and citation checks: the 13 plugin mode commands invoke their namespaced core skill and resolve bundled references from the plugin root, which restores citation-check loading (#857); Chinese APA 7 checks catch a missing in-text author abbreviation, keep ambiguity exceptions and complete reference-list author fields, and propose a reorder only with evidence of a stroke-order inversion (#882); citation checks now distinguish visible syntax errors from unverified resolution or source claims (#882); and new English, Traditional Chinese, and Korean trigger phrases route citation-check requests, with CI bounding each skill description at 1,024 code points (#858, #864). The Pi wrapper accepts string-array system prompts (#880). No schema, command model, or effort setting changes.
+2 -2
View File
@@ -61,7 +61,7 @@ You: "I want to produce a complete research paper about how agentic AI
is reshaping student learning outcome measurement"
```
This triggers the full 10-stage pipeline. Budget ~$4-6 in API costs and 2-4 hours of collaborative work.
This triggers the full 10-stage pipeline, which stops at every stage for your confirmation. API cost is an estimated US$3–7 per run at 2026-09 list prices, before cache discounts ([docs/PERFORMANCE.md](docs/PERFORMANCE.md)). Expect the collaborative work to span hours to days, depending on how closely you review each stage.
## Which mode should I use?
@@ -73,7 +73,7 @@ This triggers the full 10-stage pipeline. Budget ~$4-6 in API costs and 2-4 hour
| Write a paper from scratch | `academic-paper` full mode |
| Plan a paper chapter by chapter | `academic-paper` plan mode |
| Get my paper reviewed | `academic-paper-reviewer` full mode |
| Do everything end-to-end | `academic-pipeline` — say "I want a complete research paper" |
| Go from question to finished paper, confirming each stage | `academic-pipeline` — say "I want a complete research paper" |
## What's next?
+9 -7
View File
@@ -18,7 +18,7 @@ Un conjunto completo de skills para Claude Code dedicadas a la investigación ac
Después prueba `/ars-plan` para revisar la estructura de tu artículo mediante diálogo socrático, o ve directamente a [Instalación rápida](#instalación-rápida) si necesitas los prerrequisitos y el flujo tradicional con enlaces simbólicos.
> **La IA es tu copiloto, no el piloto.** Esta herramienta no escribe tu artículo por ti. Se ocupa del trabajo pesado: buscar referencias, formatear citas, verificar datos, comprobar la coherencia lógica. Así puedes concentrarte en lo que de verdad requiere tu cabeza: definir la pregunta, elegir el método, interpretar qué significan los datos y escribir la frase que va después de «sostengo que».
> **La IA es tu copiloto, no el piloto.** Puede redactar borradores, incluido un artículo completo en el modo full, pero las decisiones son tuyas y el pipeline se detiene en cada etapa para que las confirmes. Se ocupa del trabajo pesado: buscar referencias, formatear citas, verificar datos, comprobar la coherencia lógica. Así puedes concentrarte en lo que de verdad requiere tu cabeza: definir la pregunta, elegir el método, interpretar qué significan los datos y decidir qué va después de «sostengo que». La autoría es tuya, y respondes de cada afirmación que envíes.
>
> A diferencia de un humanizador, esta herramienta no te ayuda a ocultar que has usado IA. Te ayuda a escribir mejor. Style Calibration aprende tu voz a partir de trabajos anteriores. Writing Quality Check detecta los patrones que hacen que un texto se sienta generado por una máquina. El objetivo es la calidad, no hacer trampa.
@@ -30,7 +30,7 @@ ARS parte de la premisa de que **un investigador humano aumentado por IA evita e
[**Zhao et al.**](https://arxiv.org/abs/2605.07723) (2026-05) auditó 111 M de referencias en 2,5 M de artículos de arXiv, bioRxiv, SSRN y PMC. Su estimación conservadora es de 146.932 citas alucinadas solo en 2025, con un punto de inflexión observado a mediados de 2024; para el emparejamiento bioRxiv-PMC reportan una persistencia del 85,3 % de preprint a publicación. El artículo describe como problema abierto el uso de «citas reales desplegadas para sostener afirmaciones que las referencias citadas no sostienen realmente». ARS v3.7.1 añadió trust-chain frontmatter para la procedencia de las fuentes; v3.7.3 añadió infraestructura de localizadores (anclas de cita en tres capas) para futuras auditorías a nivel de afirmación y muestra señales de riesgo advertidas en el momento de citar (ARS llama internamente «L3» a esa brecha de fidelidad entre afirmación y fuente; es terminología de ARS, no del artículo). v3.7.x responde a los hallazgos a escala de corpus de Zhao et al.; la evaluación a escala de corpus del propio ARS sigue siendo trabajo futuro.
v3.8 cierra la segunda mitad de la brecha L3. v3.7.3 hizo que cada cita llevara un ancla de localizador; v3.8 añade una pasada de auditoría opcional (`ARS_CLAIM_AUDIT=1`) que recupera la fuente citada contra cada ancla y juzga si la afirmación está realmente sostenida. Cinco nuevas clases HIGH-WARN (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) bloquean mediante gate la salida de la hard gate terminal del formatter. La calibración se publica como un gold set de 20 tuplas con umbrales de aceptación FNR<0,15 + FPR<0,10; el plan de ramp-on se aplaza hasta tener evidencia post-calibración, según la especificación de v3.8 §5.
v3.8 cierra la segunda mitad de la brecha L3. v3.7.3 hizo que cada cita llevara un ancla de localizador; v3.8 añade una pasada de auditoría opcional (`ARS_CLAIM_AUDIT=1`) que recupera la fuente citada contra cada ancla y juzga si la afirmación está realmente sostenida. Cinco nuevas clases HIGH-WARN (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) bloquean mediante gate la salida de la hard gate terminal del formatter. Un runner de calibración se publica con un gold set sintético de 25 tuplas y umbrales de aceptación FNR<0,15 + FPR<0,10. Su test incluido ejecuta el runner con un juez simulado que devuelve las etiquetas del gold set, así que comprueba la herramienta y no un juez real; todavía no hay registrado ningún resultado de calibración con un juez real, y el plan de ramp-on espera a tenerlo, según la especificación de v3.8 §5.
[**Ren et al.**](https://arxiv.org/abs/2607.13104) (2026, *Self-Improvements in Modern Agentic Systems: A Survey*) aporta una tercera ancla, a nivel de survey. Su síntesis sobre descubrimiento científico (§7.4) concluye que los agentes de descubrimiento no pueden verificar por sí mismos novedad, corrección o reproducibilidad y que pueden apoyarse en proxies débiles, que deben gestionar evidencia entre herramientas y literaturas heterogéneas y que plantean problemas de gobernanza: "la escritura científica también puede amplificar desinformación cuando la evidencia es débil". Sus capítulos sobre el bucle de generación (§5.1–§5.2) incluyen la auditoría humana y los anclas humanas conservadas entre las salvaguardas prácticas para bucles de evaluación autogenerados, y su capítulo histórico (§2.2) registra la forma más antigua de esa misma lección: el éxito práctico de EURISKO de Lenat dependía en gran medida de que el usuario actuara como señal externa de evaluación, podando la deriva improductiva de heurísticas; una limitación que el survey documenta como persistente en los sistemas agénticos modernos. ARS cita el survey como justificación de diseño de su postura de humano en el bucle, no como prueba empírica de que los pipelines con humano en el bucle superen a los autónomos; las mejoras accionables del survey para ARS se siguen en #539–#541 y #547–#550.
@@ -86,7 +86,7 @@ El documento de arquitectura sustituye a la extensa descripción del pipeline qu
## Rendimiento y coste
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — presupuestos de tokens por modo, estimación del pipeline completo (unos 4–6 $ por un artículo de 15k palabras) y ajustes recomendados de Claude Code (Auto mode; Agent Team opcional).
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — presupuestos de tokens por modo, estimación del pipeline completo (unos 3–7 US$ por un artículo de 15k palabras a precios de lista de 2026-09, antes de descuentos de caché) y ajustes recomendados de Claude Code (Auto mode; Agent Team opcional).
## Guías y artículos
@@ -115,7 +115,9 @@ El documento de arquitectura sustituye a la extensa descripción del pipeline qu
## Escaparate: salida real del pipeline
Consulta los artefactos completos de una ejecución real del pipeline de 10 etapas: informes de revisión por pares, informes de verificación de integridad y el artículo final.
Consulta los artefactos completos de una ejecución real del pipeline: informes de revisión por pares, informes de verificación de integridad y el artículo final.
> **Un registro de marzo de 2026, no el rendimiento actual.** Esta ejecución (del 2026-03-07 al 03-08) usó academic-pipeline v2.3, antes de que ARS añadiera la comprobación con Semantic Scholar en v3.3 y la puerta determinista de citas con cuatro índices en v3.11. Sus cifras describen esa versión, y las puertas actuales no se han medido con este artículo. La autoría figura a nombre de Claude (Anthropic) porque la persona investigadora lo pidió durante el experimento. El artículo no es una publicación de Anthropic, y la posición de ARS es que la herramienta no sustituye a quien investiga ni reclama la autoría (consulta [POSITIONING.md](POSITIONING.md#what-this-is-not)).
**[Ver todos los artefactos del pipeline →](examples/showcase/)**
@@ -123,13 +125,13 @@ Consulta los artefactos completos de una ejecución real del pipeline de 10 etap
|---|---|
| [Final Paper (EN)](examples/showcase/full_paper_apa7.pdf) | Formateado en APA 7.0, compilado con LaTeX |
| [Final Paper (ZH)](examples/showcase/full_paper_zh_apa7.pdf) | Versión en chino, APA 7.0 |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Etapa 2.5: detectó 15 referencias fabricadas + 3 errores estadísticos |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Etapa 2.5: señaló 15 referencias con problemas (8 con errores bibliográficos, 6–8 probablemente fabricadas) + 3 errores estadísticos |
| [Integrity Report — Final](examples/showcase/integrity_report_stage4.5.pdf) | Etapa 4.5: cero regresiones confirmadas |
| [Peer Review Round 1](examples/showcase/stage3_review_report.pdf) | Journal-Fit Reviewer + 3 revisores + Devil's Advocate |
| [Re-Review](examples/showcase/stage3prime_rereview_report.pdf) | Verificación tras las revisiones |
| [Peer Review Round 2](examples/showcase/stage3_review_report_r2.pdf) | Revisión de seguimiento |
| [Response to Reviewers](examples/showcase/response_to_reviewers_r2.pdf) | Respuesta puntual de los autores |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Auditoría independiente de todas las referencias: encontró 21 de 68 problemas que se escaparon en 3 rondas de comprobaciones de integridad |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Auditoría de todas las referencias hecha aparte con Claude Code + WebSearch: 21 de las 68 referencias finales seguían con problemas tras 3 rondas de comprobaciones de integridad |
---
@@ -275,7 +277,7 @@ Revisión multiperspectiva con 7 agentes y **juicios narrativos atados a criteri
### Academic Pipeline (v3.22.1)
Orquestador de 10 etapas con verificación de integridad, revisión en dos fases, coaching socrático y evaluación de la colaboración. Garantías del pipeline: cada etapa requiere un checkpoint de confirmación del usuario; la verificación de integridad (Etapa 2.5 + 4.5) es OBLIGATORIA y sin bypass no registrado (toda excepción requiere que quede registrada la justificación del usuario para la Etapa 6); la Matriz de Trazabilidad R&R (Schema 11) verifica de forma independiente las afirmaciones de revisión de los autores. v3.4 añadió el Compliance Agent (PRISMA-trAIce + RAISE) en las Etapas 2.5 / 4.5. v3.5 añade el **Collaboration Depth Observer** (`collaboration_depth_agent`, solo advisory, nunca bloquea) en cada checkpoint FULL/SLIM y al completar el pipeline. Las puertas de integridad OBLIGATORIAS (2.5 / 4.5) saltan explícitamente el observador para que las comprobaciones de cumplimiento no queden diluidas. Basado en Wang & Zhang (2026), IJETHE 23:11. Matriz etapa por etapa con agentes, artefactos y puertas: consulta ARCHITECTURE.md §3.
Orquestador de 10 etapas con verificación de integridad, revisión en dos fases, coaching socrático y evaluación de la colaboración. Reglas del pipeline (protocolo que siguen los agentes, no garantías en tiempo de ejecución): cada etapa requiere un checkpoint de confirmación del usuario; la verificación de integridad (Etapa 2.5 + 4.5) es OBLIGATORIA y sin bypass no registrado (toda excepción requiere que quede registrada la justificación del usuario para la Etapa 6); la Matriz de Trazabilidad R&R (Schema 11) vincula cada observación de la revisión con el cambio que declara el equipo autor y registra si la re-revisión lo verificó. v3.4 añadió el Compliance Agent (PRISMA-trAIce + RAISE) en las Etapas 2.5 / 4.5. v3.5 añade el **Collaboration Depth Observer** (`collaboration_depth_agent`, solo advisory, nunca bloquea) en cada checkpoint FULL/SLIM y al completar el pipeline. Las puertas de integridad OBLIGATORIAS (2.5 / 4.5) saltan explícitamente el observador para que las comprobaciones de cumplimiento no queden diluidas. Basado en Wang & Zhang (2026), IJETHE 23:11. Matriz etapa por etapa con agentes, artefactos y puertas: consulta ARCHITECTURE.md §3.
---
+9 -7
View File
@@ -18,7 +18,7 @@
その後、`/ars-plan` を試してソクラテス式対話で論文構成を整理するか、前提条件と従来のシンボリックリンク方式については [クイックインストール](#クイックインストール) を参照してください。
> **AI はあなたの副操縦士であり、操縦士ではありません。** このツールはあなたの代わりに論文を書きません。参考文献の探索、引用のフォーマット、データ検証、論理的整合性チェックといった泥臭い作業を引き受けることで、本当に頭を使う必要のある部分 — 問いの定義、手法の選択、データの意味の解釈、「私はこう主張する」に続く文を書くこと — にあなたが集中できるようにします。
> **AI はあなたの副操縦士であり、操縦士ではありません。** 文章の下書きはでき、full モードでは論文全体を下書きすることもあります。しかし判断を下すのはあなたで、パイプラインは各ステージであなたの確認を待ちます。参考文献の探索、引用のフォーマット、データ検証、論理的整合性チェックといった泥臭い作業を引き受けることで、本当に頭を使う必要のある部分(問いの定義、手法の選択、データの意味の解釈、「私はこう主張する」の後に何を続けるかの判断)にあなたが集中できるようにします。著者はあなたであり、提出するすべての主張に責任を負うのもあなたです。
>
> 「humanizer」とは異なり、このツールは AI を使った事実を隠すためのものではありません。より良い文章を書くための助けです。Style Calibration は過去の作品からあなたの声を学習します。Writing Quality Check は機械的に見える文章のパターンを検出します。目的は品質であって、ごまかしではありません。
@@ -30,7 +30,7 @@ ARS は **人間の研究者を AI が支援する形式が、どちらか単独
[**Zhao ら**](https://arxiv.org/abs/2605.07723)(2026-05)は arXiv、bioRxiv、SSRN、PMC の 2.5M 論文にわたる 111M 件の参考文献を監査しました。彼らの保守的見積りでは、2025年だけで 146,932 件のハルシネーション引用が観測され、2024年中頃に変曲点が観測されています。bioRxiv-to-PMC ペアリングでは、プレプリントから出版物への持続率は 85.3% と報告されています。論文は「引用された参考文献が実際には主張していない主張を支持するために配置された実在の引用」を未解決の課題として記述しています。ARS v3.7.1 はソース来歴のための trust-chain frontmatter を追加し、v3.7.3 は将来の主張レベル監査のためのロケーターインフラストラクチャ(三層引用アンカー)を追加し、引用時に advisory リスクシグナルを表面化します(ARS は主張忠実性ギャップを内部で「L3」とラベル付けしています。これは論文の用語ではなく ARS の用語です)。v3.7.x は Zhao らのコーパス規模の発見に動機付けられています。ARS 自体のコーパス規模評価は今後の課題として残されています。
v3.8 は L3 ギャップの後半を閉じます。v3.7.3 は全引用にロケーターアンカーを持たせ、v3.8 はオプトインの監査パス(`ARS_CLAIM_AUDIT=1`)を追加します。これは各アンカーに対して引用元を取得し、主張が実際に裏付けられているかを判断します。5 つの新しい HIGH-WARN クラス(claim-not-supported、negative-constraint-violation、fabricated-reference、anchorless、constraint-violation-uncited)は、formatter ターミナルハードゲートを通じて出力を gate-refuse します。キャリブレーションは 20-tuple のゴールドセットと共に FNR<0.15 + FPR<0.10 の受容閾値で出荷されます。ramp-on 計画は v3.8 spec §5 に従いキャリブレーション後の証拠まで保留されます。
v3.8 は L3 ギャップの後半を閉じます。v3.7.3 は全引用にロケーターアンカーを持たせ、v3.8 はオプトインの監査パス(`ARS_CLAIM_AUDIT=1`)を追加します。これは各アンカーに対して引用元を取得し、主張が実際に裏付けられているかを判断します。5 つの新しい HIGH-WARN クラス(claim-not-supported、negative-constraint-violation、fabricated-reference、anchorless、constraint-violation-uncited)は、formatter ターミナルハードゲートを通じて出力を gate-refuse します。キャリブレーションランナーは 25-tuple の合成ゴールドセットと FNR<0.15 + FPR<0.10 の受容閾値と共に出荷されます。同梱のテストはゴールドラベルをそのまま返すスタブのジャッジでランナーを動かすため、検証しているのはツールであり、実際のジャッジではありません。実ジャッジによるキャリブレーション結果はまだ記録されておらず、ramp-on 計画は v3.8 spec §5 に従いその結果を待ちます。
v3.3 は [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pfister & Yoon, 2026, Google)に触発されました: Semantic Scholar API 検証、アンチリーケージプロトコル、VLM 図表検証、改訂軌跡追跡。ARS の現行実装は、数値デルタではなく、基準ごとの証拠に基づくナラティブな退行チェックを行います。型付き軌跡キャリアは未実装です。
@@ -73,7 +73,7 @@ v3.3 は [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pf
## パフォーマンス&コスト
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — モードごとのトークン予算、フルパイプライン見積り(15k 語の論文で約 $4-6)、推奨 Claude Code 設定(Auto モード; Agent Team オプション)。
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — モードごとのトークン予算、フルパイプライン見積り(15k 語の論文で、2026-09 の定価で約 US$3〜7、キャッシュ割引前)、推奨 Claude Code 設定(Auto モード; Agent Team オプション)。
## ガイド&記事
@@ -100,7 +100,9 @@ v3.3 は [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pf
## ショーケース: 実際のパイプライン出力
実際の 10 ステージパイプライン実行からの完全な成果物を参照してください — ピアレビューレポート、整合性検証レポート、最終論文:
実際のパイプライン実行からの完全な成果物(ピアレビューレポート、整合性検証レポート、最終論文)を参照してください:
> **2026 年 3 月の記録であり、現在の性能ではありません。** この実行(2026-03-07〜03-08)は academic-pipeline v2.3 によるもので、ARS が v3.3 で Semantic Scholar 照合を、v3.11 で決定論的な 4 インデックス引用ゲートを追加する前のものです。ここの数値はその版を表しており、現行のゲートはこの論文ではまだ測定されていません。著者欄が Claude(Anthropic)になっているのは、研究者がこの実験中にそう依頼したためです。この論文は Anthropic の出版物ではありません。ARS の位置づけは、ツールは研究者に取って代わらず、著者であるとも主張しないというものです([POSITIONING.md](POSITIONING.md#what-this-is-not) を参照)。
**[すべてのパイプライン成果物を見る →](examples/showcase/)**
@@ -108,13 +110,13 @@ v3.3 は [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pf
|---|---|
| [Final Paper (EN)](examples/showcase/full_paper_apa7.pdf) | APA 7.0 フォーマット、LaTeX コンパイル済み |
| [Final Paper (ZH)](examples/showcase/full_paper_zh_apa7.pdf) | 中国語版、APA 7.0 |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: 捏造参照 15 件 + 統計エラー 3 件を捕捉 |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: 問題のある参照 15 件(書誌エラー 8 件、捏造の疑い 6〜8 件)+ 統計エラー 3 件を指摘 |
| [Integrity Report — Final](examples/showcase/integrity_report_stage4.5.pdf) | Stage 4.5: ゼロリグレッションを確認 |
| [Peer Review Round 1](examples/showcase/stage3_review_report.pdf) | Journal-Fit Reviewer + 3 Reviewers + Devil's Advocate |
| [Re-Review](examples/showcase/stage3prime_rereview_report.pdf) | 改訂後の検証 |
| [Peer Review Round 2](examples/showcase/stage3_review_report_r2.pdf) | フォローアップレビュー |
| [Response to Reviewers](examples/showcase/response_to_reviewers_r2.pdf) | ポイントごとの著者回答 |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | 独立した完全参照監査: 3 回の整合性チェックで見逃された 21/68 件の問題を発見 |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Claude Code + WebSearch で別途行った全参照監査: 3 回の整合性チェック後も、最終的な 68 件の参照のうち 21 件に問題が残っていた |
---
@@ -254,7 +256,7 @@ You: "status"
### Academic Pipeline(v3.22.1)
整合性検証、二段階レビュー、ソクラテス式コーチング、コラボレーション評価を持つ 10 ステージのオーケストレーター。パイプライン保証: 各ステージにユーザー確認チェックポイントが必要。整合性検証(Stage 2.5 + 4.5)は MANDATORY であり、記録されないバイパス経路は存在しない(すべてのオーバーライドは Stage 6 のためにユーザーの理由の記録を要する)。R&R Traceability Matrix(Schema 11)は著者の改訂主張を独立に検証する。v3.4 は Stage 2.5 / 4.5 に Compliance Agent(PRISMA-trAIce + RAISE)を追加した。v3.5 はすべての FULL/SLIM チェックポイントとパイプライン完了時に **Collaboration Depth Observer**(`collaboration_depth_agent`、advisory のみ — 決してブロックしない)を追加する。MANDATORY 整合性ゲート(2.5 / 4.5)は、コンプライアンスチェックが希薄化されないよう observer を明示的にスキップする。Wang & Zhang(2026), IJETHE 23:11 に基づく。エージェント、成果物、ゲートを含むステージごとのマトリクス: ARCHITECTURE.md §3 を参照。
整合性検証、二段階レビュー、ソクラテス式コーチング、コラボレーション評価を持つ 10 ステージのオーケストレーター。パイプラインのルール(エージェントが従うプロトコルであり、実行時の保証ではない): 各ステージにユーザー確認チェックポイントが必要。整合性検証(Stage 2.5 + 4.5)は MANDATORY であり、記録されないバイパス経路は存在しない(すべてのオーバーライドは Stage 6 のためにユーザーの理由の記録を要する)。R&R Traceability Matrix(Schema 11)は各査読コメントを著者の改訂主張に対応づけ、再審査でそれが検証されたかどうかを記録する。v3.4 は Stage 2.5 / 4.5 に Compliance Agent(PRISMA-trAIce + RAISE)を追加した。v3.5 はすべての FULL/SLIM チェックポイントとパイプライン完了時に **Collaboration Depth Observer**(`collaboration_depth_agent`、advisory のみ — 決してブロックしない)を追加する。MANDATORY 整合性ゲート(2.5 / 4.5)は、コンプライアンスチェックが希薄化されないよう observer を明示的にスキップする。Wang & Zhang(2026), IJETHE 23:11 に基づく。エージェント、成果物、ゲートを含むステージごとのマトリクス: ARCHITECTURE.md §3 を参照。
---
+9 -7
View File
@@ -18,7 +18,7 @@
그런 다음 `/ars-plan`을 실행해 소크라테스식 대화로 논문 구조를 짜보거나, 사전 요건과 전통적인 심볼릭 링크 방식을 보려면 [빠른 설치](#빠른-설치)로 이동하세요.
> **AI는 부조종사이지 조종사가 아닙니다.** 이 도구는 논문을 대신 써 주지 않습니다. 참고문헌 탐색, 인용 형식 정리, 데이터 검증, 논리적 일관성 점검과 같은 반복적이고 소모적인 작업을 지원하여, 실제로 사람의 판단이 필요한 부분 — 질문 정의, 방법 선택, 데이터가 의미하는 바의 해석, 그리고 "나는 ~라고 주장한다" 다음에 오는 문장을 쓰는 일 — 에 집중할 수 있게 합니다.
> **AI는 부조종사이지 조종사가 아닙니다.** 이 도구는 글의 초안을 쓸 수 있고, full 모드에서는 논문 전체의 초안도 씁니다. 하지만 결정은 사용자가 내리며, 파이프라인은 모든 단계에서 사용자의 확인을 기다립니다. 참고문헌 탐색, 인용 형식 정리, 데이터 검증, 논리적 일관성 점검과 같은 반복적이고 소모적인 작업을 지원하여, 실제로 사람의 판단이 필요한 부분(질문 정의, 방법 선택, 데이터가 의미하는 바의 해석, 그리고 "나는 ~라고 주장한다" 다음에 무엇을 쓸지 정하는 일)에 집중할 수 있게 합니다. 저자는 사용자이며, 제출하는 모든 주장에 대한 책임도 사용자에게 있습니다.
>
> 휴머나이저(humanizer)와 달리, 이 도구는 AI를 사용했다는 사실을 숨기도록 돕지 않습니다. 더 잘 쓰도록 돕습니다. Style Calibration은 과거 작업에서 사용자의 문체를 학습합니다. Writing Quality Check는 기계가 생성한 듯한 느낌을 주는 패턴을 잡아냅니다. 목표는 품질이지 부정행위가 아닙니다.
@@ -30,7 +30,7 @@ ARS는 **AI의 지원을 받는 인간 연구자가 인간이나 AI가 단독으
[**Zhao et al.**](https://arxiv.org/abs/2605.07723) (2026-05)은 arXiv, bioRxiv, SSRN, PMC의 250만 편 논문에 걸친 1억 1,100만 건의 참고문헌을 대규모로 점검했습니다. 이들의 보수적 추정치는 2025년 한 해에만 146,932건의 환각된 인용이며, 2024년 중반에 변곡점이 관찰되었습니다. bioRxiv-to-PMC 쌍에 대해서는 85.3%의 preprint-to-published 지속성을 보고합니다. 이 논문은 "인용된 참고문헌이 실제로는 뒷받침하지 않는 주장을 지지하기 위해 배치된 진짜 인용"을 미해결 과제로 기술합니다. ARS v3.7.1은 출처 provenance를 위한 trust-chain frontmatter를 추가했고, v3.7.3은 향후 주장 수준 감사를 위한 locator 인프라(3계층 인용 앵커)를 추가하고 인용 시점에 참고용 위험 신호를 표시합니다(ARS는 이 주장-충실성 격차를 내부적으로 "L3"로 라벨링합니다. 이는 ARS 용어이며 논문의 용어가 아닙니다). v3.7.x는 Zhao et al.의 코퍼스 규모 발견에 동기를 두며, ARS 자체에 대한 코퍼스 규모 평가는 향후 과제로 남아 있습니다.
v3.8은 L3 격차의 나머지 절반을 메웁니다. v3.7.3은 모든 인용이 locator 앵커를 갖도록 했고, v3.8은 각 앵커에 대해 인용된 출처를 가져와 주장이 실제로 뒷받침되는지 판단하는 옵트인 감사 패스(`ARS_CLAIM_AUDIT=1`)를 추가합니다. 다섯 개의 새로운 HIGH-WARN 클래스(claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited)의 출력을 formatter terminal hard gate가 거부합니다. 캘리브레이션은 FNR<0.15 + FPR<0.10 합격 임계값을 갖는 20개 항목으로 구성된 골드셋으로 제공됩니다. 단계적 활성화 계획은 v3.8 명세 §5에 따라 캘리브레이션 후 증거가 나올 때까지 보류됩니다.
v3.8은 L3 격차의 나머지 절반을 메웁니다. v3.7.3은 모든 인용이 locator 앵커를 갖도록 했고, v3.8은 각 앵커에 대해 인용된 출처를 가져와 주장이 실제로 뒷받침되는지 판단하는 옵트인 감사 패스(`ARS_CLAIM_AUDIT=1`)를 추가합니다. 다섯 개의 새로운 HIGH-WARN 클래스(claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited)의 출력을 formatter terminal hard gate가 거부합니다. 캘리브레이션 러너는 25개 항목의 합성 골드셋과 FNR<0.15 + FPR<0.10 합격 임계값과 함께 제공됩니다. 함께 제공되는 테스트는 골드 레이블을 그대로 돌려주는 스텁 심판으로 러너를 실행하므로, 검증 대상은 도구 자체이지 실제 심판이 아닙니다. 실제 심판의 캘리브레이션 결과는 아직 기록되지 않았으며, 단계적 활성화 계획은 v3.8 명세 §5에 따라 그 결과를 기다립니다.
v3.3은 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfister & Yoon, 2026, Google)에서 영감을 받았습니다: Semantic Scholar API 검증, anti-leakage 프로토콜, VLM 그림 검증, 수정 궤적 추적. 현재 ARS는 수치 델타 대신 기준별 증거 기반 서술형 회귀 점검을 수행하며, typed trajectory carrier는 아직 구현되지 않았습니다.
@@ -74,7 +74,7 @@ v3.3은 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfis
## 성능 & 비용
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — 모드별 토큰 예산, 전체 파이프라인 추정치(15,000 단어 논문 기준 약 $4–6), 권장 Claude Code 설정(Auto 모드. Agent Team 선택).
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — 모드별 토큰 예산, 전체 파이프라인 추정치(15,000 단어 논문 기준, 2026-09 정가로 약 US$3–7, 캐시 할인 전), 권장 Claude Code 설정(Auto 모드. Agent Team 선택).
## 가이드 & 글
@@ -101,7 +101,9 @@ v3.3은 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfis
## 쇼케이스: 실제 파이프라인 출력
실제 10단계 파이프라인 실행에서 나온 완전한 산출물 — 동료 심사 보고서, 무결성 검증 보고서, 최종 논문 — 을 확인하세요:
실제 파이프라인 실행에서 나온 완전한 산출물(동료 심사 보고서, 무결성 검증 보고서, 최종 논문)을 확인하세요:
> **2026년 3월의 기록이며, 현재 성능이 아닙니다.** 이 실행(2026-03-07~03-08)은 academic-pipeline v2.3으로 이루어졌으며, ARS가 v3.3에서 Semantic Scholar 대조를, v3.11에서 결정론적 4개 인덱스 인용 게이트를 추가하기 전입니다. 여기의 수치는 그 버전을 나타내며, 현재 게이트는 이 논문에서 아직 측정되지 않았습니다. 저자란이 Claude(Anthropic)로 되어 있는 것은 연구자가 이 실험 중에 그렇게 요청했기 때문입니다. 이 논문은 Anthropic의 출판물이 아니며, ARS는 도구가 연구자를 대신하지 않고 저자권을 주장하지도 않는다는 입장입니다([POSITIONING.md](POSITIONING.md#what-this-is-not) 참조).
**[모든 파이프라인 산출물 둘러보기 →](examples/showcase/)**
@@ -109,13 +111,13 @@ v3.3은 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfis
|---|---|
| [Final Paper (EN)](examples/showcase/full_paper_apa7.pdf) | APA 7.0 형식, LaTeX 컴파일 |
| [Final Paper (ZH)](examples/showcase/full_paper_zh_apa7.pdf) | 중국어 버전, APA 7.0 |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: 날조된 참고문헌 15건 + 통계 오류 3건 적발 |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: 문제 있는 참고문헌 15건(서지 오류 8건, 날조 의심 6~8건) + 통계 오류 3건 지적 |
| [Integrity Report — Final](examples/showcase/integrity_report_stage4.5.pdf) | Stage 4.5: 회귀 없음 확인 |
| [Peer Review Round 1](examples/showcase/stage3_review_report.pdf) | Journal-Fit Reviewer + 리뷰어 3명 + Devil's Advocate |
| [Re-Review](examples/showcase/stage3prime_rereview_report.pdf) | 수정 후 검증 |
| [Peer Review Round 2](examples/showcase/stage3_review_report_r2.pdf) | 후속 심사 |
| [Response to Reviewers](examples/showcase/response_to_reviewers_r2.pdf) | 항목별 저자 응답 |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | 독립적 전체 참고문헌 감사: 3회의 무결성 점검이 놓친 21/68 문제 발견 |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Claude Code + WebSearch로 별도 수행한 전체 참고문헌 감사: 3회의 무결성 점검 후에도 최종 참고문헌 68건 중 21건에 문제가 남아 있음 |
---
@@ -261,7 +263,7 @@ You: "status"
### Academic Pipeline (v3.22.1)
무결성 검증, 2단계 심사, 소크라테스식 코칭, 협업 평가를 갖춘 10단계 오케스트레이터. 파이프라인 보장: 모든 단계는 사용자 확인 체크포인트를 요구하며, 무결성 검증(Stage 2.5 + 4.5)은 MANDATORY이며 기록 없는 우회 경로가 없고(모든 오버라이드는 Stage 6를 위해 사용자 사유 기록을 요구), R&R Traceability Matrix(Schema 11)는 저자의 수정 주장을 독립적으로 검증합니다. v3.4는 Stage 2.5 / 4.5에 Compliance Agent(PRISMA-trAIce + RAISE)를 추가했습니다. v3.5는 모든 FULL/SLIM 체크포인트와 파이프라인 완료 시점에 **Collaboration Depth Observer**(`collaboration_depth_agent`, 자문 전용 — 절대 차단하지 않음)를 추가합니다. 필수(MANDATORY) 무결성 게이트(2.5 / 4.5)는 컴플라이언스 점검이 희석되지 않도록 observer를 명시적으로 건너뜁니다. Wang & Zhang (2026), IJETHE 23:11에 기반합니다. 에이전트·산출물·게이트를 포함한 단계별 매트릭스: ARCHITECTURE.md §3 참조.
무결성 검증, 2단계 심사, 소크라테스식 코칭, 협업 평가를 갖춘 10단계 오케스트레이터. 파이프라인 규칙(에이전트가 따르는 프로토콜이며 실행 시 보장이 아님): 모든 단계는 사용자 확인 체크포인트를 요구하며, 무결성 검증(Stage 2.5 + 4.5)은 MANDATORY이며 기록 없는 우회 경로가 없고(모든 오버라이드는 Stage 6를 위해 사용자 사유 기록을 요구), R&R Traceability Matrix(Schema 11)는 각 심사 의견을 저자의 수정 주장에 대응시키고 재심사에서 그것이 검증되었는지를 기록합니다. v3.4는 Stage 2.5 / 4.5에 Compliance Agent(PRISMA-trAIce + RAISE)를 추가했습니다. v3.5는 모든 FULL/SLIM 체크포인트와 파이프라인 완료 시점에 **Collaboration Depth Observer**(`collaboration_depth_agent`, 자문 전용 — 절대 차단하지 않음)를 추가합니다. 필수(MANDATORY) 무결성 게이트(2.5 / 4.5)는 컴플라이언스 점검이 희석되지 않도록 observer를 명시적으로 건너뜁니다. Wang & Zhang (2026), IJETHE 23:11에 기반합니다. 에이전트·산출물·게이트를 포함한 단계별 매트릭스: ARCHITECTURE.md §3 참조.
---
+9 -7
View File
@@ -18,7 +18,7 @@ A comprehensive suite of Claude Code skills for academic research, covering the
Then try `/ars-plan` to walk through your paper structure via Socratic dialogue, or jump to [Quick install](#quick-install) for prerequisites and the traditional symlink flow.
> **AI is your copilot, not the pilot.** This tool won't write your paper for you. It handles the grunt work — hunting down references, formatting citations, verifying data, checking logical consistency — so you can focus on the parts that actually require your brain: defining the question, choosing the method, interpreting what the data means, and writing the sentence after "I argue that."
> **AI is your copilot, not the pilot.** It can draft text, including a whole paper in full mode, but the decisions stay yours, and the pipeline stops for your confirmation at every stage. It handles the grunt work (hunting down references, formatting citations, verifying data, checking logical consistency) so you can focus on the parts that actually require your brain: defining the question, choosing the method, interpreting what the data means, and deciding what comes after "I argue that." You remain the author, and you answer for every claim you submit.
>
> Unlike a humanizer, this tool doesn't help you hide the fact that you used AI. It helps you write better. Style Calibration learns your voice from past work. Writing Quality Check catches the patterns that make prose feel machine-generated. The goal is quality, not cheating.
@@ -30,7 +30,7 @@ ARS is built on the premise that **a human researcher augmented by AI avoids the
[**Zhao et al.**](https://arxiv.org/abs/2605.07723) (2026-05) audited 111M references across 2.5M papers on arXiv, bioRxiv, SSRN, and PMC. Their conservative estimate is 146,932 hallucinated citations for 2025 alone, with an observed mid-2024 inflection; for the bioRxiv-to-PMC pairing they report 85.3% preprint-to-published persistence. The paper describes "real citations deployed to support claims the cited references do not actually make" as an open challenge. ARS v3.7.1 added trust-chain frontmatter for source provenance; v3.7.3 added locator infrastructure (three-layer citation anchors) for future claim-level audits and surfaces advisory risk signals at cite time (ARS labels the claim-faithfulness gap internally as "L3"; this is ARS terminology, not the paper's). v3.7.x is motivated by Zhao et al.'s corpus-scale findings; corpus-scale evaluation of ARS itself remains future work.
v3.8 closes the second half of the L3 gap. v3.7.3 made every citation carry a locator anchor; v3.8 adds an opt-in audit pass (`ARS_CLAIM_AUDIT=1`) that fetches the cited source against each anchor and judges whether the claim is actually supported. Five new HIGH-WARN classes (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) gate-refuse output through the formatter terminal hard gate. Calibration is shipped as a 20-tuple gold set with FNR<0.15 + FPR<0.10 acceptance thresholds; ramp-on plan is deferred to post-calibration evidence per v3.8 spec §5.
v3.8 closes the second half of the L3 gap. v3.7.3 made every citation carry a locator anchor; v3.8 adds an opt-in audit pass (`ARS_CLAIM_AUDIT=1`) that fetches the cited source against each anchor and judges whether the claim is actually supported. Five new HIGH-WARN classes (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) gate-refuse output through the formatter terminal hard gate. A calibration runner ships with a 25-tuple synthetic gold set and FNR<0.15 + FPR<0.10 acceptance thresholds. Its shipped test drives the runner with a stub judge that returns the gold labels, so it checks the tooling, not a live judge; no live-judge calibration result is recorded yet, and the ramp-on plan waits for one (v3.8 spec §5).
[**Ren et al.**](https://arxiv.org/abs/2607.13104) (2026, *Self-Improvements in Modern Agentic Systems: A Survey*) supplies a third, survey-level anchor. Its scientific-discovery synthesis (§7.4) concludes that discovery agents cannot easily verify novelty, correctness, or reproducibility on their own and may exploit weak proxies instead, must manage evidence across heterogeneous tools and literature, and raise governance issues — "scientific writing can also amplify misinformation when the evidence is weak." Its generation-loop chapters (§5.1–§5.2) list human auditing and retained human anchors among the practical safeguards for self-generated evaluation loops, and its historical chapter (§2.2) records the oldest form of the same lesson: the practical success of Lenat's EURISKO depended heavily on the user serving as the external evaluation signal, pruning unproductive heuristic drift — a limitation the survey notes persists in modern agentic systems. ARS cites the survey as design rationale for its human-in-the-loop stance, not as empirical proof that human-in-the-loop pipelines outperform autonomous ones; the survey's actionable deltas for ARS are tracked in #539–#541 and #547–#550.
@@ -86,7 +86,7 @@ The architecture doc supersedes the sprawling pipeline description that used to
## Performance & cost
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — per-mode token budgets, full-pipeline estimate (~$4–6 for a 15k-word paper), and recommended Claude Code settings (Auto mode; Agent Team optional).
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — per-mode token budgets, full-pipeline estimate (about US$3–7 for a 15k-word paper at 2026-09 list prices, before cache discounts), and recommended Claude Code settings (Auto mode; Agent Team optional).
## Guides & articles
@@ -115,7 +115,9 @@ The architecture doc supersedes the sprawling pipeline description that used to
## Showcase: real pipeline output
See the complete artifacts from a real 10-stage pipeline run — peer review reports, integrity verification reports, and the final paper:
See the complete artifacts from a real pipeline run, including peer review reports, integrity verification reports, and the final paper:
> **A March 2026 record, not current performance.** This run (2026-03-07 to 03-08) used academic-pipeline v2.3, before ARS v3.3 added the Semantic Scholar check and v3.11 added the deterministic four-index citation gate. Its numbers describe that version; the current gates have not been measured on this paper. The byline names Claude (Anthropic) as author because the researcher asked for that during the experiment. The paper is not an Anthropic publication, and ARS's positioning is that the tool does not replace the researcher and does not claim authorship ([POSITIONING.md](POSITIONING.md#what-this-is-not)).
**[Browse all pipeline artifacts →](examples/showcase/)**
@@ -123,13 +125,13 @@ See the complete artifacts from a real 10-stage pipeline run — peer review rep
|---|---|
| [Final Paper (EN)](examples/showcase/full_paper_apa7.pdf) | APA 7.0 formatted, LaTeX-compiled |
| [Final Paper (ZH)](examples/showcase/full_paper_zh_apa7.pdf) | Chinese version, APA 7.0 |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: caught 15 fabricated refs + 3 statistical errors |
| [Integrity Report — Pre-Review](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5: flagged 15 problem references (8 with bibliographic errors, 6–8 likely fabricated) + 3 statistical errors |
| [Integrity Report — Final](examples/showcase/integrity_report_stage4.5.pdf) | Stage 4.5: zero regressions confirmed |
| [Peer Review Round 1](examples/showcase/stage3_review_report.pdf) | Journal-Fit Reviewer + 3 Reviewers + Devil's Advocate |
| [Re-Review](examples/showcase/stage3prime_rereview_report.pdf) | Verification after revisions |
| [Peer Review Round 2](examples/showcase/stage3_review_report_r2.pdf) | Follow-up review |
| [Response to Reviewers](examples/showcase/response_to_reviewers_r2.pdf) | Point-by-point author response |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Independent full-reference audit: found 21/68 issues missed by 3 rounds of integrity checks |
| [Post-Publication Audit Report](examples/showcase/post_publication_audit_2026-03-09.pdf) | Full-reference audit run separately with Claude Code + WebSearch: 21 of the 68 final references still had problems after 3 rounds of integrity checks |
---
@@ -275,7 +277,7 @@ Per-agent responsibilities and per-stage artifacts now live in [`docs/ARCHITECTU
### Academic Pipeline (v3.22.1)
10-stage orchestrator with integrity verification, two-stage review, Socratic coaching, and collaboration evaluation. Pipeline guarantees: every stage requires user confirmation checkpoint; integrity verification (Stage 2.5 + 4.5) is MANDATORY with no unrecorded bypass (every override requires user reasoning recorded for Stage 6); R&R Traceability Matrix (Schema 11) independently verifies author revision claims. v3.4 added the Compliance Agent (PRISMA-trAIce + RAISE) at Stage 2.5 / 4.5. v3.5 adds the **Collaboration Depth Observer** (`collaboration_depth_agent`, advisory only — never blocks) at every FULL/SLIM checkpoint and at pipeline completion. MANDATORY integrity gates (2.5 / 4.5) explicitly skip the observer so compliance checks are not diluted. Based on Wang & Zhang (2026), IJETHE 23:11. Stage-by-stage matrix with agents, artifacts, and gates: see ARCHITECTURE.md §3.
10-stage orchestrator with integrity verification, two-stage review, Socratic coaching, and collaboration evaluation. Pipeline rules (protocol the agents follow, not runtime guarantees): every stage requires user confirmation checkpoint; integrity verification (Stage 2.5 + 4.5) is MANDATORY with no unrecorded bypass (every override requires user reasoning recorded for Stage 6); R&R Traceability Matrix (Schema 11) maps each reviewer concern to the author's revision claim and records whether the re-review verified it. v3.4 added the Compliance Agent (PRISMA-trAIce + RAISE) at Stage 2.5 / 4.5. v3.5 adds the **Collaboration Depth Observer** (`collaboration_depth_agent`, advisory only — never blocks) at every FULL/SLIM checkpoint and at pipeline completion. MANDATORY integrity gates (2.5 / 4.5) explicitly skip the observer so compliance checks are not diluted. Based on Wang & Zhang (2026), IJETHE 23:11. Stage-by-stage matrix with agents, artifacts, and gates: see ARCHITECTURE.md §3.
---
+9 -7
View File
@@ -18,7 +18,7 @@
安装后运行 `/ars-plan`,ARS 会用苏格拉底式对话帮你规划章节结构。需要前置条件或传统 symlink 安装,请看 [快速安装](#快速安装)。
> **AI 是你的副驾驶,不是机长。** 这个工具不会替你写论文。它处理繁琐工作:搜文献、排格式、验数据、查逻辑一致性。这样你就能专注在真正需要思考的事上:定义问题、选择方法、解读数据意义、写出「我认为」后面那句话。
> **AI 是你的副驾驶,不是机长。** 它可以起草文字,full mode 甚至会起草整篇论文;但决定权在你,pipeline 每个阶段都会停下来等你确认。它处理繁琐工作(搜文献、排格式、验数据、查逻辑一致性),让你专注在真正需要思考的事上:定义问题、选择方法、解读数据意义、决定「我认为」后面要接什么。作者是你,提交的每一个主张都由你负责。
>
> 和 humanizer 不同,这个工具不是帮你隐藏使用 AI 协作的事实,而是帮你把关文章质量。风格校准会从你过去的文章中学习你的声音,写作质量检查会识别让文字读起来像机器生成的模式。目标是质量,不是掩饰。
@@ -30,7 +30,7 @@ ARS 建立在这个前提上:**人类研究者 + AI 的组合,比纯自动
[**Zhao 等人**](https://arxiv.org/abs/2605.07723)(2026-05)盘点了 arXiv、bioRxiv、SSRN、PMC 上 250 万篇论文中的 1.11 亿条引用,保守估计 2025 年单年就有 146,932 条幻觉引用,并观察到 2024 年中是上升的拐点;bioRxiv-to-PMC 这条配对的「预印本进入正式发表版本」幻觉存活率达 85.3%。他们把「真实引用被用来支撑被引文献其实没有提出的主张」描述为当前未解的问题。ARS v3.7.1 为来源 provenance 加上 trust-chain frontmatter,v3.7.3 为未来的 claim-level 审计铺设 locator 基础设施(三层引用 anchor),并在引用阶段呈现 advisory 风险信号(ARS 内部把这条 claim-faithfulness 缺口标记为「L3」,此为 ARS 的用词,不是论文的用词)。v3.7.x 的设计动机来自 Zhao 等人的 corpus-scale 发现;ARS 本身的 corpus-scale 评估仍是未来工作。
v3.8 补上 L3 缺口的另一半。v3.7.3 让每一条引用都带 locator anchor,v3.8 在这个基础上加一道 opt-in 审计(`ARS_CLAIM_AUDIT=1`):获取每个 anchor 指向的原始文本,判断论文里的 claim 是否真有被该引用支撑。五类新的 HIGH-WARN annotation(claim-not-supported、negative-constraint-violation、fabricated-reference、anchorless、constraint-violation-uncited)会在 formatter terminal hard gate 直接阻止输出。Calibration 随 release 提供 20 条 gold set,采用 FNR<0.15、FPR<0.10 双阈值;正式放大投入前要先有 calibration 证据(v3.8 spec §5)。
v3.8 补上 L3 缺口的另一半。v3.7.3 让每一条引用都带 locator anchor,v3.8 在这个基础上加一道 opt-in 审计(`ARS_CLAIM_AUDIT=1`):获取每个 anchor 指向的原始文本,判断论文里的 claim 是否真有被该引用支撑。五类新的 HIGH-WARN annotation(claim-not-supported、negative-constraint-violation、fabricated-reference、anchorless、constraint-violation-uncited)会在 formatter terminal hard gate 直接阻止输出。Calibration runner 随 release 附一组 25 条的合成 gold set,采用 FNR<0.15、FPR<0.10 双阈值。随附的测试使用直接返回标准答案的替身裁判,所以它验证的是工具本身,不是真正的 AI 裁判;目前还没有真实裁判的 calibration 结果,正式放大投入要等这份证据(v3.8 spec §5)。
v3.3 的灵感来自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pfister & Yoon, 2026, Google):Semantic Scholar API 验证、反泄露协议、VLM 图表验证、修订轨迹追踪。ARS 当前以分类式、证据锚定的准则轨迹实现最后一项,不计算分数差。
@@ -73,7 +73,7 @@ v3.3 的灵感来自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(So
## 性能与费用
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — 各模式 token 预算、完整 pipeline 估算(一篇 15k 字论文约 ~$4–6),以及建议的 Claude Code 设置(Auto 模式;Agent Team 选用)。
**👉 [docs/PERFORMANCE.md](docs/PERFORMANCE.md)** — 各模式 token 预算、完整 pipeline 估算(一篇 15k 字论文,按 2026-09 牌价约 US$3–7,未计 cache 折扣),以及建议的 Claude Code 设置(Auto 模式;Agent Team 选用)。
## 使用指南与文章
@@ -100,7 +100,9 @@ v3.3 的灵感来自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(So
## 实际产出展示
查看完整 10 阶段 pipeline 的实际产出 — 包含**同行评审报告、学术诚信验证报告、完稿论文**:
查看一次完整 pipeline 运行的实际产出,包含**同行评审报告、学术诚信验证报告、完稿论文**:
> **这是 2026 年 3 月的记录,不代表现在的表现。** 这次运行(2026-03-07 至 03-08)使用的是 academic-pipeline v2.3,当时 ARS 还没有在 v3.3 加入 Semantic Scholar 核对,也还没有在 v3.11 加入确定性的四索引引用闸门。这里的数字描述的是那个版本,现行闸门还没有在这篇论文上测量过。作者栏写 Claude(Anthropic),是因为研究者在这次实验中这样要求。这篇论文不是 Anthropic 的出版物;ARS 的定位是工具不取代研究者,也不主张作者身份(见 [POSITIONING.md](POSITIONING.md#what-this-is-not))。
**[浏览所有 pipeline 产出 →](examples/showcase/)**
@@ -108,13 +110,13 @@ v3.3 的灵感来自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(So
|---|---|
| [完稿论文(英文)](examples/showcase/full_paper_apa7.pdf) | APA 7.0 格式,LaTeX 编译 |
| [完稿论文(中文)](examples/showcase/full_paper_zh_apa7.pdf) | 中文版,APA 7.0 |
| [学术诚信报告 — 审稿前](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5:发现 15 个虚构引用 + 3 个统计错误 |
| [学术诚信报告 — 审稿前](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5:标出 15 条有问题的引用(8 条书目错误、6–8 条疑似虚构)+ 3 个统计错误 |
| [学术诚信报告 — 最终](examples/showcase/integrity_report_stage4.5.pdf) | Stage 4.5:确认零回归 |
| [同行评审第一轮](examples/showcase/stage3_review_report.pdf) | Journal-Fit Reviewer + 3 审查者 + 魔鬼代言人 |
| [再审](examples/showcase/stage3prime_rereview_report.pdf) | 修订后验证审查 |
| [同行评审第二轮](examples/showcase/stage3_review_report_r2.pdf) | 跟踪审查 |
| [回复审查意见](examples/showcase/response_to_reviewers_r2.pdf) | 逐点回复 |
| [出版后审计报告](examples/showcase/post_publication_audit_2026-03-09.pdf) | 独立全引用审计:发现 21/68 篇问题,在 3 轮学术诚信审查后仍被漏掉 |
| [出版后审计报告](examples/showcase/post_publication_audit_2026-03-09.pdf) | 另行用 Claude Code + WebSearch 审计全部引用:经过 3 轮学术诚信审查后,最后 68 条引用中仍有 21 条有问题 |
---
@@ -254,7 +256,7 @@ ARS Stage 2 写作 → 用验证过的实验结果撰写论文
### Academic Pipeline (v3.22.1)
10 阶段调度器,含学术诚信验证、两阶段审查、苏格拉底指导、协作质量评估。Pipeline 保证:每个阶段都需用户确认 checkpoint;学术诚信验证(Stage 2.5 + 4.5)为 MANDATORY 且没有不留记录的绕过路径(所有覆写都须记录用户理由、供 Stage 6 使用);R&R 追溯矩阵(Schema 11)独立验证作者修订主张。v3.4 添加 Compliance Agent(PRISMA-trAIce + RAISE)于 Stage 2.5 / 4.5。v3.5 添加 **协作深度观察员**(`collaboration_depth_agent`,仅咨询性质、永不阻挡流程)于每一次 FULL/SLIM checkpoint 与 pipeline 完成时。MANDATORY 学术诚信闸门(2.5 / 4.5)明确跳过观察员,避免稀释合规检查。理论基础:Wang & Zhang (2026), IJETHE 23:11。逐阶段矩阵(agent、产出物、闸门):见 ARCHITECTURE.md §3。
10 阶段调度器,含学术诚信验证、两阶段审查、苏格拉底指导、协作质量评估。Pipeline 规则(由 agent 按流程遵守,不是运行时保证):每个阶段都需用户确认 checkpoint;学术诚信验证(Stage 2.5 + 4.5)为 MANDATORY 且没有不留记录的绕过路径(所有覆写都须记录用户理由、供 Stage 6 使用);R&R 追溯矩阵(Schema 11)把每一项审查意见对应到作者的修订主张,并记录复审是否验证通过。v3.4 添加 Compliance Agent(PRISMA-trAIce + RAISE)于 Stage 2.5 / 4.5。v3.5 添加 **协作深度观察员**(`collaboration_depth_agent`,仅咨询性质、永不阻挡流程)于每一次 FULL/SLIM checkpoint 与 pipeline 完成时。MANDATORY 学术诚信闸门(2.5 / 4.5)明确跳过观察员,避免稀释合规检查。理论基础:Wang & Zhang (2026), IJETHE 23:11。逐阶段矩阵(agent、产出物、闸门):见 ARCHITECTURE.md §3。
---
+9 -7
View File
@@ -18,7 +18,7 @@
裝完跑 `/ars-plan`,ARS 會用蘇格拉底對話幫你規劃章節結構。需要前置條件或傳統 symlink 安裝請看 [快速安裝](#快速安裝)。
> **AI 是你的副駕駛,不是機長。** 這工具不會幫你寫論文。它處理苦工 — 搜文獻、排格式、驗數據、查邏輯一致性 — 讓你專注在真正需要你腦子的事:定義問題、選方法、詮釋數據的意義、寫出「我認為」後面那句話。
> **AI 是你的副駕駛,不是機長。** 它可以起草文字,full mode 甚至會起草整篇論文;但決定權在你,pipeline 每個階段都會停下來等你確認。它處理苦工(搜文獻、排格式、驗數據、查邏輯一致性),讓你專注在真正需要你腦子的事:定義問題、選方法、詮釋數據的意義、決定「我認為」後面要接什麼。作者是你,送出的每一個主張都由你負責。
>
> 跟 humanizer 不同,這工具不是幫你隱藏用 AI 協作的事實,而是幫你把關文章品質。風格校準從你過去的文章學習你的聲音,寫作品質檢查抓出讓文字讀起來像機器產的模式。目標是品質,不是遮掩。
@@ -30,7 +30,7 @@ ARS 建立在這個前提上:**人類研究者 + AI 的組合,比純自動
[**Zhao 等人**](https://arxiv.org/abs/2605.07723)(2026-05)盤點了 arXiv、bioRxiv、SSRN、PMC 上 250 萬篇論文裡的 1.11 億筆引用,保守估計 2025 年單年就有 146,932 筆幻覺引用,並觀察到 2024 年中是上升的拐點;bioRxiv-to-PMC 這條配對的「預印本進到正式發表」幻覺存活率達 85.3%。他們把「真實引用被用來支撐被引文獻其實沒有提出的主張」描述為當前未解的問題。ARS v3.7.1 為來源 provenance 加上 trust-chain frontmatter,v3.7.3 為未來的 claim-level 稽核鋪上 locator 基礎建設(三層引用 anchor),並在引用時段帶出 advisory 風險訊號(ARS 內部把這條 claim-faithfulness 缺口標記為「L3」,此為 ARS 的用詞,不是論文的用詞)。v3.7.x 的設計動機來自 Zhao 等人的 corpus-scale 發現;ARS 本身的 corpus-scale 評估仍是未來工作。
v3.8 補上 L3 缺口的另一半。v3.7.3 讓每一筆引用都帶 locator anchor,v3.8 在這個基礎上加一道 opt-in 稽核(`ARS_CLAIM_AUDIT=1`):抓回每一個 anchor 指向的原始文本,判斷論文裡的 claim 是否真有被該引用支撐。五類新的 HIGH-WARN annotation(claim-not-supported、negative-constraint-violation、fabricated-reference、anchorless、constraint-violation-uncited)會在 formatter terminal hard gate 直接攔下輸出。Calibration 隨 release 出 20 筆 gold set,採 FNR<0.15、FPR<0.10 雙閾值;正式放大投入前要先有 calibration 證據(v3.8 spec §5)。
v3.8 補上 L3 缺口的另一半。v3.7.3 讓每一筆引用都帶 locator anchor,v3.8 在這個基礎上加一道 opt-in 稽核(`ARS_CLAIM_AUDIT=1`):抓回每一個 anchor 指向的原始文本,判斷論文裡的 claim 是否真有被該引用支撐。五類新的 HIGH-WARN annotation(claim-not-supported、negative-constraint-violation、fabricated-reference、anchorless、constraint-violation-uncited)會在 formatter terminal hard gate 直接攔下輸出。Calibration runner 隨 release 附一組 25 筆的合成 gold set,採 FNR<0.15、FPR<0.10 雙閾值。隨附的測試用的是直接回傳標準答案的替身裁判,所以它驗證的是工具本身,不是真正的 AI 裁判;目前還沒有真裁判的 calibration 結果,正式放大投入要等這份證據(v3.8 spec §5)。
[**Ren 等人**](https://arxiv.org/abs/2607.13104)(2026,*Self-Improvements in Modern Agentic Systems: A Survey*)補上第三個、survey 層級的錨點。其科學發現章節的綜合結論(§7.4)指出:發現型 agent 難以自行驗證 novelty、正確性與可重現性,反而可能鑽弱代理指標的漏洞;證據管理必須跨異質工具與文獻維持;並帶有治理疑慮——「證據薄弱時,科學寫作也會放大錯誤資訊」。其生成迴圈章節(§5.1–§5.2)把人工稽核與保留人類標註列為自生成評估迴圈的實務防護;歷史章節(§2.2)則記下同一課題最早的版本:Lenat 的 EURISKO 的實務成功高度依賴使用者充當外部評估訊號、修剪無效的 heuristic 漂移——survey 明言此限制延續到現代 agentic 系統。ARS 引用這篇 survey 作為 human-in-the-loop 立場的設計依據,而非「人機協作必然勝過全自動」的實證證明;survey 對 ARS 可落地的增量記錄在 #539–#541 與 #547–#550。
@@ -79,7 +79,7 @@ v3.3 的靈感來自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(So
## 效能與費用
**👉 [docs/PERFORMANCE.zh-TW.md](docs/PERFORMANCE.zh-TW.md)** — 各模式 token 預算、完整 pipeline 估算(一篇 15k 字論文約 ~$4–6),以及建議的 Claude Code 設定(Auto 模式;Agent Team 選用)。
**👉 [docs/PERFORMANCE.zh-TW.md](docs/PERFORMANCE.zh-TW.md)** — 各模式 token 預算、完整 pipeline 估算(一篇 15k 字論文,以 2026-09 牌價計約 US$3–7,未計 cache 折扣),以及建議的 Claude Code 設定(Auto 模式;Agent Team 選用)。
## 使用指南與文章
@@ -106,7 +106,9 @@ v3.3 的靈感來自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(So
## 實際產出展示
查看完整 10 階段 pipeline 的實際產出 — 包含**同儕審查報告、誠信驗證報告、完稿論文**:
查看一次完整 pipeline 執行的實際產出,包含**同儕審查報告、誠信驗證報告、完稿論文**:
> **這是 2026 年 3 月的紀錄,不代表現在的表現。** 這次執行(2026-03-07 至 03-08)用的是 academic-pipeline v2.3,當時 ARS 還沒有在 v3.3 加入 Semantic Scholar 核對,也還沒有在 v3.11 加入確定性的四索引引用閘門。這裡的數字描述的是那個版本,現行閘門還沒有在這篇論文上量過。作者欄寫 Claude(Anthropic),是因為研究者在這次實驗中這樣要求。這篇論文不是 Anthropic 的出版品;ARS 的定位是工具不取代研究者,也不主張作者身分(見 [POSITIONING.md](POSITIONING.md#what-this-is-not))。
**[瀏覽所有 pipeline 產出 →](examples/showcase/)**
@@ -114,13 +116,13 @@ v3.3 的靈感來自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(So
|---|---|
| [完稿論文(英文)](examples/showcase/full_paper_apa7.pdf) | APA 7.0 格式,LaTeX 編譯 |
| [完稿論文(中文)](examples/showcase/full_paper_zh_apa7.pdf) | 中文版,APA 7.0 |
| [誠信報告 — 審稿前](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5:抓出 15 個虛構引用 + 3 個統計錯誤 |
| [誠信報告 — 審稿前](examples/showcase/integrity_report_stage2.5.pdf) | Stage 2.5:標出 15 筆有問題的引用(8 筆書目錯誤、6–8 筆疑似虛構)+ 3 個統計錯誤 |
| [誠信報告 — 最終](examples/showcase/integrity_report_stage4.5.pdf) | Stage 4.5:確認零回歸 |
| [同儕審查第一輪](examples/showcase/stage3_review_report.pdf) | Journal-Fit Reviewer + 3 審查者 + 魔鬼代言人 |
| [複審](examples/showcase/stage3prime_rereview_report.pdf) | 修訂後驗證審查 |
| [同儕審查第二輪](examples/showcase/stage3_review_report_r2.pdf) | 追蹤審查 |
| [回覆審查意見](examples/showcase/response_to_reviewers_r2.pdf) | 逐點回覆 |
| [出版後稽核報告](examples/showcase/post_publication_audit_2026-03-09.pdf) | 獨立全引用稽核:發現 21/68 篇問題,通過了 3 輪誠信審查仍漏網 |
| [出版後稽核報告](examples/showcase/post_publication_audit_2026-03-09.pdf) | 另行以 Claude Code + WebSearch 稽核全部引用:經過 3 輪誠信審查後,最後 68 筆引用中仍有 21 筆有問題 |
---
@@ -260,7 +262,7 @@ ARS Stage 2 寫作 → 用驗證過的實驗結果撰寫論文
### Academic Pipeline (v3.22.1)
10 階段調度器,含誠信驗證、兩階段審查、蘇格拉底指導、協作品質評估。Pipeline 保證:每個階段都需使用者確認 checkpoint;誠信驗證(Stage 2.5 + 4.5)為 MANDATORY 且沒有不留紀錄的繞過路徑(所有覆寫都須記錄使用者理由、供 Stage 6 使用);R&R 追溯矩陣(Schema 11)獨立驗證作者修訂宣稱。v3.4 新增 Compliance Agent(PRISMA-trAIce + RAISE)於 Stage 2.5 / 4.5。v3.5 新增 **協作深度觀察員**(`collaboration_depth_agent`,僅諮詢性質、永不阻擋流程)於每一次 FULL/SLIM checkpoint 與 pipeline 完成時。MANDATORY 誠信閘門(2.5 / 4.5)明確跳過觀察員,避免稀釋合規檢查。理論基礎:Wang & Zhang (2026), IJETHE 23:11。逐階段矩陣(agent、產出物、閘門):見 ARCHITECTURE.md §3。
10 階段調度器,含誠信驗證、兩階段審查、蘇格拉底指導、協作品質評估。Pipeline 規則(由 agent 依流程遵守,不是執行期保證):每個階段都需使用者確認 checkpoint;誠信驗證(Stage 2.5 + 4.5)為 MANDATORY 且沒有不留紀錄的繞過路徑(所有覆寫都須記錄使用者理由、供 Stage 6 使用);R&R 追溯矩陣(Schema 11)把每一項審查意見對應到作者的修訂宣稱,並記錄複審是否驗證通過。v3.4 新增 Compliance Agent(PRISMA-trAIce + RAISE)於 Stage 2.5 / 4.5。v3.5 新增 **協作深度觀察員**(`collaboration_depth_agent`,僅諮詢性質、永不阻擋流程)於每一次 FULL/SLIM checkpoint 與 pipeline 完成時。MANDATORY 誠信閘門(2.5 / 4.5)明確跳過觀察員,避免稀釋合規檢查。理論基礎:Wang & Zhang (2026), IJETHE 23:11。逐階段矩陣(agent、產出物、閘門):見 ARCHITECTURE.md §3。
---
+1 -1
View File
@@ -107,7 +107,7 @@ flowchart TD
| **3' → 4' Residual Coaching** | `academic-paper-reviewer` (Journal-Fit Reviewer Socratic sub-stage) | VERIFIED_ONLY | Residual-issue dialogue | eic_agent | 🧑 **Decision-heavy checkpoint:** Socratic dialogue on trade-offs for residual issues (max 5 rounds). User may skip. Source: `two_stage_review_protocol.md` |
| **4'. RE-REVISE** | `academic-paper` v3.3.1 (revision) | REDACTED | Final Revised Draft (terminal; advances to 4.5) | draft_writer_agent; revision_coach_agent; **👁 collaboration_depth_agent (v3.5.0, advisory)** | 🧑 **Decision-heavy checkpoint:** user confirms content frozen. No further review loop permitted. 👁 Observer runs post-checkpoint; never blocks |
| **4.5 FINAL INTEGRITY** | `academic-pipeline` v3.22.1 (gate) | VERIFIED_ONLY | Updated Material Passport (`verification_status: VERIFIED`) + `repro_lock` declared — populated or explicit `null` (honest opt-out); Claim Verification Report (**final-check mode: 100% of E1 registered claims; semantic extraction completeness unknown** per `claim_verification_protocol.md`) | integrity_verification_agent (deeper re-run of 7 modes); state_tracker_agent. **👁 collaboration_depth_agent: SKIPPED (MANDATORY gate — observer dilution explicitly prevented)** | ✓ **Integrity gate** + user ack. Zero named gate defects within the registered/sampled populations; no skip permitted. Any mode SUSPECTED at 2.5 must be CLEAR or user-Overridden by 4.5. `repro_lock` is **not** read by the integrity gate at runtime (per `artifact_reproducibility_pattern.md`); if populated, `stochasticity_declaration` must be verbatim and is validated by the standalone `check_repro_lock.py` — this is post-hoc documentation, not a runtime block or global correctness certificate |
| **4→5 CLAIM-AUDIT** (v3.8, opt-in via `ARS_CLAIM_AUDIT=1`) | `academic-pipeline` v3.22.1 (gate) | VERIFIED_ONLY | `claim_audit_results[]` + `claim_drifts[]` + `uncited_assertions[]` + `constraint_violations[]` + `audit_sampling_summaries[]` aggregates; reads `claim_intent_manifests[]` (writer-side pre-commitment baseline). Emits 5 HIGH-WARN annotation classes consumed by Stage 5 formatter REFUSE rules 6-10 | claim_ref_alignment_audit_agent (Stage 4→5 dispatch slot, after v3.7.1 cite finalizer, before formatter hard gate) | ✓ **Audit gate** (default OFF for v3.8.0). Per-citation LLM-as-judge against retrieved excerpt; 8-row finalizer matrix discriminates paywall (LOW-WARN) / fabricated (HIGH-WARN) / anchorless (HIGH-WARN) / audit_tool_failure (MED-WARN) via `ref_retrieval_method`. Calibration runner (`scripts/test_claim_audit_calibration.py`) gates with FNR<0.15 + FPR<0.10 on the shipped 20-tuple gold set. Spec: `docs/design/2026-05-15-issue-103-claim-alignment-audit-spec.md` |
| **4→5 CLAIM-AUDIT** (v3.8, opt-in via `ARS_CLAIM_AUDIT=1`) | `academic-pipeline` v3.22.1 (gate) | VERIFIED_ONLY | `claim_audit_results[]` + `claim_drifts[]` + `uncited_assertions[]` + `constraint_violations[]` + `audit_sampling_summaries[]` aggregates; reads `claim_intent_manifests[]` (writer-side pre-commitment baseline). Emits 5 HIGH-WARN annotation classes consumed by Stage 5 formatter REFUSE rules 6-10 | claim_ref_alignment_audit_agent (Stage 4→5 dispatch slot, after v3.7.1 cite finalizer, before formatter hard gate) | ✓ **Audit gate** (default OFF for v3.8.0). Per-citation LLM-as-judge against retrieved excerpt; 8-row finalizer matrix discriminates paywall (LOW-WARN) / fabricated (HIGH-WARN) / anchorless (HIGH-WARN) / audit_tool_failure (MED-WARN) via `ref_retrieval_method`. Calibration runner (`scripts/test_claim_audit_calibration.py`) checks the FNR<0.15 + FPR<0.10 gate on the shipped 25-tuple gold set with a stub judge that returns the gold labels, so it validates the tooling; no live-judge calibration result is recorded. Spec: `docs/design/2026-05-15-issue-103-claim-alignment-audit-spec.md` |
| **5. FINALIZE** | `academic-paper` v3.3.1 (format-convert / disclosure) | VERIFIED_ONLY | Publication-ready MD; DOCX (Pandoc, if available); LaTeX (user confirms); PDF (tectonic); default venue AI applicability/status bundle (`REQUIRED` / `ACTION_ONLY` / `NOT_REQUIRED` / `UNKNOWN` plus typed halt) or policy-anchor-specific render | formatter_agent | 🧑 **Decision-heavy checkpoint:** user selects format before render. The disclosure output must match the selected venue or policy anchor (15-entry venue database: ICLR / NeurIPS / Nature / Science / ACL / EMNLP + medical-publishing targets incl. ICMJE, NEJM, The Lancet, JAMA, BMJ, PLOS, Frontiers, and two Chinese-language policy targets — one publisher-wide and one journal; see `venue_disclosure_policies.md`). **v3.8 terminal hard gate (formatter_agent REFUSE rules 6-10)** refuses output on any unresolved `[HIGH-WARN-CLAIM-NOT-SUPPORTED]` / `[HIGH-WARN-NEGATIVE-CONSTRAINT-VIOLATION]` / `[HIGH-WARN-FABRICATED-REFERENCE]` / `[HIGH-WARN-CLAIM-AUDIT-ANCHORLESS]` / `[HIGH-WARN-CONSTRAINT-VIOLATION-UNCITED]` annotation when `ARS_CLAIM_AUDIT=1` was set upstream. **v3.10 rule 11** refuses on any `severity=HIGH-BLOCK` terminal-policy token (generic; co-emitted by the finalizer under a strict `terminal_policies` mode). **v3.11 rule 12 (#182)** refuses on a `lookup_verified == false` citation-existence row ONLY under `terminal_policies.citation_existence == strict` — default advisory passes (`/ars-mark-read`-ack-able); the narrowed ID-keyed `false` never fires on a title-only-unmatched `unresolvable` citation |
| **6. PROCESS SUMMARY** | `academic-pipeline` v3.22.1 | VERIFIED_ONLY | Paper Creation Process Record (MD + PDF); AI Self-Reflection Report (concession rate, sycophancy risk, health alerts, Failure Mode Audit Log); narrative criterion-regression notes when recorded; **Collaboration Depth Chapter (v3.5.0)** summarising the per-checkpoint observer reports from `collaboration_depth_history[]` | state_tracker_agent; pipeline_orchestrator_agent; **👁 collaboration_depth_agent (v3.5.0, pipeline-completion dispatch — final advisory report)** | 🧑 **Decision-heavy checkpoint:** language confirmed with user. Collaboration quality evaluated. No typed criterion-trajectory visualization is claimed until its producer/validator is implemented. Post-publication audit report (if peer-review published). 👁 Observer runs final pipeline-completion dispatch; never blocks |
+8 -6
View File
@@ -1,9 +1,11 @@
# Showcase: Full Academic Pipeline Output
Complete artifacts from a real 10-stage academic pipeline run, demonstrating the end-to-end quality assurance process.
Complete artifacts from a real academic pipeline run, demonstrating the end-to-end quality assurance process.
**Paper**: *From Snapshots to Trajectories: How Agentic AI Will Redefine Student Learning Outcomes and Transform Student Success Measurement — Implications for Taiwan's Next Cycle of Institutional Accreditation*
> **Version note.** This run (2026-03-07 to 03-08) used academic-pipeline v2.3, before ARS v3.3 added the Semantic Scholar check and v3.11 added the deterministic four-index citation gate. The numbers below describe that version; the current gates have not been measured on this paper. The byline names Claude (Anthropic) as author because the researcher asked for that during the experiment. The paper is not an Anthropic publication, and ARS's positioning is that the tool does not replace the researcher and does not claim authorship ([POSITIONING.md](../../POSITIONING.md#what-this-is-not)).
## Paper Outputs
| File | Description |
@@ -15,16 +17,16 @@ Complete artifacts from a real 10-stage academic pipeline run, demonstrating the
| File | Stage | Verdict |
|------|-------|---------|
| [integrity_report_stage2.5.pdf](integrity_report_stage2.5.pdf) | Pre-Review (Stage 2.5) | FAIL — found 15 fabricated refs + 3 statistical errors |
| [integrity_report_stage2.5.pdf](integrity_report_stage2.5.pdf) | Pre-Review (Stage 2.5) | FAIL — flagged 15 problem references (8 with bibliographic errors, 6–8 likely fabricated) + 3 statistical errors |
| [integrity_reverification_stage2.5.pdf](integrity_reverification_stage2.5.pdf) | Re-verification (Stage 2.5) | PASS — all 22 issues fixed |
| [integrity_report_stage4.5.pdf](integrity_report_stage4.5.pdf) | Post-Revision (Stage 4.5) | PASS — zero regressions, 3 new refs verified |
| [post_publication_audit_2026-03-09.pdf](post_publication_audit_2026-03-09.pdf) | Post-Publication Audit | 21/68 refs corrected — manual WebSearch stress test |
| [post_publication_audit_2026-03-09.pdf](post_publication_audit_2026-03-09.pdf) | Post-Publication Audit | 21/68 refs had problems, 19 removed or corrected — Claude Code + WebSearch audit, run separately from paper generation |
## Peer Review Reports
| File | Description |
|------|-------------|
| [stage3_review_report.pdf](stage3_review_report.pdf) | Round 1: Journal-Fit Reviewer + 3 Reviewers + Devil's Advocate (5 independent reviews) |
| [stage3_review_report.pdf](stage3_review_report.pdf) | Round 1: Journal-Fit Reviewer + 3 Reviewers + Devil's Advocate (5 role-separated reviews, all generated by Claude; not independent reviewers) |
| [stage3prime_rereview_report.pdf](stage3prime_rereview_report.pdf) | Re-review after Round 1 revisions |
| [stage3_review_report_r2.pdf](stage3_review_report_r2.pdf) | Round 2: Follow-up review |
| [response_to_reviewers_r2.pdf](response_to_reviewers_r2.pdf) | Author's point-by-point response to reviewers |
@@ -51,9 +53,9 @@ Stage 4: Final Revision
Stage 4.5: Final Integrity Verification ← MANDATORY, confirmed 0 regressions
Stage 5: Finalization (LaTeX → PDF, bilingual)
Stage 6: Process Summary
Post-Publication: Manual WebSearch audit of all 68 refs → 21 issues found & fixed
Post-Publication: Claude Code + WebSearch audit of all 68 refs → 21 issues found, 19 removed or corrected
```
## Post-Publication Audit (2026-03-09)
After the pipeline completed, a manual WebSearch audit of all 68 references revealed 21 issues (31%) that survived three rounds of automated integrity checking. This stress test led to the integrity verification agent v2.0 overhaul (Anti-Hallucination Mandate, gray-zone elimination, known hallucination pattern library). See [post_publication_audit_2026-03-09.pdf](post_publication_audit_2026-03-09.pdf) for the full report.
After the pipeline completed, Claude Code with WebSearch, run separately from paper generation, audited all 68 references and found problems in 21 of them (31% of the reference list) that had survived three rounds of automated integrity checking: 4 not found (likely fabricated), 6 with wrong authors, 7 with wrong metadata (year, title, pages, journal, or DOI), and 4 minor format or title issues. Afterward, the 4 were removed from the paper and 15 of the other 17 were corrected. The audit PDF calls the 31% a false-negative rate; it is the share of final references with a problem, not a miss rate over all problem references. At least seven of the 21 (cited as Banihashem, El-Banna, Gandara, Kestin, Stanford SCALE, Tao, and Temper) had been flagged at Stage 2.5 and marked FIXED at re-verification, yet the audit still found errors in them, such as wrong authors or an altered title. This stress test led to the integrity verification agent v2.0 overhaul (Anti-Hallucination Mandate, gray-zone elimination, known hallucination pattern library). See [post_publication_audit_2026-03-09.pdf](post_publication_audit_2026-03-09.pdf) for the full report.