mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-12/22): document and diagram understanding three eras
This commit is contained in:
@@ -0,0 +1,95 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 520" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.reg { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
|
||||
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="480" y="24" text-anchor="middle" class="title">Document AI — three eras from OCR pipeline to VLM-native</text>
|
||||
|
||||
<rect x="30" y="50" width="900" height="210" class="box"/>
|
||||
|
||||
<rect x="50" y="80" width="280" height="160" class="hot"/>
|
||||
<text x="190" y="102" text-anchor="middle" class="step">Era 1: OCR pipeline</text>
|
||||
<text x="190" y="124" text-anchor="middle" class="small">Tesseract / TrOCR detect</text>
|
||||
<text x="190" y="140" text-anchor="middle" class="small">LayoutLMv3 layout</text>
|
||||
<text x="190" y="156" text-anchor="middle" class="small">table recognizer</text>
|
||||
<text x="190" y="172" text-anchor="middle" class="small">regex + domain rules</text>
|
||||
<text x="190" y="196" text-anchor="middle" class="step">pros: cheap, deterministic</text>
|
||||
<text x="190" y="216" text-anchor="middle" class="small">cons: brittle on new formats</text>
|
||||
|
||||
<rect x="340" y="80" width="280" height="160" class="cool"/>
|
||||
<text x="480" y="102" text-anchor="middle" class="step">Era 2: OCR-free specialists</text>
|
||||
<text x="480" y="124" text-anchor="middle" class="small">Donut: image -> JSON</text>
|
||||
<text x="480" y="140" text-anchor="middle" class="small">Nougat: paper -> LaTeX</text>
|
||||
<text x="480" y="156" text-anchor="middle" class="small">DocLLM: layout-aware gen</text>
|
||||
<text x="480" y="172" text-anchor="middle" class="small">swin / ViT encoder</text>
|
||||
<text x="480" y="196" text-anchor="middle" class="step">pros: single model</text>
|
||||
<text x="480" y="216" text-anchor="middle" class="small">cons: domain-specific</text>
|
||||
|
||||
<rect x="630" y="80" width="280" height="160" class="cold"/>
|
||||
<text x="770" y="102" text-anchor="middle" class="step">Era 3: VLM-native</text>
|
||||
<text x="770" y="124" text-anchor="middle" class="small">Qwen2.5-VL native res</text>
|
||||
<text x="770" y="140" text-anchor="middle" class="small">PaliGemma 2 doc-trained</text>
|
||||
<text x="770" y="156" text-anchor="middle" class="small">Claude 4.7 at 2576px</text>
|
||||
<text x="770" y="172" text-anchor="middle" class="small">frontier proprietary</text>
|
||||
<text x="770" y="196" text-anchor="middle" class="step">pros: no pipeline</text>
|
||||
<text x="770" y="216" text-anchor="middle" class="small">cons: cost + hallucination</text>
|
||||
|
||||
<rect x="30" y="280" width="900" height="230" class="box"/>
|
||||
<text x="480" y="302" text-anchor="middle" class="head">benchmarks + 2026 recipe picker</text>
|
||||
|
||||
<g transform="translate(60, 320)">
|
||||
<text x="0" y="15" class="step">benchmark</text>
|
||||
<text x="200" y="15" class="step">OCR+LLMv3</text>
|
||||
<text x="320" y="15" class="step">Nougat</text>
|
||||
<text x="420" y="15" class="step">PaliGemma 2</text>
|
||||
<text x="560" y="15" class="step">Claude 4.7</text>
|
||||
|
||||
<text x="0" y="40" class="small">DocVQA</text>
|
||||
<text x="200" y="40" class="small">83.0</text>
|
||||
<text x="320" y="40" class="small">77.3</text>
|
||||
<text x="420" y="40" class="small">88.4</text>
|
||||
<text x="560" y="40" class="small">95.1</text>
|
||||
|
||||
<text x="0" y="60" class="small">ChartQA</text>
|
||||
<text x="200" y="60" class="small">-</text>
|
||||
<text x="320" y="60" class="small">-</text>
|
||||
<text x="420" y="60" class="small">85.1</text>
|
||||
<text x="560" y="60" class="small">92.2</text>
|
||||
|
||||
<text x="0" y="80" class="small">Math LaTeX</text>
|
||||
<text x="200" y="80" class="small">-</text>
|
||||
<text x="320" y="80" class="small">90.5</text>
|
||||
<text x="420" y="80" class="small">82.0</text>
|
||||
<text x="560" y="80" class="small">94.3</text>
|
||||
|
||||
<text x="0" y="100" class="small">handwriting</text>
|
||||
<text x="200" y="100" class="small">65</text>
|
||||
<text x="320" y="100" class="small">-</text>
|
||||
<text x="420" y="100" class="small">80</text>
|
||||
<text x="560" y="100" class="small">92</text>
|
||||
</g>
|
||||
|
||||
<rect x="620" y="320" width="290" height="170" class="reg"/>
|
||||
<text x="765" y="342" text-anchor="middle" class="step">2026 picker</text>
|
||||
<text x="765" y="362" text-anchor="middle" class="small">10M invoices/day -> Era 1</text>
|
||||
<text x="765" y="378" text-anchor="middle" class="small">scientific papers -> Nougat + VLM</text>
|
||||
<text x="765" y="394" text-anchor="middle" class="small">mixed handwriting -> VLM-native</text>
|
||||
<text x="765" y="410" text-anchor="middle" class="small">regulated -> hybrid cross-check</text>
|
||||
<text x="765" y="436" text-anchor="middle" class="step">frontier gap</text>
|
||||
<text x="765" y="456" text-anchor="middle" class="small">open 7B VLM: ~88 DocVQA</text>
|
||||
<text x="765" y="472" text-anchor="middle" class="small">Claude 4.7: ~95, near-human</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.4 KiB |
@@ -0,0 +1,126 @@
|
||||
"""Document AI stack toy — LayoutLMv3-style inputs + Donut schema + token budgets.
|
||||
|
||||
Stdlib. Produces the three-stream LayoutLM input (text, bbox, patch-ids) for a
|
||||
toy page, generates a Donut-style JSON schema, and compares total input token
|
||||
counts across (OCR-pipeline, Donut, Nougat, VLM-native).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass
|
||||
|
||||
|
||||
@dataclass
|
||||
class Token:
|
||||
text: str
|
||||
bbox: tuple[int, int, int, int]
|
||||
|
||||
|
||||
def mock_page() -> list[Token]:
|
||||
"""A synthetic invoice page."""
|
||||
return [
|
||||
Token("INVOICE", (100, 50, 300, 80)),
|
||||
Token("ACME Co.", (100, 100, 250, 130)),
|
||||
Token("Item", (100, 200, 200, 230)),
|
||||
Token("Widget A", (100, 240, 250, 270)),
|
||||
Token("Price", (400, 200, 500, 230)),
|
||||
Token("$120.00", (400, 240, 500, 270)),
|
||||
Token("Total", (400, 400, 500, 430)),
|
||||
Token("$1,245.00", (400, 440, 550, 470)),
|
||||
]
|
||||
|
||||
|
||||
def layoutlm_input(tokens: list[Token], patch_grid: tuple[int, int] = (16, 16)) -> dict:
|
||||
"""Produce the three-stream input: text, bbox, patch-ids."""
|
||||
text_ids = [hash(t.text) % 10000 for t in tokens]
|
||||
bbox_stream = [t.bbox for t in tokens]
|
||||
n_patches = patch_grid[0] * patch_grid[1]
|
||||
patch_ids = list(range(n_patches))
|
||||
return {"text_ids": text_ids, "bbox_stream": bbox_stream,
|
||||
"patch_ids": patch_ids}
|
||||
|
||||
|
||||
def donut_schema(task: str = "invoice") -> dict:
|
||||
schemas = {
|
||||
"invoice": {
|
||||
"vendor": "<string>",
|
||||
"invoice_number": "<string>",
|
||||
"line_items": [
|
||||
{"description": "<string>", "quantity": "<int>", "price": "<float>"}
|
||||
],
|
||||
"total": "<float>",
|
||||
"currency": "<string>",
|
||||
},
|
||||
"form": {
|
||||
"form_id": "<string>",
|
||||
"fields": [
|
||||
{"name": "<string>", "value": "<string>", "confidence": "<float>"}
|
||||
],
|
||||
},
|
||||
}
|
||||
return schemas.get(task, {})
|
||||
|
||||
|
||||
def token_budget() -> None:
|
||||
print("\nINPUT TOKEN BUDGET PER PAGE (A4 at 300 DPI, ~2500x3500 px)")
|
||||
print("-" * 60)
|
||||
rows = [
|
||||
("OCR pipeline + LayoutLMv3", 512, "text + bbox + small image"),
|
||||
("Donut (OCR-free)", 4096, "swin encoder, ~4k patches"),
|
||||
("Nougat (paper pages)", 4096, "896x896, 4-tile AnyRes"),
|
||||
("VLM AnyRes 4-tile (LLaVA)", 2916, "336 tiles + thumbnail"),
|
||||
("VLM native 2048 (Qwen2.5-VL)", 8192, "native resolution"),
|
||||
("VLM native 2576 (Claude 4.7)", 12000, "frontier, best accuracy"),
|
||||
]
|
||||
print(f" {'stack':<28}{'tokens':<10} note")
|
||||
for name, toks, note in rows:
|
||||
print(f" {name:<28}{toks:<10} {note}")
|
||||
|
||||
|
||||
def demo_pipeline_output() -> None:
|
||||
print("\nLAYOUTLMv3-STYLE INPUT (invoice page)")
|
||||
print("-" * 60)
|
||||
tokens = mock_page()
|
||||
data = layoutlm_input(tokens)
|
||||
print(f" text_ids[0:4] : {data['text_ids'][:4]}...")
|
||||
print(f" bbox_stream[0:2] : {data['bbox_stream'][:2]}")
|
||||
print(f" patch_ids count : {len(data['patch_ids'])}")
|
||||
|
||||
print("\nDONUT SCHEMA (invoice)")
|
||||
print("-" * 60)
|
||||
schema = donut_schema("invoice")
|
||||
print(json.dumps(schema, indent=2))
|
||||
|
||||
|
||||
def eras_table() -> None:
|
||||
print("\nTHREE ERAS OF DOCUMENT AI")
|
||||
print("-" * 60)
|
||||
rows = [
|
||||
("Era 1 OCR pipeline", "Tesseract, TrOCR, LayoutLMv3", "deterministic"),
|
||||
("Era 2 OCR-free", "Donut, Nougat, DocLLM", "generalist less"),
|
||||
("Era 3 VLM-native", "Qwen2.5-VL, PaliGemma 2, Claude 4.7", "frontier 2026"),
|
||||
]
|
||||
for era, examples, trait in rows:
|
||||
print(f" {era:<20}{examples:<36}{trait}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 60)
|
||||
print("DOCUMENT AND DIAGRAM UNDERSTANDING (Phase 12, Lesson 22)")
|
||||
print("=" * 60)
|
||||
|
||||
demo_pipeline_output()
|
||||
token_budget()
|
||||
eras_table()
|
||||
|
||||
print("\nRECIPE PICKER")
|
||||
print("-" * 60)
|
||||
print(" 10M invoices/day : OCR pipeline + LayoutLMv3, cheap")
|
||||
print(" scientific papers : Nougat for math, VLM for figures")
|
||||
print(" mixed + handwriting : VLM-native (PaliGemma 2 or Qwen2.5-VL)")
|
||||
print(" regulated : OCR + VLM cross-check, auditable")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,171 @@
|
||||
# Document and Diagram Understanding
|
||||
|
||||
> Documents are not photos. A PDF, scientific paper, invoice, or handwritten form has layout, tables, diagrams, footnotes, headers, and semantic structure that plain image understanding cannot capture. The pre-VLM stack was a pipeline: Tesseract OCR + LayoutLMv3 + table-extraction heuristics. The VLM wave replaced that with OCR-free models — Donut (2022), Nougat (2023), DocLLM (2023) — that emit structured markup directly. By 2026 the frontier is just "feed the page image to Claude Opus 4.7 at 2576px native," and the structured-markup output comes for free. This lesson reads the three-era arc of document AI.
|
||||
|
||||
**Type:** Build
|
||||
**Languages:** Python (stdlib, layout-aware document parser skeleton)
|
||||
**Prerequisites:** Phase 12 · 05 (LLaVA), Phase 5 (NLP)
|
||||
**Time:** ~180 minutes
|
||||
|
||||
## Learning Objectives
|
||||
|
||||
- Explain the three eras of document AI: OCR pipeline, OCR-free, VLM-native.
|
||||
- Describe LayoutLMv3's three input streams: text, layout (bbox), image patches, with unified masking.
|
||||
- Compare Donut (OCR-free, image → markup), Nougat (scientific paper → LaTeX), DocLLM (layout-aware generative), PaliGemma 2 (VLM-native).
|
||||
- Pick a document model for a new task (invoices, scientific papers, handwritten forms, Chinese receipts).
|
||||
|
||||
## The Problem
|
||||
|
||||
"Understand this PDF" is deceptively hard. The information sits in:
|
||||
|
||||
- Text content (90% of the signal).
|
||||
- Layout (headers, footnotes, sidebars, two-column format).
|
||||
- Tables (rows, columns, merged cells).
|
||||
- Figures and diagrams.
|
||||
- Handwritten annotations.
|
||||
- Fonts and typography (title vs body).
|
||||
|
||||
Raw OCR dumps the text and loses the rest. A system that cares about invoices needs to know "Total: $1,245" came from the bottom-right, not from a footnote.
|
||||
|
||||
## The Concept
|
||||
|
||||
### Era 1 — OCR pipeline (pre-2021)
|
||||
|
||||
The classic stack:
|
||||
|
||||
1. PDF → image per page.
|
||||
2. Tesseract (or commercial OCR) extracts text with per-word bounding boxes.
|
||||
3. Layout analyzer identifies blocks (header, table, paragraph).
|
||||
4. Table structure recognizer parses tables.
|
||||
5. Domain rules + regex extract fields.
|
||||
|
||||
Works for clean printed text. Breaks on handwriting, skewed scans, complex tables, non-English scripts. Every failure mode requires a custom exception path.
|
||||
|
||||
### TrOCR (2021)
|
||||
|
||||
TrOCR (Li et al., arXiv:2109.10282) replaced Tesseract's classic CNN-CTC with a transformer encoder-decoder trained on synthetic + real text images. Clean win on handwritten and multilingual text. Still a pipeline (detector then TrOCR then layout), but the OCR step improved dramatically.
|
||||
|
||||
### Era 2 — OCR-free (2022-2023)
|
||||
|
||||
The first OCR-free models said: skip detection entirely, map image pixels to structured output directly.
|
||||
|
||||
Donut (Kim et al., arXiv:2111.15664):
|
||||
- Encoder-decoder transformer, encoder is Swin-B.
|
||||
- Output is JSON for form understanding, markdown for summarization, or any task-specific schema.
|
||||
- No OCR, no layout, no detection.
|
||||
|
||||
Nougat (Blecher et al., arXiv:2308.13418):
|
||||
- Trained specifically on scientific papers.
|
||||
- Output is LaTeX / markdown.
|
||||
- Handles equations, multi-column layout, figures.
|
||||
- The model every arXiv-parser calls.
|
||||
|
||||
These are specialists, not generalists. Donut on a scientific paper fails; Nougat on an invoice fails.
|
||||
|
||||
### LayoutLMv3 (2022)
|
||||
|
||||
A different track. LayoutLMv3 (Huang et al., arXiv:2204.08387) keeps OCR but adds layout understanding:
|
||||
|
||||
- Three input streams: OCR text tokens, per-token 2D bounding boxes, image patches.
|
||||
- Masked training objective across all three modalities (masked text, masked patches, masked layout).
|
||||
- Downstream: classification, entity extraction, table QA.
|
||||
|
||||
LayoutLMv3 is the peak of OCR-based document understanding. Strong on forms and invoices. Requires OCR upstream. Best pre-VLM accuracy on standardized document benchmarks.
|
||||
|
||||
### DocLLM (2023)
|
||||
|
||||
DocLLM (Wang et al., arXiv:2401.00908) is LayoutLM's generative sibling. Generates free-form answers conditioned on layout tokens. Better for QA on documents; still depends on OCR input.
|
||||
|
||||
### Era 3 — VLM-native (2024+)
|
||||
|
||||
2024 VLMs became good enough to replace the pipeline entirely. Feed the full page image at high resolution to a VLM, ask the question, get an answer.
|
||||
|
||||
- LLaVA-NeXT 336-tile AnyRes works for small documents.
|
||||
- Qwen2.5-VL dynamic-resolution handles 2048+ pixels natively.
|
||||
- Claude Opus 4.7 supports 2576px documents.
|
||||
- PaliGemma 2 (April 2025) trains specifically for documents + handwriting.
|
||||
|
||||
The gap between VLM-native and OCR-pipeline closed rapidly. By 2026, VLM-native wins on:
|
||||
|
||||
- Scene text (hand-written + printed, mixed scripts).
|
||||
- Complex tables with merged cells.
|
||||
- Math equations embedded in text.
|
||||
- Figures with text annotations.
|
||||
|
||||
OCR pipelines still win on:
|
||||
|
||||
- Pure-scan workloads at massive scale where per-page latency matters.
|
||||
- Pipeline reliability (deterministic failures vs VLM hallucinations).
|
||||
- Regulated environments requiring auditable OCR output.
|
||||
|
||||
### The Claude 4.7 / GPT-5 frontier
|
||||
|
||||
At 2576-pixel native input, frontier VLMs do document understanding at near-human accuracy. The benchmark numbers from early 2026:
|
||||
|
||||
- DocVQA: Claude 4.7 ~95.1, PaliGemma 2 ~88.4, Nougat ~77.3, pipelined LayoutLMv3 ~83.
|
||||
- ChartQA: Claude 4.7 ~92.2, GPT-4V ~78.
|
||||
- VisualMRC: Claude 4.7 ~94.
|
||||
|
||||
The closed-model gap is mostly resolution and base-LLM scale. Open models at 7B are a few points behind but catching up.
|
||||
|
||||
### Math equations and LaTeX output
|
||||
|
||||
Scientific papers need exact LaTeX output for equations. Nougat was trained on this. VLMs trained with LaTeX targets (Qwen2.5-VL-Math, Nougat derivatives) produce usable LaTeX. Without explicit LaTeX training, VLMs produce readable but imprecise transcriptions.
|
||||
|
||||
For scientific-paper pipelines in 2026: chain Nougat on the PDF, then a VLM on tricky pages.
|
||||
|
||||
### Handwriting
|
||||
|
||||
Still the hardest sub-task. Mixed printed + handwritten (doctors' notes, filled forms) is where OCR pipelines still beat VLMs for cost. Handwritten-only VLMs are improving (Claude 4.7, PaliGemma 2).
|
||||
|
||||
### 2026 recipe
|
||||
|
||||
For a new document-AI project:
|
||||
|
||||
- Pure-printed invoices at scale: LayoutLMv3 + rules, cost-efficient.
|
||||
- Mixed documents (scientific + handwritten + forms): VLM-native (PaliGemma 2 or Qwen2.5-VL).
|
||||
- Full arXiv ingestion: Nougat for math, VLM for figures.
|
||||
- Regulatory: OCR pipeline + VLM validator for cross-check.
|
||||
|
||||
## Use It
|
||||
|
||||
`code/main.py`:
|
||||
|
||||
- A toy layout-aware tokenizer: given (text, bbox) pairs, produces the LayoutLMv3-style input.
|
||||
- A Donut-style task schema generator: JSON template for forms.
|
||||
- A comparison of token budgets per page across OCR-pipeline, Donut, Nougat, and VLM-native.
|
||||
|
||||
## Ship It
|
||||
|
||||
This lesson produces `outputs/skill-document-ai-stack-picker.md`. Given a document-AI project (domain, scale, quality, regulatory), picks between OCR pipeline, OCR-free specialist, and VLM-native.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. Your project is 10M invoices per day. Which stack minimizes cost-per-page without losing accuracy?
|
||||
|
||||
2. Why does LayoutLMv3 outperform pure-CLIP-VLMs on form QA but underperform at scene-text? What does the bbox stream give up?
|
||||
|
||||
3. Nougat generates LaTeX. Propose a test case where VLM-native output beats Nougat on LaTeX fidelity, and a case where Nougat wins.
|
||||
|
||||
4. Read PaliGemma 2 paper (Google, 2024). What was the key training-data addition that lifted document accuracy vs PaliGemma 1?
|
||||
|
||||
5. Design a regulatory-safe hybrid: OCR pipeline as primary, VLM as secondary cross-check. How do you resolve disagreement?
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|-----------------|------------------------|
|
||||
| OCR pipeline | "Tesseract-style" | Stage-wise stack: detect -> OCR -> layout -> rules; deterministic, fragile |
|
||||
| OCR-free | "Donut-style" | Image-to-output transformer that skips explicit OCR; single model |
|
||||
| Layout-aware | "LayoutLM" | Input includes per-token bbox coordinates; unified masking across modalities |
|
||||
| VLM-native | "Frontier VLM" | Feed page image directly to Claude/GPT/Qwen VLM at high resolution; no pipeline |
|
||||
| DocVQA | "Doc benchmark" | Document VQA standard; most-cited score |
|
||||
| Markup output | "LaTeX / MD" | Structured output format instead of free-form text; enables downstream automation |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Li et al. — TrOCR (arXiv:2109.10282)](https://arxiv.org/abs/2109.10282)
|
||||
- [Blecher et al. — Nougat (arXiv:2308.13418)](https://arxiv.org/abs/2308.13418)
|
||||
- [Huang et al. — LayoutLMv3 (arXiv:2204.08387)](https://arxiv.org/abs/2204.08387)
|
||||
- [Kim et al. — Donut (arXiv:2111.15664)](https://arxiv.org/abs/2111.15664)
|
||||
- [Wang et al. — DocLLM (arXiv:2401.00908)](https://arxiv.org/abs/2401.00908)
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
---
|
||||
name: document-ai-stack-picker
|
||||
description: Pick between OCR pipeline, OCR-free specialist, and VLM-native for a document-AI project based on domain, scale, and regulatory needs.
|
||||
version: 1.0.0
|
||||
phase: 12
|
||||
lesson: 22
|
||||
tags: [document-ai, ocr, donut, nougat, paligemma, vlm-native]
|
||||
---
|
||||
|
||||
Given a document-AI project (domain: invoices / scientific papers / forms / mixed; scale: pages per day; quality bar; regulatory needs), pick a stack and produce a reference config.
|
||||
|
||||
Produce:
|
||||
|
||||
1. Stack pick. Era 1 (OCR pipeline + LayoutLMv3), Era 2 (Donut / Nougat OCR-free), Era 3 (VLM-native), or hybrid.
|
||||
2. Per-page cost estimate. Token count and latency at the chosen stack.
|
||||
3. Accuracy expectation. DocVQA + ChartQA + domain-specific benchmarks.
|
||||
4. Handwriting strategy. VLM-native for cost-insensitive; dedicated TrOCR + routing for scale.
|
||||
5. Math / LaTeX output. Nougat for scientific papers; VLM for other.
|
||||
6. Regulatory fallback. Hybrid with cross-check audit log.
|
||||
|
||||
Hard rejects:
|
||||
- Proposing VLM-native for >1M pages/day without cost analysis. Token cost at 2576px per page is significant.
|
||||
- Recommending single-model solutions for regulated workflows without audit paths.
|
||||
- Claiming Nougat handles scanned invoices. It does not — it is scientific-paper specialist.
|
||||
|
||||
Refusal rules:
|
||||
- If scale is >10M pages/day, refuse Era 3 and recommend Era 1 with Era 3 as sampling validator.
|
||||
- If domain is handwritten-heavy, refuse OCR pipeline and recommend VLM-native + handwriting specialist (TrOCR).
|
||||
- If LaTeX fidelity is required for equations, require Nougat in the loop.
|
||||
|
||||
Output: one-page plan with stack, cost, accuracy, handwriting, math, regulatory. End with arXiv 2308.13418 (Nougat), 2204.08387 (LayoutLMv3), 2111.15664 (Donut).
|
||||
Reference in New Issue
Block a user