mirror of
https://github.com/nashsu/llm_wiki.git
synced 2026-10-02 02:44:34 +08:00
Spec for adding image extraction from PDF/PPTX/DOCX, vision-LLM captioning, and indexing of captions through the existing RAG pipeline. Phased delivery; Phase 5 (multimodal embedding) deferred. See plans/multimodal-images.md for the full design, current-state audit, open questions, and per-phase implementation notes.