feat(phase-12/24): multimodal RAG and cross-modal retrieval

This commit is contained in:
Rohit Ghumare
2026-04-24 12:40:05 +01:00
parent 480c2708e2
commit cd2ebc77bd
5 changed files with 433 additions and 0 deletions
@@ -0,0 +1,93 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 520" font-family="Georgia, 'Times New Roman', serif">
<defs>
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
.reg { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="480" y="24" text-anchor="middle" class="title">Multimodal RAG — cross-modal retrieve, fuse, ground, generate</text>
<rect x="30" y="50" width="900" height="220" class="box"/>
<text x="480" y="72" text-anchor="middle" class="head">query -&gt; decompose -&gt; 3 retrievers -&gt; fuse -&gt; VLM generator</text>
<rect x="60" y="90" width="180" height="60" class="reg"/>
<text x="150" y="112" text-anchor="middle" class="step">query</text>
<text x="150" y="130" text-anchor="middle" class="small">"quiet vegan brunch</text>
<text x="150" y="145" text-anchor="middle" class="small">with natural light"</text>
<path d="M 245 120 L 285 100" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M 245 120 L 285 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M 245 140 L 285 240" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<rect x="290" y="80" width="170" height="50" class="hot"/>
<text x="375" y="100" text-anchor="middle" class="step">text retriever</text>
<text x="375" y="120" text-anchor="middle" class="small">reviews / menus</text>
<rect x="290" y="150" width="170" height="50" class="cool"/>
<text x="375" y="170" text-anchor="middle" class="step">image retriever</text>
<text x="375" y="190" text-anchor="middle" class="small">CLIP / SigLIP photos</text>
<rect x="290" y="220" width="170" height="50" class="cold"/>
<text x="375" y="240" text-anchor="middle" class="step">audio retriever</text>
<text x="375" y="260" text-anchor="middle" class="small">CLAP ambient clips</text>
<path d="M 465 105 L 505 150" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M 465 175 L 505 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<path d="M 465 245 L 505 200" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<rect x="510" y="130" width="170" height="90" class="reg"/>
<text x="595" y="152" text-anchor="middle" class="step">score fusion</text>
<text x="595" y="172" text-anchor="middle" class="small">weighted sum</text>
<text x="595" y="188" text-anchor="middle" class="small">or MoE gate</text>
<text x="595" y="206" text-anchor="middle" class="small">top-k candidates</text>
<path d="M 685 175 L 725 175" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
<rect x="730" y="130" width="180" height="90" class="cool"/>
<text x="820" y="152" text-anchor="middle" class="step">VLM generator</text>
<text x="820" y="172" text-anchor="middle" class="small">Qwen2.5-VL / Claude</text>
<text x="820" y="188" text-anchor="middle" class="small">grounded citations</text>
<text x="820" y="206" text-anchor="middle" class="small">per source</text>
<rect x="30" y="290" width="900" height="220" class="box"/>
<text x="480" y="312" text-anchor="middle" class="head">the three surveys of 2025</text>
<rect x="60" y="330" width="260" height="170" class="hot"/>
<text x="190" y="352" text-anchor="middle" class="step">Abootorabi et al.</text>
<text x="190" y="368" text-anchor="middle" class="small">arXiv:2502.08826</text>
<text x="190" y="388" text-anchor="middle" class="small">comprehensive taxonomy</text>
<text x="190" y="404" text-anchor="middle" class="small">retrieval / fusion / generation</text>
<text x="190" y="420" text-anchor="middle" class="small">broadest coverage</text>
<text x="190" y="446" text-anchor="middle" class="step">start here if new</text>
<text x="190" y="466" text-anchor="middle" class="caption">names all subproblems</text>
<rect x="340" y="330" width="260" height="170" class="cool"/>
<text x="470" y="352" text-anchor="middle" class="step">Mei et al.</text>
<text x="470" y="368" text-anchor="middle" class="small">arXiv:2504.08748</text>
<text x="470" y="388" text-anchor="middle" class="small">sub-task benchmarks</text>
<text x="470" y="404" text-anchor="middle" class="small">failure modes cataloged</text>
<text x="470" y="420" text-anchor="middle" class="small">useful for eval design</text>
<text x="470" y="446" text-anchor="middle" class="step">read for evals</text>
<text x="470" y="466" text-anchor="middle" class="caption">per-metric decomposition</text>
<rect x="620" y="330" width="290" height="170" class="cold"/>
<text x="765" y="352" text-anchor="middle" class="step">Zhao et al.</text>
<text x="765" y="368" text-anchor="middle" class="small">arXiv:2503.18016</text>
<text x="765" y="388" text-anchor="middle" class="small">vision-focused RAG</text>
<text x="765" y="404" text-anchor="middle" class="small">strong on ColPali-family</text>
<text x="765" y="420" text-anchor="middle" class="small">visual-only emphasis</text>
<text x="765" y="446" text-anchor="middle" class="step">read for vision RAG</text>
<text x="765" y="466" text-anchor="middle" class="caption">complements lesson 23</text>
</svg>

After

Width:  |  Height:  |  Size: 5.8 KiB

@@ -0,0 +1,153 @@
"""Multimodal RAG toy — three retrievers + score fusion + grounded generator.
Stdlib. A synthetic restaurant corpus with text reviews, image-feature tags,
and audio-ambiance scores. Runs three retrievers, fuses scores, emits a stub
answer with citations. Demonstrates agentic reformulation on low-confidence.
"""
from __future__ import annotations
from dataclasses import dataclass
@dataclass
class Restaurant:
id: str
name: str
review_text: str
image_tags: list[str]
ambient_db: float
CORPUS = [
Restaurant("r1", "Sunday Plant Bistro",
"best vegan brunch, quiet mornings, lots of windows", ["natural_light", "minimal"], 38),
Restaurant("r2", "Orange Grove Cafe",
"all-day vegan brunch, noisy music, industrial style", ["industrial"], 68),
Restaurant("r3", "Vine & Leaf",
"vegan lunch, dim lighting", ["warm_lighting"], 55),
Restaurant("r4", "Morning Glow",
"vegan brunch, airy space, lots of sun", ["natural_light", "airy"], 42),
Restaurant("r5", "Steak Central",
"steakhouse, loud atmosphere", ["dark"], 72),
]
def text_retrieve(query: str) -> dict[str, float]:
"""Crude keyword matching for the query against review text."""
keywords = [w.lower() for w in query.split() if len(w) > 2]
scores = {}
for r in CORPUS:
text = r.review_text.lower()
s = sum(text.count(k) for k in keywords)
scores[r.id] = s / len(keywords) if keywords else 0
return scores
def image_retrieve(query: str) -> dict[str, float]:
q = query.lower()
tag_hints = []
if "light" in q or "sun" in q:
tag_hints.append("natural_light")
if "airy" in q or "spacious" in q:
tag_hints.append("airy")
if "minimal" in q:
tag_hints.append("minimal")
scores = {}
for r in CORPUS:
s = sum(1.0 for t in tag_hints if t in r.image_tags)
scores[r.id] = s / max(1, len(tag_hints))
return scores
def audio_retrieve(query: str) -> dict[str, float]:
q = query.lower()
scores = {}
if "quiet" in q or "calm" in q:
for r in CORPUS:
scores[r.id] = max(0.0, 1.0 - r.ambient_db / 80.0)
else:
for r in CORPUS:
scores[r.id] = 0.5
return scores
def fuse(scores_list: list[dict[str, float]], weights: list[float]) -> dict[str, float]:
fused = {}
for r in CORPUS:
s = 0.0
for w, scores in zip(weights, scores_list):
s += w * scores.get(r.id, 0)
fused[r.id] = s
return fused
def top_k(scored: dict[str, float], k: int = 3) -> list[tuple[str, float]]:
return sorted(scored.items(), key=lambda x: -x[1])[:k]
def grounded_generate(query: str, ranked: list[tuple[str, float]]) -> str:
lines = [f"Answer for: '{query}'"]
for i, (rid, score) in enumerate(ranked, 1):
r = next(x for x in CORPUS if x.id == rid)
lines.append(
f" {i}. {r.name} (score {score:.2f})"
f" [review {rid}] [img tags {r.image_tags}] [ambient {r.ambient_db}dB]")
return "\n".join(lines)
def agentic_loop(query: str, confidence_floor: float = 0.8) -> str:
t = text_retrieve(query)
i = image_retrieve(query)
a = audio_retrieve(query)
fused = fuse([t, i, a], [0.3, 0.4, 0.3])
top = top_k(fused, k=3)
confidence = top[0][1] if top else 0
trace = [f"round 1: top={top[0]} confidence={confidence:.2f}"]
if confidence < confidence_floor:
trace.append(" confidence low; reformulating query")
query2 = query + " bright windows low noise"
i2 = image_retrieve(query2)
a2 = audio_retrieve(query2)
fused = fuse([t, i2, a2], [0.3, 0.5, 0.2])
top = top_k(fused, k=3)
trace.append(f"round 2: top={top[0]} confidence={top[0][1]:.2f}")
return "\n".join(trace) + "\n\n" + grounded_generate(query, top)
def surveys_table() -> None:
print("\n2025 MULTIMODAL RAG SURVEYS")
print("-" * 60)
rows = [
("Abootorabi et al.", "Feb 2025", "comprehensive taxonomy"),
("Mei et al.", "Apr 2025", "sub-task benchmarks + failure modes"),
("Zhao et al.", "Mar 2025", "vision-focused, strong on ColPali"),
]
for name, date, note in rows:
print(f" {name:<22}{date:<10}{note}")
def main() -> None:
print("=" * 60)
print("MULTIMODAL RAG (Phase 12, Lesson 24)")
print("=" * 60)
query = "find me a quiet vegan brunch with natural light"
print(f"\nQUERY: {query}")
print("-" * 60)
result = agentic_loop(query, confidence_floor=0.7)
print(result)
surveys_table()
print("\nFUSION STRATEGIES")
print("-" * 60)
print(" score fusion : weighted sum, simple, fast")
print(" MoE fusion : gating routes to experts, learnable, trains")
print(" attention : small network weights retrieved items")
print(" default: score fusion + slight bias toward dominant modality")
if __name__ == "__main__":
main()
@@ -0,0 +1,156 @@
# Multimodal RAG and Cross-Modal Retrieval
> Vision-native document RAG is one slice. Production multimodal RAG goes wider — retrieving across text, images, audio, and video for workflows like trip planning ("find me a quiet vegan brunch with natural light"), medical triage ("what injury matches this photo + these notes"), e-commerce ("outfits similar to this selfie, in my size"), and field service ("diagnose this engine sound plus photo of the part"). Three 2025 surveys — Abootorabi et al., Mei et al., Zhao et al. — codified the sub-problems: cross-modal retrieval, retrieval fusion, generation grounding, multimodal evaluation. This lesson reads the surveys and designs a production pipeline.
**Type:** Build
**Languages:** Python (stdlib, cross-modal retriever with fusion + grounded generator)
**Prerequisites:** Phase 12 · 23 (ColPali), Phase 11 (RAG basics)
**Time:** ~180 minutes
## Learning Objectives
- Design cross-modal retrieval: text → image, image → text, audio → video, etc.
- Compare three fusion strategies: score fusion, attention-based fusion, MoE fusion.
- Explain generation grounding: what "cite your sources" looks like when sources are a mix of modalities.
- Name the three canonical multimodal RAG surveys of 2025 and their sub-problem taxonomy.
## The Problem
Single-modality RAG is a solved pattern: embed query, embed chunks, retrieve, stuff into LLM. Multimodal RAG requires:
1. Multiple retrieval heads (each modality needs embeddings in a compatible space).
2. Fusion of retrieval results across modalities.
3. Generation grounding that cites sources across modalities.
4. Evaluation metrics that cover cross-modal signal.
The 2025 surveys all arrive at the same taxonomy.
## The Concept
### Cross-modal retrieval
Retrieve documents of modality B given a query of modality A. Three patterns:
1. Shared embedding space. CLIP and CLAP produce text + image / text + audio embeddings in a shared space. Cosine similarity across modalities works directly. Limited to CLIP-trained pairs.
2. Per-modality encoder + translation. Text encoder + image encoder + a small translator module mapping between spaces. Sen2Sen by Gupta et al. and other 2024 designs. Flexible but adds complexity.
3. VLM as encoder. Use a VLM's hidden states as the retrieval representation. Any modality the VLM supports works. Higher quality, more expensive.
Choice: CLIP / SigLIP 2 for text+image; CLAP for text+audio; VLM-hidden-states for cross-modal at frontier quality.
### Fusion strategies
You retrieved 10 results: 5 images, 3 text passages, 2 audio clips. How do you merge?
Score fusion (cheapest). Each modality has its own retriever, each returns scores. Normalize scores within-modality then sum. Simple, often works.
Attention-based fusion. Concatenate all retrieved items, let a small attention network weight them. Needs training.
MoE fusion. Gating network routes to modality-specific experts. Different query types route differently — a visual question weights images higher.
Production default: score fusion with a slight bias toward the query's dominant modality. Upgrade to MoE if A/B shows clear wins on your domain.
### Generation grounding
The LLM should cite which retrieved item drove each claim. For multi-modal:
- Text source: standard citation `[1]`.
- Image source: `[img 3]` with a short caption.
- Audio: `[audio 2 at 0:34]`.
Train the generator with grounding-aware data: each claim in the training target is tagged with the source index. At inference, the model naturally emits citations.
### The 2025 surveys
Abootorabi et al. (arXiv:2502.08826, "Ask in Any Modality"): taxonomy for multimodal RAG. Covers retrieval, fusion, generation. Broadest coverage.
Mei et al. (arXiv:2504.08748, "A Survey of Multimodal RAG"): focuses on sub-task benchmarks and failure modes. Useful for evaluation design.
Zhao et al. (arXiv:2503.18016): vision-focused survey. Strong on ColPali-family work.
Reading all three gives you the state of the art as of spring 2025. Most of the sub-problems are still open.
### MuRAG — the foundational paper
MuRAG (Chen et al., 2022) was the first multimodal RAG. Retrieved image + text from a multimodal KB, generated answers. Showed feasibility before the VLM wave. Modern systems (REACT, VisRAG, M3DocRAG) build on it.
### A production trip-planner example
Query: "find me a quiet vegan brunch with natural light."
Pipeline:
1. Decompose query. "quiet" → audio/review keyword; "vegan brunch" → menu item; "natural light" → image feature.
2. Retrieve per modality:
- Text retrieval on reviews: "vegan brunch, quiet ambiance."
- Image retrieval on restaurant photos: "natural light, airy."
- Audio retrieval on ambient-sound clips: "low decibel, no music."
3. Fuse scores. Each restaurant has a composite score.
4. Top-k restaurants → VLM generator with all evidence → answer with citations.
This is well beyond text-RAG. Each modality adds signal that text alone misses.
### Agentic multimodal RAG
Multi-hop: if the first retrieval does not return high-confidence answers, the LLM reformulates and retrieves again. Agentic RAG patterns from Phase 14 apply here. Examples:
- Retrieve initial top-10 → LLM asks "too noisy, filter for <40 dB" → re-retrieve.
- Retrieve images → LLM sees one has a menu → retrieve the menu text → answer.
Adds complexity but handles queries that single-shot retrieval cannot.
### Evaluation
Cross-modal evaluation is still immature. Common proxies:
- Recall@k per modality.
- Fused top-k accuracy.
- Human-judged end-to-end satisfaction.
- Task-specific (bookings completed, purchases made).
No standard benchmark spans all modalities. Most papers evaluate on domain-specific tasks.
## Use It
`code/main.py`:
- Three mock retrievers (text, image, audio) operating on a shared corpus of restaurants.
- Score fusion that combines modality scores with configurable weights.
- A generator stub that emits a final answer with citations.
- A simple agentic loop that reformulates the query if confidence is low.
## Ship It
This lesson produces `outputs/skill-multimodal-rag-designer.md`. Given a product spec with a multimodal query flow, designs retrievers, fusion, generator, and evaluation.
## Exercises
1. Propose a medical-triage multimodal RAG: query = photo of injury + text symptoms. What modalities retrieve from what KB?
2. Score fusion is a simple weighted sum. What failure mode does it have that MoE fusion avoids?
3. Read Abootorabi et al.'s taxonomy (Section 3). What are the three canonical sub-problems and how do they map to your chosen product?
4. Design an eval spec for a trip-planner multimodal RAG. What metrics cover image recall, audio recall, and composite correctness?
5. Agentic multi-hop RAG has a latency tax per round-trip. At what query difficulty does the accuracy gain justify the latency?
## Key Terms
| Term | What people say | What it actually means |
|------|-----------------|------------------------|
| Cross-modal retrieval | "Query one modality, retrieve another" | Text query retrieves images; image query retrieves text; requires a shared space or translator |
| Score fusion | "Combine scores" | Weighted sum of per-modality retrieval scores; simplest fusion |
| MoE fusion | "Modality-routed experts" | Gating network picks which modality's scores to trust per query |
| Grounded generation | "Cite your sources" | Each claim in the answer tagged with the source index |
| MuRAG | "First multimodal RAG" | 2022 paper that established the multimodal RAG pattern |
| Agentic multi-hop | "Reformulate and retry" | LLM re-queries retrievers when first-pass confidence is low |
## Further Reading
- [Abootorabi et al. — Ask in Any Modality (arXiv:2502.08826)](https://arxiv.org/abs/2502.08826)
- [Mei et al. — A Survey of Multimodal RAG (arXiv:2504.08748)](https://arxiv.org/abs/2504.08748)
- [Zhao et al. — Vision RAG Survey (arXiv:2503.18016)](https://arxiv.org/abs/2503.18016)
- [Chen et al. — MuRAG (arXiv:2210.02928)](https://arxiv.org/abs/2210.02928)
- [Liu et al. — REACT (arXiv:2301.10382)](https://arxiv.org/abs/2301.10382)
@@ -0,0 +1,31 @@
---
name: multimodal-rag-designer
description: Design a production multimodal RAG across text, images, audio, video with retrievers, fusion strategy, and grounded generator.
version: 1.0.0
phase: 12
lesson: 24
tags: [multimodal-rag, cross-modal-retrieval, fusion, grounded-generation]
---
Given a multimodal product query flow (which modalities in the query, which in the corpus), design retrievers, fusion, and generation.
Produce:
1. Per-modality retrievers. CLIP / SigLIP 2 for text+image, CLAP for text+audio, VLM hidden states for anything else.
2. Fusion pick. Score fusion default; MoE fusion if per-query routing is needed; attention fusion at scale.
3. Grounded generator. Qwen2.5-VL or Claude 4.7 with training on source-tagged outputs.
4. Evaluation. Recall@k per modality + fused top-k accuracy + human-judged end-to-end.
5. Agentic multi-hop. When to re-query; confidence threshold to trigger.
6. Storage estimate. Per-modality vector counts and compression.
Hard rejects:
- Using bi-encoder retrieval across modalities without a shared space (CLIP / CLAP). Scores are meaningless.
- Proposing MoE fusion without training data. MoE needs supervision to route correctly.
- Claiming score-fusion weights transfer across domains. They do not.
Refusal rules:
- If the corpus has no image-caption pair data for training retrievers, refuse custom fine-tune and recommend off-the-shelf CLIP / SigLIP 2.
- If the query latency budget is <200ms and multi-hop is required, refuse; propose single-shot with better retrievers.
- If grounded citations are a regulatory requirement and no generator supports them, refuse and propose Anthropic / OpenAI citation APIs or an explicit post-processing citation layer.
Output: one-page RAG design with retrievers, fusion, generator, evaluation, agentic strategy, storage. End with arXiv 2502.08826, 2504.08748, 2503.18016.