mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-12/24): multimodal RAG and cross-modal retrieval
This commit is contained in:
@@ -0,0 +1,93 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 520" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.reg { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
|
||||
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="480" y="24" text-anchor="middle" class="title">Multimodal RAG — cross-modal retrieve, fuse, ground, generate</text>
|
||||
|
||||
<rect x="30" y="50" width="900" height="220" class="box"/>
|
||||
<text x="480" y="72" text-anchor="middle" class="head">query -> decompose -> 3 retrievers -> fuse -> VLM generator</text>
|
||||
|
||||
<rect x="60" y="90" width="180" height="60" class="reg"/>
|
||||
<text x="150" y="112" text-anchor="middle" class="step">query</text>
|
||||
<text x="150" y="130" text-anchor="middle" class="small">"quiet vegan brunch</text>
|
||||
<text x="150" y="145" text-anchor="middle" class="small">with natural light"</text>
|
||||
|
||||
<path d="M 245 120 L 285 100" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
<path d="M 245 120 L 285 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
<path d="M 245 140 L 285 240" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="290" y="80" width="170" height="50" class="hot"/>
|
||||
<text x="375" y="100" text-anchor="middle" class="step">text retriever</text>
|
||||
<text x="375" y="120" text-anchor="middle" class="small">reviews / menus</text>
|
||||
|
||||
<rect x="290" y="150" width="170" height="50" class="cool"/>
|
||||
<text x="375" y="170" text-anchor="middle" class="step">image retriever</text>
|
||||
<text x="375" y="190" text-anchor="middle" class="small">CLIP / SigLIP photos</text>
|
||||
|
||||
<rect x="290" y="220" width="170" height="50" class="cold"/>
|
||||
<text x="375" y="240" text-anchor="middle" class="step">audio retriever</text>
|
||||
<text x="375" y="260" text-anchor="middle" class="small">CLAP ambient clips</text>
|
||||
|
||||
<path d="M 465 105 L 505 150" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
<path d="M 465 175 L 505 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
<path d="M 465 245 L 505 200" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="510" y="130" width="170" height="90" class="reg"/>
|
||||
<text x="595" y="152" text-anchor="middle" class="step">score fusion</text>
|
||||
<text x="595" y="172" text-anchor="middle" class="small">weighted sum</text>
|
||||
<text x="595" y="188" text-anchor="middle" class="small">or MoE gate</text>
|
||||
<text x="595" y="206" text-anchor="middle" class="small">top-k candidates</text>
|
||||
|
||||
<path d="M 685 175 L 725 175" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="730" y="130" width="180" height="90" class="cool"/>
|
||||
<text x="820" y="152" text-anchor="middle" class="step">VLM generator</text>
|
||||
<text x="820" y="172" text-anchor="middle" class="small">Qwen2.5-VL / Claude</text>
|
||||
<text x="820" y="188" text-anchor="middle" class="small">grounded citations</text>
|
||||
<text x="820" y="206" text-anchor="middle" class="small">per source</text>
|
||||
|
||||
<rect x="30" y="290" width="900" height="220" class="box"/>
|
||||
<text x="480" y="312" text-anchor="middle" class="head">the three surveys of 2025</text>
|
||||
|
||||
<rect x="60" y="330" width="260" height="170" class="hot"/>
|
||||
<text x="190" y="352" text-anchor="middle" class="step">Abootorabi et al.</text>
|
||||
<text x="190" y="368" text-anchor="middle" class="small">arXiv:2502.08826</text>
|
||||
<text x="190" y="388" text-anchor="middle" class="small">comprehensive taxonomy</text>
|
||||
<text x="190" y="404" text-anchor="middle" class="small">retrieval / fusion / generation</text>
|
||||
<text x="190" y="420" text-anchor="middle" class="small">broadest coverage</text>
|
||||
<text x="190" y="446" text-anchor="middle" class="step">start here if new</text>
|
||||
<text x="190" y="466" text-anchor="middle" class="caption">names all subproblems</text>
|
||||
|
||||
<rect x="340" y="330" width="260" height="170" class="cool"/>
|
||||
<text x="470" y="352" text-anchor="middle" class="step">Mei et al.</text>
|
||||
<text x="470" y="368" text-anchor="middle" class="small">arXiv:2504.08748</text>
|
||||
<text x="470" y="388" text-anchor="middle" class="small">sub-task benchmarks</text>
|
||||
<text x="470" y="404" text-anchor="middle" class="small">failure modes cataloged</text>
|
||||
<text x="470" y="420" text-anchor="middle" class="small">useful for eval design</text>
|
||||
<text x="470" y="446" text-anchor="middle" class="step">read for evals</text>
|
||||
<text x="470" y="466" text-anchor="middle" class="caption">per-metric decomposition</text>
|
||||
|
||||
<rect x="620" y="330" width="290" height="170" class="cold"/>
|
||||
<text x="765" y="352" text-anchor="middle" class="step">Zhao et al.</text>
|
||||
<text x="765" y="368" text-anchor="middle" class="small">arXiv:2503.18016</text>
|
||||
<text x="765" y="388" text-anchor="middle" class="small">vision-focused RAG</text>
|
||||
<text x="765" y="404" text-anchor="middle" class="small">strong on ColPali-family</text>
|
||||
<text x="765" y="420" text-anchor="middle" class="small">visual-only emphasis</text>
|
||||
<text x="765" y="446" text-anchor="middle" class="step">read for vision RAG</text>
|
||||
<text x="765" y="466" text-anchor="middle" class="caption">complements lesson 23</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.8 KiB |
@@ -0,0 +1,153 @@
|
||||
"""Multimodal RAG toy — three retrievers + score fusion + grounded generator.
|
||||
|
||||
Stdlib. A synthetic restaurant corpus with text reviews, image-feature tags,
|
||||
and audio-ambiance scores. Runs three retrievers, fuses scores, emits a stub
|
||||
answer with citations. Demonstrates agentic reformulation on low-confidence.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
|
||||
|
||||
@dataclass
|
||||
class Restaurant:
|
||||
id: str
|
||||
name: str
|
||||
review_text: str
|
||||
image_tags: list[str]
|
||||
ambient_db: float
|
||||
|
||||
|
||||
CORPUS = [
|
||||
Restaurant("r1", "Sunday Plant Bistro",
|
||||
"best vegan brunch, quiet mornings, lots of windows", ["natural_light", "minimal"], 38),
|
||||
Restaurant("r2", "Orange Grove Cafe",
|
||||
"all-day vegan brunch, noisy music, industrial style", ["industrial"], 68),
|
||||
Restaurant("r3", "Vine & Leaf",
|
||||
"vegan lunch, dim lighting", ["warm_lighting"], 55),
|
||||
Restaurant("r4", "Morning Glow",
|
||||
"vegan brunch, airy space, lots of sun", ["natural_light", "airy"], 42),
|
||||
Restaurant("r5", "Steak Central",
|
||||
"steakhouse, loud atmosphere", ["dark"], 72),
|
||||
]
|
||||
|
||||
|
||||
def text_retrieve(query: str) -> dict[str, float]:
|
||||
"""Crude keyword matching for the query against review text."""
|
||||
keywords = [w.lower() for w in query.split() if len(w) > 2]
|
||||
scores = {}
|
||||
for r in CORPUS:
|
||||
text = r.review_text.lower()
|
||||
s = sum(text.count(k) for k in keywords)
|
||||
scores[r.id] = s / len(keywords) if keywords else 0
|
||||
return scores
|
||||
|
||||
|
||||
def image_retrieve(query: str) -> dict[str, float]:
|
||||
q = query.lower()
|
||||
tag_hints = []
|
||||
if "light" in q or "sun" in q:
|
||||
tag_hints.append("natural_light")
|
||||
if "airy" in q or "spacious" in q:
|
||||
tag_hints.append("airy")
|
||||
if "minimal" in q:
|
||||
tag_hints.append("minimal")
|
||||
scores = {}
|
||||
for r in CORPUS:
|
||||
s = sum(1.0 for t in tag_hints if t in r.image_tags)
|
||||
scores[r.id] = s / max(1, len(tag_hints))
|
||||
return scores
|
||||
|
||||
|
||||
def audio_retrieve(query: str) -> dict[str, float]:
|
||||
q = query.lower()
|
||||
scores = {}
|
||||
if "quiet" in q or "calm" in q:
|
||||
for r in CORPUS:
|
||||
scores[r.id] = max(0.0, 1.0 - r.ambient_db / 80.0)
|
||||
else:
|
||||
for r in CORPUS:
|
||||
scores[r.id] = 0.5
|
||||
return scores
|
||||
|
||||
|
||||
def fuse(scores_list: list[dict[str, float]], weights: list[float]) -> dict[str, float]:
|
||||
fused = {}
|
||||
for r in CORPUS:
|
||||
s = 0.0
|
||||
for w, scores in zip(weights, scores_list):
|
||||
s += w * scores.get(r.id, 0)
|
||||
fused[r.id] = s
|
||||
return fused
|
||||
|
||||
|
||||
def top_k(scored: dict[str, float], k: int = 3) -> list[tuple[str, float]]:
|
||||
return sorted(scored.items(), key=lambda x: -x[1])[:k]
|
||||
|
||||
|
||||
def grounded_generate(query: str, ranked: list[tuple[str, float]]) -> str:
|
||||
lines = [f"Answer for: '{query}'"]
|
||||
for i, (rid, score) in enumerate(ranked, 1):
|
||||
r = next(x for x in CORPUS if x.id == rid)
|
||||
lines.append(
|
||||
f" {i}. {r.name} (score {score:.2f})"
|
||||
f" [review {rid}] [img tags {r.image_tags}] [ambient {r.ambient_db}dB]")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def agentic_loop(query: str, confidence_floor: float = 0.8) -> str:
|
||||
t = text_retrieve(query)
|
||||
i = image_retrieve(query)
|
||||
a = audio_retrieve(query)
|
||||
fused = fuse([t, i, a], [0.3, 0.4, 0.3])
|
||||
top = top_k(fused, k=3)
|
||||
confidence = top[0][1] if top else 0
|
||||
|
||||
trace = [f"round 1: top={top[0]} confidence={confidence:.2f}"]
|
||||
if confidence < confidence_floor:
|
||||
trace.append(" confidence low; reformulating query")
|
||||
query2 = query + " bright windows low noise"
|
||||
i2 = image_retrieve(query2)
|
||||
a2 = audio_retrieve(query2)
|
||||
fused = fuse([t, i2, a2], [0.3, 0.5, 0.2])
|
||||
top = top_k(fused, k=3)
|
||||
trace.append(f"round 2: top={top[0]} confidence={top[0][1]:.2f}")
|
||||
return "\n".join(trace) + "\n\n" + grounded_generate(query, top)
|
||||
|
||||
|
||||
def surveys_table() -> None:
|
||||
print("\n2025 MULTIMODAL RAG SURVEYS")
|
||||
print("-" * 60)
|
||||
rows = [
|
||||
("Abootorabi et al.", "Feb 2025", "comprehensive taxonomy"),
|
||||
("Mei et al.", "Apr 2025", "sub-task benchmarks + failure modes"),
|
||||
("Zhao et al.", "Mar 2025", "vision-focused, strong on ColPali"),
|
||||
]
|
||||
for name, date, note in rows:
|
||||
print(f" {name:<22}{date:<10}{note}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 60)
|
||||
print("MULTIMODAL RAG (Phase 12, Lesson 24)")
|
||||
print("=" * 60)
|
||||
|
||||
query = "find me a quiet vegan brunch with natural light"
|
||||
print(f"\nQUERY: {query}")
|
||||
print("-" * 60)
|
||||
result = agentic_loop(query, confidence_floor=0.7)
|
||||
print(result)
|
||||
|
||||
surveys_table()
|
||||
|
||||
print("\nFUSION STRATEGIES")
|
||||
print("-" * 60)
|
||||
print(" score fusion : weighted sum, simple, fast")
|
||||
print(" MoE fusion : gating routes to experts, learnable, trains")
|
||||
print(" attention : small network weights retrieved items")
|
||||
print(" default: score fusion + slight bias toward dominant modality")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,156 @@
|
||||
# Multimodal RAG and Cross-Modal Retrieval
|
||||
|
||||
> Vision-native document RAG is one slice. Production multimodal RAG goes wider — retrieving across text, images, audio, and video for workflows like trip planning ("find me a quiet vegan brunch with natural light"), medical triage ("what injury matches this photo + these notes"), e-commerce ("outfits similar to this selfie, in my size"), and field service ("diagnose this engine sound plus photo of the part"). Three 2025 surveys — Abootorabi et al., Mei et al., Zhao et al. — codified the sub-problems: cross-modal retrieval, retrieval fusion, generation grounding, multimodal evaluation. This lesson reads the surveys and designs a production pipeline.
|
||||
|
||||
**Type:** Build
|
||||
**Languages:** Python (stdlib, cross-modal retriever with fusion + grounded generator)
|
||||
**Prerequisites:** Phase 12 · 23 (ColPali), Phase 11 (RAG basics)
|
||||
**Time:** ~180 minutes
|
||||
|
||||
## Learning Objectives
|
||||
|
||||
- Design cross-modal retrieval: text → image, image → text, audio → video, etc.
|
||||
- Compare three fusion strategies: score fusion, attention-based fusion, MoE fusion.
|
||||
- Explain generation grounding: what "cite your sources" looks like when sources are a mix of modalities.
|
||||
- Name the three canonical multimodal RAG surveys of 2025 and their sub-problem taxonomy.
|
||||
|
||||
## The Problem
|
||||
|
||||
Single-modality RAG is a solved pattern: embed query, embed chunks, retrieve, stuff into LLM. Multimodal RAG requires:
|
||||
|
||||
1. Multiple retrieval heads (each modality needs embeddings in a compatible space).
|
||||
2. Fusion of retrieval results across modalities.
|
||||
3. Generation grounding that cites sources across modalities.
|
||||
4. Evaluation metrics that cover cross-modal signal.
|
||||
|
||||
The 2025 surveys all arrive at the same taxonomy.
|
||||
|
||||
## The Concept
|
||||
|
||||
### Cross-modal retrieval
|
||||
|
||||
Retrieve documents of modality B given a query of modality A. Three patterns:
|
||||
|
||||
1. Shared embedding space. CLIP and CLAP produce text + image / text + audio embeddings in a shared space. Cosine similarity across modalities works directly. Limited to CLIP-trained pairs.
|
||||
|
||||
2. Per-modality encoder + translation. Text encoder + image encoder + a small translator module mapping between spaces. Sen2Sen by Gupta et al. and other 2024 designs. Flexible but adds complexity.
|
||||
|
||||
3. VLM as encoder. Use a VLM's hidden states as the retrieval representation. Any modality the VLM supports works. Higher quality, more expensive.
|
||||
|
||||
Choice: CLIP / SigLIP 2 for text+image; CLAP for text+audio; VLM-hidden-states for cross-modal at frontier quality.
|
||||
|
||||
### Fusion strategies
|
||||
|
||||
You retrieved 10 results: 5 images, 3 text passages, 2 audio clips. How do you merge?
|
||||
|
||||
Score fusion (cheapest). Each modality has its own retriever, each returns scores. Normalize scores within-modality then sum. Simple, often works.
|
||||
|
||||
Attention-based fusion. Concatenate all retrieved items, let a small attention network weight them. Needs training.
|
||||
|
||||
MoE fusion. Gating network routes to modality-specific experts. Different query types route differently — a visual question weights images higher.
|
||||
|
||||
Production default: score fusion with a slight bias toward the query's dominant modality. Upgrade to MoE if A/B shows clear wins on your domain.
|
||||
|
||||
### Generation grounding
|
||||
|
||||
The LLM should cite which retrieved item drove each claim. For multi-modal:
|
||||
|
||||
- Text source: standard citation `[1]`.
|
||||
- Image source: `[img 3]` with a short caption.
|
||||
- Audio: `[audio 2 at 0:34]`.
|
||||
|
||||
Train the generator with grounding-aware data: each claim in the training target is tagged with the source index. At inference, the model naturally emits citations.
|
||||
|
||||
### The 2025 surveys
|
||||
|
||||
Abootorabi et al. (arXiv:2502.08826, "Ask in Any Modality"): taxonomy for multimodal RAG. Covers retrieval, fusion, generation. Broadest coverage.
|
||||
|
||||
Mei et al. (arXiv:2504.08748, "A Survey of Multimodal RAG"): focuses on sub-task benchmarks and failure modes. Useful for evaluation design.
|
||||
|
||||
Zhao et al. (arXiv:2503.18016): vision-focused survey. Strong on ColPali-family work.
|
||||
|
||||
Reading all three gives you the state of the art as of spring 2025. Most of the sub-problems are still open.
|
||||
|
||||
### MuRAG — the foundational paper
|
||||
|
||||
MuRAG (Chen et al., 2022) was the first multimodal RAG. Retrieved image + text from a multimodal KB, generated answers. Showed feasibility before the VLM wave. Modern systems (REACT, VisRAG, M3DocRAG) build on it.
|
||||
|
||||
### A production trip-planner example
|
||||
|
||||
Query: "find me a quiet vegan brunch with natural light."
|
||||
|
||||
Pipeline:
|
||||
|
||||
1. Decompose query. "quiet" → audio/review keyword; "vegan brunch" → menu item; "natural light" → image feature.
|
||||
2. Retrieve per modality:
|
||||
- Text retrieval on reviews: "vegan brunch, quiet ambiance."
|
||||
- Image retrieval on restaurant photos: "natural light, airy."
|
||||
- Audio retrieval on ambient-sound clips: "low decibel, no music."
|
||||
3. Fuse scores. Each restaurant has a composite score.
|
||||
4. Top-k restaurants → VLM generator with all evidence → answer with citations.
|
||||
|
||||
This is well beyond text-RAG. Each modality adds signal that text alone misses.
|
||||
|
||||
### Agentic multimodal RAG
|
||||
|
||||
Multi-hop: if the first retrieval does not return high-confidence answers, the LLM reformulates and retrieves again. Agentic RAG patterns from Phase 14 apply here. Examples:
|
||||
|
||||
- Retrieve initial top-10 → LLM asks "too noisy, filter for <40 dB" → re-retrieve.
|
||||
- Retrieve images → LLM sees one has a menu → retrieve the menu text → answer.
|
||||
|
||||
Adds complexity but handles queries that single-shot retrieval cannot.
|
||||
|
||||
### Evaluation
|
||||
|
||||
Cross-modal evaluation is still immature. Common proxies:
|
||||
|
||||
- Recall@k per modality.
|
||||
- Fused top-k accuracy.
|
||||
- Human-judged end-to-end satisfaction.
|
||||
- Task-specific (bookings completed, purchases made).
|
||||
|
||||
No standard benchmark spans all modalities. Most papers evaluate on domain-specific tasks.
|
||||
|
||||
## Use It
|
||||
|
||||
`code/main.py`:
|
||||
|
||||
- Three mock retrievers (text, image, audio) operating on a shared corpus of restaurants.
|
||||
- Score fusion that combines modality scores with configurable weights.
|
||||
- A generator stub that emits a final answer with citations.
|
||||
- A simple agentic loop that reformulates the query if confidence is low.
|
||||
|
||||
## Ship It
|
||||
|
||||
This lesson produces `outputs/skill-multimodal-rag-designer.md`. Given a product spec with a multimodal query flow, designs retrievers, fusion, generator, and evaluation.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. Propose a medical-triage multimodal RAG: query = photo of injury + text symptoms. What modalities retrieve from what KB?
|
||||
|
||||
2. Score fusion is a simple weighted sum. What failure mode does it have that MoE fusion avoids?
|
||||
|
||||
3. Read Abootorabi et al.'s taxonomy (Section 3). What are the three canonical sub-problems and how do they map to your chosen product?
|
||||
|
||||
4. Design an eval spec for a trip-planner multimodal RAG. What metrics cover image recall, audio recall, and composite correctness?
|
||||
|
||||
5. Agentic multi-hop RAG has a latency tax per round-trip. At what query difficulty does the accuracy gain justify the latency?
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|-----------------|------------------------|
|
||||
| Cross-modal retrieval | "Query one modality, retrieve another" | Text query retrieves images; image query retrieves text; requires a shared space or translator |
|
||||
| Score fusion | "Combine scores" | Weighted sum of per-modality retrieval scores; simplest fusion |
|
||||
| MoE fusion | "Modality-routed experts" | Gating network picks which modality's scores to trust per query |
|
||||
| Grounded generation | "Cite your sources" | Each claim in the answer tagged with the source index |
|
||||
| MuRAG | "First multimodal RAG" | 2022 paper that established the multimodal RAG pattern |
|
||||
| Agentic multi-hop | "Reformulate and retry" | LLM re-queries retrievers when first-pass confidence is low |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Abootorabi et al. — Ask in Any Modality (arXiv:2502.08826)](https://arxiv.org/abs/2502.08826)
|
||||
- [Mei et al. — A Survey of Multimodal RAG (arXiv:2504.08748)](https://arxiv.org/abs/2504.08748)
|
||||
- [Zhao et al. — Vision RAG Survey (arXiv:2503.18016)](https://arxiv.org/abs/2503.18016)
|
||||
- [Chen et al. — MuRAG (arXiv:2210.02928)](https://arxiv.org/abs/2210.02928)
|
||||
- [Liu et al. — REACT (arXiv:2301.10382)](https://arxiv.org/abs/2301.10382)
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
---
|
||||
name: multimodal-rag-designer
|
||||
description: Design a production multimodal RAG across text, images, audio, video with retrievers, fusion strategy, and grounded generator.
|
||||
version: 1.0.0
|
||||
phase: 12
|
||||
lesson: 24
|
||||
tags: [multimodal-rag, cross-modal-retrieval, fusion, grounded-generation]
|
||||
---
|
||||
|
||||
Given a multimodal product query flow (which modalities in the query, which in the corpus), design retrievers, fusion, and generation.
|
||||
|
||||
Produce:
|
||||
|
||||
1. Per-modality retrievers. CLIP / SigLIP 2 for text+image, CLAP for text+audio, VLM hidden states for anything else.
|
||||
2. Fusion pick. Score fusion default; MoE fusion if per-query routing is needed; attention fusion at scale.
|
||||
3. Grounded generator. Qwen2.5-VL or Claude 4.7 with training on source-tagged outputs.
|
||||
4. Evaluation. Recall@k per modality + fused top-k accuracy + human-judged end-to-end.
|
||||
5. Agentic multi-hop. When to re-query; confidence threshold to trigger.
|
||||
6. Storage estimate. Per-modality vector counts and compression.
|
||||
|
||||
Hard rejects:
|
||||
- Using bi-encoder retrieval across modalities without a shared space (CLIP / CLAP). Scores are meaningless.
|
||||
- Proposing MoE fusion without training data. MoE needs supervision to route correctly.
|
||||
- Claiming score-fusion weights transfer across domains. They do not.
|
||||
|
||||
Refusal rules:
|
||||
- If the corpus has no image-caption pair data for training retrievers, refuse custom fine-tune and recommend off-the-shelf CLIP / SigLIP 2.
|
||||
- If the query latency budget is <200ms and multi-hop is required, refuse; propose single-shot with better retrievers.
|
||||
- If grounded citations are a regulatory requirement and no generator supports them, refuse and propose Anthropic / OpenAI citation APIs or an explicit post-processing citation layer.
|
||||
|
||||
Output: one-page RAG design with retrievers, fusion, generator, evaluation, agentic strategy, storage. End with arXiv 2502.08826, 2504.08748, 2503.18016.
|
||||
Reference in New Issue
Block a user