Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

103 lines
4.0 KiB
JSON

{
"lesson": "23-chunking-strategies-rag",
"title": "Chunking Strategies for RAG",
"questions": [
{
"stage": "pre",
"question": "Why is chunking strategy as important as the embedding model in RAG?",
"options": [
"Chunking shrinks the model",
"Smaller chunks train faster",
"Chunking is required by FAISS",
"Chunk boundaries determine whether the answer is even retrievable; bad chunks defeat any embedding"
],
"correct": 3,
"explanation": "Vectara's 2025 study showed chunking quality matches or exceeds embedding-model impact on retrieval quality."
},
{
"stage": "pre",
"question": "What does LangChain's RecursiveCharacterTextSplitter try in order?",
"options": [
"Try splitting on paragraph breaks, then newlines, then sentence boundaries, then spaces",
"Split on token IDs",
"Split on whitespace only",
"Always split on character N"
],
"correct": 0,
"explanation": "Recursive splitting falls back through paragraph -> newline -> sentence -> space to preserve structure."
},
{
"stage": "check",
"question": "Why does the parent-document pattern improve answer quality?",
"options": [
"It removes embeddings",
"Children give precise retrieval; returning the larger parent block preserves the surrounding context the reader needs",
"It avoids tokenization",
"It uses fewer GPU cycles"
],
"correct": 1,
"explanation": "Retrieve by small child chunks for precision, then expand to the parent for context."
},
{
"stage": "check",
"question": "What does Anthropic's 'contextual retrieval' add to each chunk before indexing?",
"options": [
"An LLM-generated 50-100 word summary placing the chunk in the document's overall context",
"Random noise",
"A POS tag",
"A language code"
],
"correct": 0,
"explanation": "Contextual retrieval prepends an LLM-written situating summary to each chunk; ~35-50% recall gain."
},
{
"stage": "check",
"question": "Which 2026 finding contradicts the conventional wisdom about chunk overlap?",
"options": [
"Overlap should be 50%",
"Overlap is required for BM25",
"Overlap improves contextual retrieval only",
"Empirical 2026 benchmarks show overlap often provides zero measurable benefit while doubling index cost"
],
"correct": 3,
"explanation": "Newer studies (SPLADE+Mistral on NQ) show chunk overlap rarely helps and inflates index size."
},
{
"stage": "post",
"question": "Which chunk size does NVIDIA's 2026 benchmark associate with factoid queries?",
"options": [
"Roughly 256-512 tokens",
"2048-4096 tokens",
"8192 tokens",
"64 tokens"
],
"correct": 0,
"explanation": "Factoid queries benefit from smaller chunks (256-512 tokens) that concentrate the answer signal."
},
{
"stage": "post",
"question": "Why is a min-token floor important when using semantic chunking?",
"options": [
"Without a floor, semantic chunking can produce tiny 40-token fragments that hurt retrieval",
"Required by BERT",
"Floors prevent overlap",
"Floors speed up inference"
],
"correct": 0,
"explanation": "Semantic chunkers can over-segment; enforcing a min size prevents low-signal fragments."
},
{
"stage": "post",
"question": "What does 'late chunking' do differently from traditional chunking?",
"options": [
"Chunks after retrieval",
"Embeds the whole document at the token level first, then pools token embeddings into chunk vectors to preserve cross-chunk context",
"Skips chunking entirely",
"Uses BM25 instead"
],
"correct": 1,
"explanation": "Late chunking embeds first and pools second, preserving contextual interactions across chunk boundaries."
}
]
}