mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
103 lines
4.0 KiB
JSON
103 lines
4.0 KiB
JSON
{
|
|
"lesson": "23-chunking-strategies-rag",
|
|
"title": "Chunking Strategies for RAG",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "Why is chunking strategy as important as the embedding model in RAG?",
|
|
"options": [
|
|
"Chunking shrinks the model",
|
|
"Smaller chunks train faster",
|
|
"Chunking is required by FAISS",
|
|
"Chunk boundaries determine whether the answer is even retrievable; bad chunks defeat any embedding"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Vectara's 2025 study showed chunking quality matches or exceeds embedding-model impact on retrieval quality."
|
|
},
|
|
{
|
|
"stage": "pre",
|
|
"question": "What does LangChain's RecursiveCharacterTextSplitter try in order?",
|
|
"options": [
|
|
"Try splitting on paragraph breaks, then newlines, then sentence boundaries, then spaces",
|
|
"Split on token IDs",
|
|
"Split on whitespace only",
|
|
"Always split on character N"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Recursive splitting falls back through paragraph -> newline -> sentence -> space to preserve structure."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why does the parent-document pattern improve answer quality?",
|
|
"options": [
|
|
"It removes embeddings",
|
|
"Children give precise retrieval; returning the larger parent block preserves the surrounding context the reader needs",
|
|
"It avoids tokenization",
|
|
"It uses fewer GPU cycles"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Retrieve by small child chunks for precision, then expand to the parent for context."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What does Anthropic's 'contextual retrieval' add to each chunk before indexing?",
|
|
"options": [
|
|
"An LLM-generated 50-100 word summary placing the chunk in the document's overall context",
|
|
"Random noise",
|
|
"A POS tag",
|
|
"A language code"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Contextual retrieval prepends an LLM-written situating summary to each chunk; ~35-50% recall gain."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Which 2026 finding contradicts the conventional wisdom about chunk overlap?",
|
|
"options": [
|
|
"Overlap should be 50%",
|
|
"Overlap is required for BM25",
|
|
"Overlap improves contextual retrieval only",
|
|
"Empirical 2026 benchmarks show overlap often provides zero measurable benefit while doubling index cost"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Newer studies (SPLADE+Mistral on NQ) show chunk overlap rarely helps and inflates index size."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which chunk size does NVIDIA's 2026 benchmark associate with factoid queries?",
|
|
"options": [
|
|
"Roughly 256-512 tokens",
|
|
"2048-4096 tokens",
|
|
"8192 tokens",
|
|
"64 tokens"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Factoid queries benefit from smaller chunks (256-512 tokens) that concentrate the answer signal."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Why is a min-token floor important when using semantic chunking?",
|
|
"options": [
|
|
"Without a floor, semantic chunking can produce tiny 40-token fragments that hurt retrieval",
|
|
"Required by BERT",
|
|
"Floors prevent overlap",
|
|
"Floors speed up inference"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Semantic chunkers can over-segment; enforcing a min size prevents low-signal fragments."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "What does 'late chunking' do differently from traditional chunking?",
|
|
"options": [
|
|
"Chunks after retrieval",
|
|
"Embeds the whole document at the token level first, then pools token embeddings into chunk vectors to preserve cross-chunk context",
|
|
"Skips chunking entirely",
|
|
"Uses BM25 instead"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Late chunking embeds first and pools second, preserving contextual interactions across chunk boundaries."
|
|
}
|
|
]
|
|
}
|