Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

79 lines
3.5 KiB
JSON

{
"lesson": "65-hybrid-retrieval-bm25-dense",
"title": "Hybrid Retrieval with BM25 and Dense Embeddings",
"questions": [
{
"stage": "pre",
"question": "Why does a production RAG endpoint need both lexical and semantic retrieval at once?",
"options": [
"Lexical search is required by law for compliance",
"Semantic retrieval is always more accurate",
"It halves the index size",
"The query distribution mixes literal-identifier and paraphrased queries; one modality fails on each class"
],
"correct": 3,
"explanation": "BM25 dominates on literal symbols; dense dominates on paraphrases. A real endpoint receives both classes from the same users."
},
{
"stage": "pre",
"question": "What is the BM25 parameter b at the default 0.75?",
"options": [
"The number of fields in the document",
"A query-time exponent on cosine similarity",
"A multiplier on the IDF term",
"A length-normalization knob: longer documents are penalized but not linearly"
],
"correct": 3,
"explanation": "b interpolates between no length normalization (0) and full normalization (1); 0.75 is the published recommendation."
},
{
"stage": "check",
"question": "What does Reciprocal Rank Fusion compute for each candidate?",
"options": [
"Average of the BM25 and dense scores after min-max normalization",
"Sum of 1 / (k + rank) across each ranked list that contains the candidate",
"Product of cosine similarities across modalities",
"The median rank across modalities"
],
"correct": 1,
"explanation": "RRF sums 1 / (k + rank) contributions per ranked list; k = 60 is the published default."
},
{
"stage": "check",
"question": "Why is rank-based fusion preferred over a linear interpolation of BM25 and cosine scores?",
"options": [
"Cosine scores are unbounded",
"BM25 scores are unbounded and corpus-dependent; cosine is bounded in -1 to 1; the score scales are not comparable across corpora",
"BM25 scores cannot be sorted",
"Linear interpolation requires GPU support"
],
"correct": 1,
"explanation": "Score interpolation breaks every reindex because the score distributions shift; ranks are comparable across modalities by construction."
},
{
"stage": "check",
"question": "What happens to RRF when you make the constant k very small?",
"options": [
"The fusion becomes equivalent to a linear score sum",
"The top-1 contribution dominates the fusion; deep ranks contribute almost nothing",
"Out-of-corpus terms get a positive score",
"The fusion ignores the second modality"
],
"correct": 1,
"explanation": "Small k concentrates the rank-1 contribution; large k flattens the contribution curve and lets deep candidates vote."
},
{
"stage": "post",
"question": "Which failure mode does field weighting through token repetition specifically avoid?",
"options": [
"Adding a second score-aggregation step at query time; the math stays identical and per-field-score combination logic is not needed",
"Out-of-vocabulary embeddings",
"BM25 IDF collapsing to zero",
"Cosine similarity going negative"
],
"correct": 0,
"explanation": "Repeating tokens at index time multiplies term frequency, which lets the BM25 scoring math stay one summation rather than a separate field-score combinator."
}
]
}