mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
103 lines
3.8 KiB
JSON
103 lines
3.8 KiB
JSON
{
|
|
"lesson": "13-question-answering",
|
|
"title": "Question Answering Systems",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "What does extractive QA predict?",
|
|
"options": [
|
|
"Start and end token indices of the answer span within a given passage",
|
|
"A retrieved passage ID",
|
|
"A generated natural-language answer",
|
|
"A confidence score only"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Extractive QA outputs the span of the passage that contains the answer."
|
|
},
|
|
{
|
|
"stage": "pre",
|
|
"question": "What two components define a basic RAG pipeline?",
|
|
"options": [
|
|
"Tokenizer and POS tagger",
|
|
"An encoder and a decoder trained jointly",
|
|
"A reranker and a translator",
|
|
"A retriever (find relevant passages) and a reader (extract or generate the answer)"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "RAG = retriever (finds relevant context) plus reader (answers from it)."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "On SQuAD, what does Exact Match (EM) measure?",
|
|
"options": [
|
|
"Token-level F1",
|
|
"Whether the prediction matches the reference exactly after normalization (lowercase, strip punctuation, remove articles)",
|
|
"Edit distance",
|
|
"Per-word overlap"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "EM is strict equality after a defined normalization step; partial matches score zero."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What does deepset/roberta-base-squad2 add over a SQuAD 1.1 model?",
|
|
"options": [
|
|
"Bigger context window",
|
|
"Multilingual support",
|
|
"Cross-lingual retrieval",
|
|
"Training on unanswerable questions so the model can predict a null answer"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "SQuAD 2.0 includes unanswerable items; models trained on it can predict 'no answer'."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Which RAGAS dimension targets hallucinations specifically?",
|
|
"options": [
|
|
"Context precision",
|
|
"Answer relevance",
|
|
"Faithfulness, measured by NLI entailment between answer claims and retrieved context",
|
|
"Context recall"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "Faithfulness checks each answer claim against retrieved context via NLI entailment."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Why should you measure retrieval recall before evaluating reader accuracy?",
|
|
"options": [
|
|
"Reader latency depends on it",
|
|
"If the correct passage is not in the top-k, the reader cannot succeed regardless of how good it is",
|
|
"Recall determines ROUGE",
|
|
"Required by transformers"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "A reader cannot answer when the right passage is missing; retrieval recall bounds reader performance."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which prompt pattern reduces hallucinations in RAG generation?",
|
|
"options": [
|
|
"Removing the question",
|
|
"Including more passages",
|
|
"Asking the model to be creative",
|
|
"Telling the model to answer only from the provided context and to reply 'I don't know' when the context is insufficient"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Grounding + explicit refusal instructions cuts hallucination rates substantially."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "When is extractive QA still preferred over generative RAG in 2026?",
|
|
"options": [
|
|
"Regulated domains (legal, medical, audit) where literal quotation from authoritative sources is required",
|
|
"Open-domain trivia",
|
|
"Conversational QA",
|
|
"Multilingual support"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Extractive QA gives verbatim quotes from an authoritative corpus, which compliance contexts demand."
|
|
}
|
|
]
|
|
}
|