mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
103 lines
3.9 KiB
JSON
103 lines
3.9 KiB
JSON
{
|
|
"lesson": "29-dialogue-state-tracking",
|
|
"title": "Dialogue State Tracking",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "What does dialogue state tracking maintain across turns?",
|
|
"options": [
|
|
"A free-text history only",
|
|
"A POS-tagged transcript",
|
|
"An embedding of the conversation",
|
|
"A slot-value dictionary representing the user's current goal, updated after every turn"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "DST keeps a structured slot-value map that the backend can act on."
|
|
},
|
|
{
|
|
"stage": "pre",
|
|
"question": "What does Joint Goal Accuracy (JGA) measure?",
|
|
"options": [
|
|
"Fraction of turns where every slot is exactly correct (all-or-nothing)",
|
|
"Cosine similarity",
|
|
"Latency per turn",
|
|
"Average slot accuracy"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "JGA is the strict per-turn match across all slots; per-slot accuracy is more lenient."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why does regenerating the whole state from history each turn handle user corrections naturally?",
|
|
"options": [
|
|
"It uses fewer tokens",
|
|
"It runs on GPU",
|
|
"It avoids embeddings",
|
|
"Reading the full history lets the model re-derive the final state including 'actually...' corrections without explicit rollback logic"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Full-history regeneration absorbs corrections by recomputing the final state from the entire conversation."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Which 2026 pattern gives a guaranteed-valid slot dict in 5 lines of code?",
|
|
"options": [
|
|
"Hand-written regex",
|
|
"LLM + Instructor + Pydantic schema with constrained or validated output",
|
|
"TF-IDF classifier",
|
|
"BM25 retrieval"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Pydantic schema + Instructor validates the LLM's state output against the slot ontology automatically."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why version your DST schema?",
|
|
"options": [
|
|
"Adding new slots post-hoc invalidates older training data and breaks longitudinal evaluation",
|
|
"Reduce token count",
|
|
"Required by JSON",
|
|
"Speed gains"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Unversioned schema changes silently break training data alignment and eval comparability."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Why must DST for compliance-sensitive domains include a rule-based check alongside LLM extraction?",
|
|
"options": [
|
|
"LLMs are slower",
|
|
"Rules avoid embeddings",
|
|
"LLM-only DST can mis-extract destructive parameters (amount, account, date); a rules layer enforces deterministic constraints",
|
|
"Rules are multilingual"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "Compliance domains require deterministic enforcement; rules catch slot errors that LLMs introduce."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "What is the cost concern with regenerating state on every turn via LLM?",
|
|
"options": [
|
|
"Embedding drift",
|
|
"More embeddings",
|
|
"Re-reading the full history each turn yields O(n^2) total token usage; cap or summarize older turns",
|
|
"Cosine costs"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "Full-history regeneration is quadratic in turns; cap history or use rolling summaries."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Why are explicit confirmation flows required before destructive backend actions?",
|
|
"options": [
|
|
"Confirmation increases JGA",
|
|
"Even good DST has nonzero slot-error rates; a deterministic confirmation prevents wrong-account or wrong-amount actions",
|
|
"Required by tokenizers",
|
|
"Latency"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Destructive actions need user confirmation because DST is never error-free."
|
|
}
|
|
]
|
|
}
|