Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

103 lines
3.9 KiB
JSON

{
"lesson": "29-dialogue-state-tracking",
"title": "Dialogue State Tracking",
"questions": [
{
"stage": "pre",
"question": "What does dialogue state tracking maintain across turns?",
"options": [
"A free-text history only",
"A POS-tagged transcript",
"An embedding of the conversation",
"A slot-value dictionary representing the user's current goal, updated after every turn"
],
"correct": 3,
"explanation": "DST keeps a structured slot-value map that the backend can act on."
},
{
"stage": "pre",
"question": "What does Joint Goal Accuracy (JGA) measure?",
"options": [
"Fraction of turns where every slot is exactly correct (all-or-nothing)",
"Cosine similarity",
"Latency per turn",
"Average slot accuracy"
],
"correct": 0,
"explanation": "JGA is the strict per-turn match across all slots; per-slot accuracy is more lenient."
},
{
"stage": "check",
"question": "Why does regenerating the whole state from history each turn handle user corrections naturally?",
"options": [
"It uses fewer tokens",
"It runs on GPU",
"It avoids embeddings",
"Reading the full history lets the model re-derive the final state including 'actually...' corrections without explicit rollback logic"
],
"correct": 3,
"explanation": "Full-history regeneration absorbs corrections by recomputing the final state from the entire conversation."
},
{
"stage": "check",
"question": "Which 2026 pattern gives a guaranteed-valid slot dict in 5 lines of code?",
"options": [
"Hand-written regex",
"LLM + Instructor + Pydantic schema with constrained or validated output",
"TF-IDF classifier",
"BM25 retrieval"
],
"correct": 1,
"explanation": "Pydantic schema + Instructor validates the LLM's state output against the slot ontology automatically."
},
{
"stage": "check",
"question": "Why version your DST schema?",
"options": [
"Adding new slots post-hoc invalidates older training data and breaks longitudinal evaluation",
"Reduce token count",
"Required by JSON",
"Speed gains"
],
"correct": 0,
"explanation": "Unversioned schema changes silently break training data alignment and eval comparability."
},
{
"stage": "post",
"question": "Why must DST for compliance-sensitive domains include a rule-based check alongside LLM extraction?",
"options": [
"LLMs are slower",
"Rules avoid embeddings",
"LLM-only DST can mis-extract destructive parameters (amount, account, date); a rules layer enforces deterministic constraints",
"Rules are multilingual"
],
"correct": 2,
"explanation": "Compliance domains require deterministic enforcement; rules catch slot errors that LLMs introduce."
},
{
"stage": "post",
"question": "What is the cost concern with regenerating state on every turn via LLM?",
"options": [
"Embedding drift",
"More embeddings",
"Re-reading the full history each turn yields O(n^2) total token usage; cap or summarize older turns",
"Cosine costs"
],
"correct": 2,
"explanation": "Full-history regeneration is quadratic in turns; cap history or use rolling summaries."
},
{
"stage": "post",
"question": "Why are explicit confirmation flows required before destructive backend actions?",
"options": [
"Confirmation increases JGA",
"Even good DST has nonzero slot-error rates; a deterministic confirmation prevents wrong-account or wrong-amount actions",
"Required by tokenizers",
"Latency"
],
"correct": 1,
"explanation": "Destructive actions need user confirmation because DST is never error-free."
}
]
}