Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

79 lines
2.6 KiB
JSON

{
"lesson": "84-refusal-evaluation",
"title": "Capstone 84 — Refusal Evaluation",
"questions": [
{
"stage": "pre",
"question": "What does under-refusal measure?",
"options": [
"How often the model refuses prompts labeled safe",
"The cost of the API call per prompt",
"How often the model answers prompts labeled unsafe",
"How many tokens the model emits per response"
],
"correct": 2,
"explanation": ""
},
{
"stage": "pre",
"question": "Why use a mock LLM rather than a real model for this lesson?",
"options": [
"Because real models are not allowed in capstones",
"Because mock LLMs are more accurate than real ones",
"So the same input yields the same output across runs, making code changes attributable to known causes",
"Because the lesson requires GPU hardware"
],
"correct": 2,
"explanation": ""
},
{
"stage": "check",
"question": "Given safe prompt count = 30, refused safe count = 5, what is over-refusal rate?",
"options": [
"0.30",
"0.17",
"0.50",
"0.05"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "What does Expected Calibration Error (ECE) measure here?",
"options": [
"The fraction of refusals on unsafe prompts",
"The gap between the model's stated confidence and its observed accuracy, binned across confidence ranges",
"The total token count across all responses",
"The number of regex rules that fired"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "Why does the framework join unsafe prompts against the lesson 82 taxonomy?",
"options": [
"To re-run the lesson 82 corpus loader",
"To raise the severity of every prompt by one",
"To report under-refusal per attack category so the team sees which boundary the model leaks on",
"To remove duplicates"
],
"correct": 2,
"explanation": ""
},
{
"stage": "post",
"question": "Why does the framework return both under-refusal and over-refusal rather than a single safety score?",
"options": [
"Because under-refusal is only used in CI",
"Because they are two opposite errors and a single number hides whichever one is worse on a given build",
"Because over-refusal is computed by a different team",
"Because Python dataclasses prefer multiple fields"
],
"correct": 1,
"explanation": ""
}
]
}