Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

68 lines
3.3 KiB
JSON

[
{
"id": "nb-pre-1",
"stage": "pre",
"question": "What is the 'naive' assumption in Naive Bayes?",
"options": [
"The prior probability of each class is equal",
"All features are conditionally independent given the class label",
"The data is normally distributed",
"The model has no parameters to learn"
],
"correct": 1,
"explanation": "Naive Bayes assumes every feature is independent of every other feature, conditioned on the class. This is mathematically wrong (e.g., 'machine' and 'learning' co-occur) but works well in practice."
},
{
"id": "nb-pre-2",
"stage": "pre",
"question": "What does Laplace smoothing prevent in Naive Bayes?",
"options": [
"Zero probabilities for words never seen in a class during training",
"Class imbalance in the dataset",
"Slow training on high-dimensional data",
"Overfitting to large datasets"
],
"correct": 0,
"explanation": "Without smoothing, a word that never appeared in 'spam' training emails would get P(word|spam) = 0, making the entire product zero regardless of other strong evidence. Laplace smoothing adds 1 to each count."
},
{
"id": "nb-post-1",
"stage": "post",
"question": "The naive independence assumption is clearly wrong for text. Why does Naive Bayes still classify well?",
"options": [
"It only works on very small vocabularies where independence holds",
"It only works when features are truly independent",
"Modern implementations secretly remove the independence assumption",
"Classification only needs correct class rankings, not correct probability estimates, and the assumption introduces stable errors that affect all classes similarly"
],
"correct": 3,
"explanation": "NB needs to rank classes correctly, not estimate exact probabilities. The independence assumption is high bias but low variance, making it stable with limited data. Correlated features double-count evidence for the correct class too."
},
{
"id": "nb-post-2",
"stage": "post",
"question": "When should you use Multinomial NB versus Gaussian NB?",
"options": [
"Multinomial for word count/frequency features, Gaussian for continuous real-valued features",
"They are interchangeable -- always use whichever is faster",
"Multinomial for binary data, Gaussian for multi-class problems",
"Multinomial for regression, Gaussian for classification"
],
"correct": 0,
"explanation": "Multinomial NB models feature counts (word frequencies in text). Gaussian NB assumes features follow normal distributions, suitable for continuous features like measurements or sensor readings."
},
{
"id": "nb-post-3",
"stage": "post",
"question": "An email contains 'free' twice and 'money' once. In Multinomial NB with log probabilities, how is the spam score computed?",
"options": [
"P(spam) * P(free|spam) * P(money|spam)",
"log P(spam) * 2 * log P(free|spam)",
"log P(spam) + log P(free|spam) + log P(money|spam)",
"log P(spam) + 2 * log P(free|spam) + 1 * log P(money|spam)"
],
"correct": 3,
"explanation": "Multinomial NB multiplies the word likelihoods raised to their count. In log space: log P(spam) + 2*log P(free|spam) + 1*log P(money|spam). The word count acts as an exponent."
}
]