Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

79 lines
3.2 KiB
JSON

{
"lesson": "32-token-positional-embeddings",
"title": "Token and Positional Embeddings",
"questions": [
{
"stage": "pre",
"question": "What shape does the embedding stage produce from (B, T) ids?",
"options": [
"(B, T, D) where D is the model dimension",
"(B, V)",
"(B, T, V) where V is the vocabulary size",
"(B, T)"
],
"correct": 0,
"explanation": "Each id is mapped to a D-dimensional vector. The batch keeps B and T and gains a new feature dimension D."
},
{
"stage": "check",
"question": "How are token and positional embeddings composed in this lesson?",
"options": [
"By an extra linear projection",
"By taking the elementwise maximum",
"By elementwise sum, broadcasting the position vector across the batch",
"By concatenation along the feature axis"
],
"correct": 2,
"explanation": "The two tensors of shape (B, T, D) and (T, D) are summed. Broadcasting replicates the positional table across the batch axis."
},
{
"stage": "check",
"question": "How many learnable parameters does a SinusoidalPositionalEmbedding contribute?",
"options": [
"max_context_length * d_model",
"Zero",
"max_context_length",
"d_model"
],
"correct": 1,
"explanation": "The sinusoidal table is computed from a formula and stored as a buffer. It contributes no parameters."
},
{
"stage": "check",
"question": "Which positional encoding cannot accept a sequence longer than its construction limit?",
"options": [
"Sinusoidal positional embedding",
"Neither",
"Both",
"Learned positional embedding"
],
"correct": 2,
"explanation": "In this lesson implementation, both classes enforce max_context_length at forward time. Learned is inherently table-bounded; sinusoidal here is also bounded because it uses a prebuilt table."
},
{
"stage": "post",
"question": "Why does the sinusoidal scheme use both sin and cos on adjacent feature pairs?",
"options": [
"So that the vector at position p+k is a linear function of the vector at position p",
"Because cos by itself cannot be backpropagated through",
"Because sin alone is not periodic",
"Because it doubles the parameter count"
],
"correct": 0,
"explanation": "Pairing sin and cos at the same wavelength means a position shift by k is a rotation, which is a linear map. Attention can then learn relative offsets through linear projections."
},
{
"stage": "post",
"question": "Why does the embedding backward pass only update rows of the token table that appeared in the batch?",
"options": [
"Rows not touched in the forward pass produce zero gradient by the chain rule",
"The optimizer skips rows it has not seen",
"PyTorch only stores gradients for non-zero entries",
"Token ids are stored in a sparse data structure"
],
"correct": 0,
"explanation": "A row that did not contribute to any output has zero derivative for the loss, so its gradient is zero for that step."
}
]
}