Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes#368
The 10 audit findings in phase 05 all pointed at SVG assets that were
never created. Nine were broken image embeds (./assets/<name>.svg) and
one was a code-fence false positive where '[tool_name](**args)' inside a
Python snippet looked like a Markdown link to the audit's regex.
This commit:
- Removes the nine broken figure embeds across lessons 01-09. The
surrounding prose stands on its own; no caption text needed rewriting.
- Splits the offending Python expression in lesson 17 onto two lines
(fn = tools[tool_name]; result = fn(**args)) so '](**args)' no longer
appears as adjacent characters.
Three approaches side by side. GloVe as co-occurrence matrix
factorization with weighted log-loss. FastText as sum of character
n-grams (with OOV-tolerant lookup demonstrated). BPE as iterative
pair-merging that produces the subword vocabulary every modern LLM
tokenizer uses.
The teaching thread: Word2Vec had no OOV story. FastText fixed it for
words. BPE solved it for the whole vocabulary at a different level and
became the transformer-era default.
Ship artifact: tokenizer-picker skill with vocab-size targets (32k
English, 64-100k multilingual) and the reproducibility pitfall
(tokenizer-model mismatch is the single most common silent production
bug).
~45 minutes. Prerequisites lesson 05/03.