mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
103 lines
4.2 KiB
JSON
103 lines
4.2 KiB
JSON
{
|
|
"lesson": "18-multilingual-nlp",
|
|
"title": "Multilingual NLP",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "What does zero-shot cross-lingual transfer mean?",
|
|
"options": [
|
|
"Translating without a translation model",
|
|
"Tokenizing with zero merges",
|
|
"Fine-tune a multilingual model on one source language and evaluate on a different language with no target-language labels",
|
|
"Training with zero examples"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "Zero-shot transfer: train on the source language, run on the target without target-language supervision."
|
|
},
|
|
{
|
|
"stage": "pre",
|
|
"question": "Which model family ships as the standard 100-language cross-lingual baseline?",
|
|
"options": [
|
|
"GPT-2",
|
|
"GloVe",
|
|
"DistilBERT",
|
|
"XLM-R (e.g. XLM-RoBERTa-base, 270M)"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "XLM-R is the canonical 100-language pretrained baseline for cross-lingual classification."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why does English-as-source not always give the best transfer for a non-English target?",
|
|
"options": [
|
|
"English has too little data",
|
|
"Language similarity (typology, script, morphology) predicts transfer quality; a closer high-resource source can outperform English",
|
|
"English uses BPE",
|
|
"English is too short"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Typologically related sources (e.g. Hindi for Indic targets) often outperform English as a fine-tune source."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What is the 'fertility tax' for low-resource languages?",
|
|
"options": [
|
|
"Smaller models train slower",
|
|
"Tokenizers cannot handle Unicode",
|
|
"BPE refuses to train",
|
|
"Low-resource text tokenizes into more subwords per word than English, consuming context window, latency, and capacity"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Long-tail languages tokenize at much higher fertility, eating context and training efficiency."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why is per-language evaluation required, not aggregated accuracy?",
|
|
"options": [
|
|
"Aggregates run faster",
|
|
"Aggregates only work on classification",
|
|
"Aggregates ignore tokenization",
|
|
"Aggregate numbers hide long-tail languages where a multilingual model can be far worse than its mean suggests"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Aggregate accuracy masks poor performance on low-resource languages; per-language scores expose it."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Why is fine-tuning learning rate critical when adapting a multilingual model with few-shot data?",
|
|
"options": [
|
|
"High LR can collapse the multilingual alignment and effectively reduce the model to English-only",
|
|
"Lower LR wastes GPU",
|
|
"It changes the vocabulary",
|
|
"Required by tokenizers"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Excessive LR drifts the shared representation; conservative LR (~2e-5) preserves cross-lingual structure."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which mitigation directly addresses tokenizer fertility for long-tail scripts?",
|
|
"options": [
|
|
"Skip stopwords",
|
|
"Use byte-fallback (SentencePiece byte_fallback=True) or a tokenizer with broader script coverage (e.g. XLM-V)",
|
|
"More training epochs",
|
|
"Lower batch size"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Byte fallback and broader-vocab tokenizers reduce fertility and OOV for low-resource scripts."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "When is a monolingual model from scratch worth trying instead of a multilingual one?",
|
|
"options": [
|
|
"Only for translation",
|
|
"Whenever the tokenizer is BPE",
|
|
"Always for English",
|
|
"When the target language has enough data to train a monolingual model that beats the multilingual baseline; test before assuming"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Sometimes monolingual training beats multilingual for high-resource targets; empirical comparison decides."
|
|
}
|
|
]
|
|
}
|