mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
79 lines
3.3 KiB
JSON
79 lines
3.3 KiB
JSON
{
|
|
"lesson": "61-cross-attention-fusion",
|
|
"title": "Cross-Attention Fusion",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "What is the operational difference between early fusion and late fusion in vision-language modeling?",
|
|
"options": [
|
|
"Late fusion uses more memory",
|
|
"Early fusion is illegal",
|
|
"Early fusion concatenates image and text tokens into one sequence; late fusion keeps them separate and bridges via cross-attention at each block",
|
|
"Early fusion runs first in wall-clock time"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "Chameleon and Emu3 are early-fusion; Flamingo and BLIP-2 are late-fusion. The architectural choice changes mask shapes and KV caching."
|
|
},
|
|
{
|
|
"stage": "pre",
|
|
"question": "In a decoder block with both self-attention and cross-attention, which one uses a causal mask?",
|
|
"options": [
|
|
"Neither",
|
|
"Both",
|
|
"Cross-attention only",
|
|
"Self-attention only; the image is fully observed and cross-attention has no temporal order"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "Text generation is autoregressive, so text self-attention is causal. Image tokens are all visible before any text is decoded."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why is the cross-attention KV cache built once per image?",
|
|
"options": [
|
|
"Image keys and values do not change as text is decoded, so projecting memory once and reusing the K and V across all decode steps is the whole inference speedup",
|
|
"It avoids using the GPU",
|
|
"PyTorch requires caching",
|
|
"Caching saves disk space"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "The vision encoder runs once. Its KV projection is computed once. Each text token reuses the cache for free."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What is the output shape of cross-attention when text length is Nt=10 and image length is Nv=197 with hidden=256?",
|
|
"options": [
|
|
"(B, 10, 256)",
|
|
"(B, 207, 256)",
|
|
"(B, 10, 197)",
|
|
"(B, 197, 256)"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Cross-attention output shape follows the query stream, so output is (B, Nt, hidden) regardless of key length."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why is no mask used on cross-attention?",
|
|
"options": [
|
|
"Masks always cause NaN",
|
|
"Masks are too expensive",
|
|
"The image is fully observed before text generation begins; every text position may attend to every patch",
|
|
"PyTorch does not support cross masks"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "There is no temporal order on image patches and no leakage to prevent, so cross-attention sees the whole image."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which extension would you add to recover Flamingo-style stability during training?",
|
|
"options": [
|
|
"Remove the FFN",
|
|
"Insert a learned tanh gate on the cross-attention residual so the model can start from text-only behavior and grow into the image stream",
|
|
"Drop the self-attention layer",
|
|
"Use only one attention head"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "Flamingo's tanh gate starts at zero, recovers text-only behavior at init, and gives the model a smooth ramp into cross-attention."
|
|
}
|
|
]
|
|
}
|