mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
91 lines
3.5 KiB
JSON
91 lines
3.5 KiB
JSON
{
|
|
"lesson": "04-tree-of-thoughts-lats",
|
|
"title": "Tree of Thoughts and LATS: Deliberate Search",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "Why does chain-of-thought struggle on Game of 24?",
|
|
"options": [
|
|
"A linear walk cannot backtrack when an early step is wrong, so later steps compound the error",
|
|
"CoT requires a calculator tool that GPT-4 lacks",
|
|
"The model cannot multiply integers",
|
|
"The prompt is too short"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "Without branching, a wrong early subexpression poisons the rest of the chain; the paper measures only 4 percent for CoT."
|
|
},
|
|
{
|
|
"stage": "pre",
|
|
"question": "What is a node in a Tree of Thoughts search?",
|
|
"options": [
|
|
"A tool registered with the runtime",
|
|
"A weight update during fine-tuning",
|
|
"A token produced by the model",
|
|
"A coherent intermediate step or thought, with K possible child expansions"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "ToT treats reasoning as a tree where each node is an intermediate thought that can expand into K children."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Which of these three is NOT one of LATS's roles for the LLM?",
|
|
"options": [
|
|
"Self-reflector that writes reflections on failure",
|
|
"Optimizer that updates model weights between rollouts",
|
|
"Value function that scores partial trajectories",
|
|
"Policy that proposes next actions"
|
|
],
|
|
"correct": 1,
|
|
"explanation": "LATS is gradient-free; the three LLM roles are policy, value, and self-reflector. There are no weight updates."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Name the four MCTS phases the lesson lists.",
|
|
"options": [
|
|
"Sample, Score, Sort, Submit",
|
|
"Plan, Execute, Reflect, Stop",
|
|
"Select, Expand, Simulate, Backpropagate",
|
|
"Search, Synthesize, Synthesize-Again, Stop"
|
|
],
|
|
"correct": 2,
|
|
"explanation": "MCTS proceeds in select, expand, simulate, backpropagate per iteration."
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "In UCT, what is the role of the exploration constant c?",
|
|
"options": [
|
|
"It scales the value estimate Q",
|
|
"It sets the maximum tree depth",
|
|
"It controls the number of rollouts",
|
|
"It weights the exploration term sqrt(ln N / n) against the exploitation term Q"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "c balances exploitation (Q) against the exploration term; tune per task."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "When is search actively harmful compared to a single trajectory?",
|
|
"options": [
|
|
"Whenever tokens are cheap",
|
|
"When the task is code generation",
|
|
"When the task involves multiple correct answers",
|
|
"When the evaluator is noisy and there is a single right answer, so the search converges on a good-scoring wrong answer"
|
|
],
|
|
"correct": 3,
|
|
"explanation": "A noisy value function plus a single correct answer is exactly when search overfits to the noise."
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Roughly how much more token usage should you budget for ToT on Game of 24 compared with CoT?",
|
|
"options": [
|
|
"100x to 1000x",
|
|
"Less than CoT because of pruning",
|
|
"About 2x",
|
|
"About 10x"
|
|
],
|
|
"correct": 0,
|
|
"explanation": "The lesson cites 100-1000x token cost for ToT on Game of 24 versus CoT."
|
|
}
|
|
]
|
|
}
|