Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

91 lines
2.9 KiB
JSON

{
"lesson": "07-end-to-end-fine-tuning-pipeline",
"title": "Capstone 07 — End-to-End Fine-Tuning Pipeline (Data to SFT to DPO to Serve)",
"questions": [
{
"stage": "pre",
"question": "What does the contamination check guard against during data preparation?",
"options": [
"Tokenizer drift between SFT and DPO stages",
"Test-set leakage from public benchmarks such as MMLU-Pro and MT-Bench-v2 into training data",
"GPU driver mismatches between training and serving",
"GGUF version skew across llama.cpp builds"
],
"correct": 1,
"explanation": ""
},
{
"stage": "pre",
"question": "Why does the pipeline compose SFT then DPO (or GRPO) rather than DPO alone?",
"options": [
"DPO cannot run on quantized weights",
"Axolotl does not implement DPO",
"SFT establishes domain behavior on labeled completions while DPO or GRPO aligns the model against preference pairs or verifiable rewards",
"DPO requires a separate base model architecture"
],
"correct": 2,
"explanation": ""
},
{
"stage": "check",
"question": "What does EAGLE-3 contribute to the vLLM serving stage?",
"options": [
"An OPA policy for tool calls",
"An automatic data dedup step",
"A new prompt-caching layer",
"Draft heads that predict N tokens ahead; the target verifies in one pass for 2-3x throughput"
],
"correct": 3,
"explanation": ""
},
{
"stage": "check",
"question": "Which metric reports how well a speculative-decoding draft aligns with the target model?",
"options": [
"Acceptance rate",
"Perplexity",
"PSI",
"Coverage delta"
],
"correct": 0,
"explanation": ""
},
{
"stage": "check",
"question": "Which trio of quants does the pipeline ship for deployment flexibility?",
"options": [
"GPTQ-INT4-Marlin, AWQ-INT4, and GGUF-Q4_K_M",
"FP32, FP16, BF16",
"INT8 only across three runtimes",
"ONNX, CoreML, TFLite"
],
"correct": 0,
"explanation": ""
},
{
"stage": "post",
"question": "Which framework convention does the 2026 model card follow in this capstone?",
"options": [
"HuggingFace YAML front-matter only",
"OpenAI's model card format",
"Datasheets for Datasets",
"Model Openness Framework (MOF) 2026 template covering data, training, eval, safety, license, and reproducibility"
],
"correct": 3,
"explanation": ""
},
{
"stage": "post",
"question": "Which HPA metric is used to autoscale the serving replicas?",
"options": [
"GPU temperature",
"CPU utilization",
"Queue-wait time",
"Network egress bytes"
],
"correct": 2,
"explanation": ""
}
]
}