mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
79 lines
2.7 KiB
JSON
79 lines
2.7 KiB
JSON
{
|
|
"lesson": "03-gpu-autoscaling-kubernetes",
|
|
"title": "GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "Which signal does HPA typically scale on by default that the lesson calls broken for vLLM-style serving?",
|
|
"options": [
|
|
"P99 TTFT",
|
|
"DCGM_FI_DEV_GPU_UTIL duty cycle",
|
|
"Queue depth",
|
|
"KV cache utilization"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Which problem does gang scheduling in KAI Scheduler primarily prevent?",
|
|
"options": [
|
|
"Tokenizer GIL contention",
|
|
"The partial-allocation trap where 7 of 8 GPUs sit idle waiting on the eighth",
|
|
"Cold-start latency",
|
|
"GPU memory fragmentation"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why is Karpenter's default consolidationPolicy WhenEmptyOrUnderutilized dangerous for inference GPU pools?",
|
|
"options": [
|
|
"It only consolidates spot instances",
|
|
"It prevents Karpenter from provisioning new nodes",
|
|
"It terminates running GPU nodes to migrate pods, which evicts running requests and reloads weights",
|
|
"It refuses to scale up under burst"
|
|
],
|
|
"correct": 2,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Roughly how much faster is Karpenter at provisioning a GPU node compared to Cluster Autoscaler?",
|
|
"options": [
|
|
"About 5% faster",
|
|
"The same",
|
|
"Roughly 40% faster (~45-60s vs ~90-120s)",
|
|
"About 10x slower"
|
|
],
|
|
"correct": 2,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "For disaggregated prefill / decode pods (Phase 17 · 17), which scaling signals does the lesson recommend?",
|
|
"options": [
|
|
"Cluster Autoscaler for both",
|
|
"A single HPA on duty cycle covering both pod classes",
|
|
"Manual scaling only",
|
|
"Queue depth for prefill pods and KV cache pressure for decode pods, as separate per-role HPAs"
|
|
],
|
|
"correct": 3,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which Karpenter disruption setting does the lesson recommend for an inference GPU pool to avoid evicting running jobs?",
|
|
"options": [
|
|
"Always run with spot instances and no consolidation",
|
|
"Disable Karpenter entirely",
|
|
"consolidationPolicy: WhenEmptyOrUnderutilized with consolidateAfter: 0s",
|
|
"consolidationPolicy: WhenEmpty with consolidateAfter: 1h"
|
|
],
|
|
"correct": 3,
|
|
"explanation": ""
|
|
}
|
|
]
|
|
}
|