Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

79 lines
2.7 KiB
JSON

{
"lesson": "03-gpu-autoscaling-kubernetes",
"title": "GPU Autoscaling on Kubernetes — Karpenter, KAI Scheduler, Gang Scheduling",
"questions": [
{
"stage": "pre",
"question": "Which signal does HPA typically scale on by default that the lesson calls broken for vLLM-style serving?",
"options": [
"P99 TTFT",
"DCGM_FI_DEV_GPU_UTIL duty cycle",
"Queue depth",
"KV cache utilization"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "Which problem does gang scheduling in KAI Scheduler primarily prevent?",
"options": [
"Tokenizer GIL contention",
"The partial-allocation trap where 7 of 8 GPUs sit idle waiting on the eighth",
"Cold-start latency",
"GPU memory fragmentation"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "Why is Karpenter's default consolidationPolicy WhenEmptyOrUnderutilized dangerous for inference GPU pools?",
"options": [
"It only consolidates spot instances",
"It prevents Karpenter from provisioning new nodes",
"It terminates running GPU nodes to migrate pods, which evicts running requests and reloads weights",
"It refuses to scale up under burst"
],
"correct": 2,
"explanation": ""
},
{
"stage": "check",
"question": "Roughly how much faster is Karpenter at provisioning a GPU node compared to Cluster Autoscaler?",
"options": [
"About 5% faster",
"The same",
"Roughly 40% faster (~45-60s vs ~90-120s)",
"About 10x slower"
],
"correct": 2,
"explanation": ""
},
{
"stage": "post",
"question": "For disaggregated prefill / decode pods (Phase 17 · 17), which scaling signals does the lesson recommend?",
"options": [
"Cluster Autoscaler for both",
"A single HPA on duty cycle covering both pod classes",
"Manual scaling only",
"Queue depth for prefill pods and KV cache pressure for decode pods, as separate per-role HPAs"
],
"correct": 3,
"explanation": ""
},
{
"stage": "post",
"question": "Which Karpenter disruption setting does the lesson recommend for an inference GPU pool to avoid evicting running jobs?",
"options": [
"Always run with spot instances and no consolidation",
"Disable Karpenter entirely",
"consolidationPolicy: WhenEmptyOrUnderutilized with consolidateAfter: 0s",
"consolidationPolicy: WhenEmpty with consolidateAfter: 1h"
],
"correct": 3,
"explanation": ""
}
]
}