mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
79 lines
2.7 KiB
JSON
79 lines
2.7 KiB
JSON
{
|
|
"lesson": "17-disaggregated-prefill-decode",
|
|
"title": "Disaggregated Prefill/Decode — NVIDIA Dynamo and llm-d",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "Why do prefill and decode want different optimal GPU configurations?",
|
|
"options": [
|
|
"Prefill must run on AMD and decode on NVIDIA",
|
|
"They use different model weights",
|
|
"Prefill is compute-bound on matmul throughput; decode is memory-bound on HBM bandwidth, so colocating them wastes one resource",
|
|
"Decode requires more network bandwidth"
|
|
],
|
|
"correct": 2,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What transport does NVIDIA Dynamo use to move KV cache between the prefill and decode pools?",
|
|
"options": [
|
|
"gRPC bidi only",
|
|
"NIXL (RDMA/InfiniBand when available, TCP fallback)",
|
|
"Plain HTTP",
|
|
"Shared filesystem on NFS"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "When does disaggregation NOT pay off according to the lesson?",
|
|
"options": [
|
|
"RAG with 8K+ prefixes",
|
|
"MoE workloads on Blackwell",
|
|
"Multi-tenant serving with shared system prompts",
|
|
"Prompts under 512 tokens and outputs under 200 tokens, where the KV transfer tax dominates the gain"
|
|
],
|
|
"correct": 3,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What is the core architectural difference between Dynamo and llm-d?",
|
|
"options": [
|
|
"Dynamo is open source; llm-d is closed",
|
|
"Dynamo runs on CPU; llm-d on GPU",
|
|
"Dynamo is a stack-above orchestrator over vLLM/SGLang/TRT-LLM; llm-d is Kubernetes-native with prefill/decode/router as independent Services",
|
|
"Dynamo requires AMD; llm-d requires NVIDIA"
|
|
],
|
|
"correct": 2,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which Dynamo components automatically tune the prefill:decode ratio for an SLO?",
|
|
"options": [
|
|
"Marlin kernels",
|
|
"Planner Profiler and SLA Planner",
|
|
"Cluster Autoscaler",
|
|
"Sidecar proxy and Envoy filter"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "How does disaggregation interact with cache-aware routing from Phase 17 · 11?",
|
|
"options": [
|
|
"Disaggregation disables KV cache reuse entirely",
|
|
"The cache-aware router can land a request on the decode pool already holding its prefix; on miss it flows prefill -> decode, so the two compound",
|
|
"They are mutually exclusive",
|
|
"Cache-aware routing is only for colocated serving"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
}
|
|
]
|
|
}
|