mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.
scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.
Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.
The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.
Fixes #368
79 lines
2.7 KiB
JSON
79 lines
2.7 KiB
JSON
{
|
|
"lesson": "07-tensorrt-llm-blackwell",
|
|
"title": "Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell",
|
|
"questions": [
|
|
{
|
|
"stage": "pre",
|
|
"question": "Roughly what is the per-million-tokens cost gap the lesson reports between Blackwell + TRT-LLM + Dynamo and H100 + vLLM on a comparable 120B-class workload?",
|
|
"options": [
|
|
"About 2x",
|
|
"About 7x",
|
|
"About 1.1x",
|
|
"About 100x"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Why does the lesson recommend keeping KV cache in FP8 rather than NVFP4 on Blackwell?",
|
|
"options": [
|
|
"NVFP4 KV cache is not yet supported in any engine",
|
|
"KV cache spans a wide dynamic range; FP4 quantization causes catastrophic accuracy loss in attention scores",
|
|
"FP8 uses less memory than FP4",
|
|
"FP8 is the only precision NVLink 5 supports"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "Which Blackwell feature does TRT-LLM exploit so models can be loaded without a post-training conversion step?",
|
|
"options": [
|
|
"BF16 KV cache",
|
|
"FP64 attention",
|
|
"INT2 weights via bitsandbytes",
|
|
"Day-0 FP4 weights shipped by model providers"
|
|
],
|
|
"correct": 3,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "check",
|
|
"question": "What is the dominant tradeoff of choosing the TRT-LLM stack per the lesson?",
|
|
"options": [
|
|
"It requires fully autonomous remediation",
|
|
"It cannot serve MoE models",
|
|
"It only works at small scale",
|
|
"It locks you into NVIDIA hardware — no AMD, no Intel, no ARM"
|
|
],
|
|
"correct": 3,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "Which precision combination does the lesson describe as the typical Blackwell config?",
|
|
"options": [
|
|
"Everything in BF16",
|
|
"Weights NVFP4, activations NVFP4, KV cache FP8, attention accumulator FP32",
|
|
"Weights INT8, activations FP32, KV cache INT4",
|
|
"Weights FP4, KV cache FP4, attention in INT8"
|
|
],
|
|
"correct": 1,
|
|
"explanation": ""
|
|
},
|
|
{
|
|
"stage": "post",
|
|
"question": "For reasoning-heavy workloads where NVFP4 weight conversion drops MATH accuracy a few points, what does the lesson advise?",
|
|
"options": [
|
|
"Switch to AMD MI300X",
|
|
"Disable speculative decoding",
|
|
"Validate task quality on your eval set per model; teams often use FP8 weights + FP4 activations or stay on H200 with FP8 throughout",
|
|
"Ship NVFP4 anyway because the cost win dominates"
|
|
],
|
|
"correct": 2,
|
|
"explanation": ""
|
|
}
|
|
]
|
|
}
|