Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

79 lines
2.7 KiB
JSON

{
"lesson": "07-tensorrt-llm-blackwell",
"title": "Hardware-Specialized Inference Compilation — FP8 and NVFP4 on Blackwell",
"questions": [
{
"stage": "pre",
"question": "Roughly what is the per-million-tokens cost gap the lesson reports between Blackwell + TRT-LLM + Dynamo and H100 + vLLM on a comparable 120B-class workload?",
"options": [
"About 2x",
"About 7x",
"About 1.1x",
"About 100x"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "Why does the lesson recommend keeping KV cache in FP8 rather than NVFP4 on Blackwell?",
"options": [
"NVFP4 KV cache is not yet supported in any engine",
"KV cache spans a wide dynamic range; FP4 quantization causes catastrophic accuracy loss in attention scores",
"FP8 uses less memory than FP4",
"FP8 is the only precision NVLink 5 supports"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "Which Blackwell feature does TRT-LLM exploit so models can be loaded without a post-training conversion step?",
"options": [
"BF16 KV cache",
"FP64 attention",
"INT2 weights via bitsandbytes",
"Day-0 FP4 weights shipped by model providers"
],
"correct": 3,
"explanation": ""
},
{
"stage": "check",
"question": "What is the dominant tradeoff of choosing the TRT-LLM stack per the lesson?",
"options": [
"It requires fully autonomous remediation",
"It cannot serve MoE models",
"It only works at small scale",
"It locks you into NVIDIA hardware — no AMD, no Intel, no ARM"
],
"correct": 3,
"explanation": ""
},
{
"stage": "post",
"question": "Which precision combination does the lesson describe as the typical Blackwell config?",
"options": [
"Everything in BF16",
"Weights NVFP4, activations NVFP4, KV cache FP8, attention accumulator FP32",
"Weights INT8, activations FP32, KV cache INT4",
"Weights FP4, KV cache FP4, attention in INT8"
],
"correct": 1,
"explanation": ""
},
{
"stage": "post",
"question": "For reasoning-heavy workloads where NVFP4 weight conversion drops MATH accuracy a few points, what does the lesson advise?",
"options": [
"Switch to AMD MI300X",
"Disable speculative decoding",
"Validate task quality on your eval set per model; teams often use FP8 weights + FP4 activations or stay on H200 with FP8 throughout",
"Ship NVFP4 anyway because the cost win dominates"
],
"correct": 2,
"explanation": ""
}
]
}