Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

91 lines
3.1 KiB
JSON

{
"lesson": "09-code-migration-agent",
"title": "Capstone 09 — Code Migration Agent (Repo-Level Language / Runtime Upgrade)",
"questions": [
{
"stage": "pre",
"question": "Why does the pipeline combine a deterministic substrate with an agent layer rather than using just one?",
"options": [
"Agents alone are faster than recipes",
"OpenRewrite or libcst handles 70-80% of mechanical rewrites safely and cheaply, leaving the agent for the ambiguous long tail",
"Determinism is needed only for the build system",
"Deterministic recipes are slower than LLMs"
],
"correct": 1,
"explanation": ""
},
{
"stage": "pre",
"question": "What signal does the pipeline use as ground truth for a successful migration?",
"options": [
"Agent self-grading on a rubric",
"Reviewer approval on the PR",
"Diff size below a hard limit",
"Green CI in the sandbox without a coverage regression beyond a small threshold"
],
"correct": 3,
"explanation": ""
},
{
"stage": "check",
"question": "Which budget caps does the agent loop enforce per repo?",
"options": [
"30 minutes wall-clock, $8 cost, and 20 agent turns",
"1 hour wall-clock and 100 turns",
"Unlimited time and cost; abort only on errors",
"No turn limit, $100 ceiling"
],
"correct": 0,
"explanation": ""
},
{
"stage": "check",
"question": "What gate fires when coverage drops more than about 2% after migration?",
"options": [
"The deterministic substrate replays its recipes",
"The repo gets filed under a coverage_regression failure class instead of opening a clean PR",
"The agent automatically force-pushes a fix",
"The reviewer is bypassed"
],
"correct": 1,
"explanation": ""
},
{
"stage": "check",
"question": "Why is the failure taxonomy treated as a deliverable rather than a side artifact?",
"options": [
"It replaces the test suite",
"It is required by GitHub branch protection",
"It satisfies a compliance checklist",
"It groups failed repos by class so future recipe authors can target the top failure modes"
],
"correct": 3,
"explanation": ""
},
{
"stage": "post",
"question": "Which public benchmark does the capstone target for Java 8 to 17 migration?",
"options": [
"MMLU-Pro",
"MigrationBench from Amazon",
"ViDoRe v3",
"SWE-bench Pro"
],
"correct": 1,
"explanation": ""
},
{
"stage": "post",
"question": "What does the agent integration rubric measure about the fix distribution?",
"options": [
"Number of dependencies pinned",
"The number of force-pushes per repo",
"The fraction of fixes handled by OpenRewrite versus authored by the agent layer",
"Tokenization speed of the source file"
],
"correct": 2,
"explanation": ""
}
]
}