Files
Rohit Ghumare dda194f840 fix(quiz): correct answer is always in the same position (slot B) (#381)
Every "Test Your Understanding" quiz placed the correct answer in option B.
Across the 2026 questions in 338 quiz files the correct answer sat at index 1
in 61.5% of cases (uniform would be ~25%), and 107 files had every answer at B,
making the quizzes guessable without reading them.

scripts/debias_quizzes.py rewrites each question's option order with a
deterministic, content-seeded permutation and updates the correct index to
follow the moved answer. It is idempotent: options are canonicalised to a sorted
base before permuting, so re-running produces byte-identical output. Questions
whose options reference each other by position ("all of the above", "both A and
B") are left untouched. The correct-answer value, the option set, and every
explanation are preserved exactly; only order and the index change.

Result: A 23.8% / B 26.3% / C 23.5% / D 26.4%.

The script doubles as a CI guard: `--check` exits non-zero if any quiz is not
de-biased, wired into the curriculum workflow so new lessons cannot regress.

Fixes #368
2026-08-01 14:24:15 +01:00

79 lines
2.9 KiB
JSON

{
"lesson": "14-ascii-art-visual-jailbreaks",
"title": "ASCII Art and Visual Jailbreaks",
"questions": [
{
"stage": "pre",
"question": "Why do token-level and semantic-level safety filters miss ArtPrompt-style attacks?",
"options": [
"ArtPrompt always uses base64 encoding",
"ArtPrompt requires fine-tuning the model",
"ArtPrompt operates at the visual recognition level: the filter sees harmless punctuation while the model reads the rendered letters as a word",
"ArtPrompt is only effective on multimodal models"
],
"correct": 2,
"explanation": ""
},
{
"stage": "check",
"question": "What are the two steps of ArtPrompt as described in Jiang et al. (ACL 2024)?",
"options": [
"Tokenizer mismatch, then context truncation",
"Random search, then perplexity boosting",
"Translation to French, then back-translation",
"Word identification of safety-relevant tokens, then ASCII-art substitution and cloaked-prompt generation"
],
"correct": 3,
"explanation": ""
},
{
"stage": "check",
"question": "Why does perplexity-filter (PPL) defense fail on ArtPrompt?",
"options": [
"Perplexity filters are disabled by default",
"Perplexity is not measurable on ASCII art",
"Perplexity always equals zero on cloaked input",
"ASCII art has high perplexity, but so does any legitimate structured input; thresholds that block the attack also block legitimate content"
],
"correct": 3,
"explanation": ""
},
{
"stage": "check",
"question": "What does the ViTC benchmark measure?",
"options": [
"The model's training-data toxicity",
"The model's tool-calling accuracy",
"The model's ability to read non-semantic visual prompts (ASCII art, wingdings, similar encoded text)",
"The model's ability to compress vision data"
],
"correct": 2,
"explanation": ""
},
{
"stage": "post",
"question": "What does StructuralSleight generalize ArtPrompt to?",
"options": [
"Image inputs only",
"Long-context shot stuffing",
"Adversarial token suffixes via gradient search",
"Uncommon Text-Encoded Structures (UTES): trees, graphs, nested JSON, CSV-in-JSON, diff-style code, etc."
],
"correct": 3,
"explanation": ""
},
{
"stage": "post",
"question": "What is the capability-safety tradeoff implied by ViTC results?",
"options": [
"Larger context windows reduce ArtPrompt success",
"Visual recognition is unrelated to ArtPrompt success",
"The better the model reads non-semantic visual text, the more effective ArtPrompt-style attacks are on it",
"Smaller models are always more secure"
],
"correct": 2,
"explanation": ""
}
]
}