mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Two-sided refusal metrics: under-refusal (answered unsafe), over-refusal (refused safe), accuracy, ECE calibration, per-category under-refusal join against the lesson 82 taxonomy. Three mock LLM policies (strict, leaky, over-cautious) demonstrate the framework detects opposite failure modes. Labeled corpus: 25 unsafe prompts tagged with taxonomy ids, 30 safe prompts non-overlapping with the lesson 83 benign set. Includes 15 unittest cases, quiz with 6 questions, skill output, ECE binning implementation, refusal phrase classifier.