Files
Rohit Ghumare e103f54af6 feat(phase-19/84): refusal evaluation with mock LLM policies
Two-sided refusal metrics: under-refusal (answered unsafe), over-refusal
(refused safe), accuracy, ECE calibration, per-category under-refusal join
against the lesson 82 taxonomy. Three mock LLM policies (strict, leaky,
over-cautious) demonstrate the framework detects opposite failure modes.
Labeled corpus: 25 unsafe prompts tagged with taxonomy ids, 30 safe prompts
non-overlapping with the lesson 83 benign set.

Includes 15 unittest cases, quiz with 6 questions, skill output, ECE binning
implementation, refusal phrase classifier.
2026-05-26 19:32:21 +01:00
..