mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Two-sided refusal metrics: under-refusal (answered unsafe), over-refusal (refused safe), accuracy, ECE calibration, per-category under-refusal join against the lesson 82 taxonomy. Three mock LLM policies (strict, leaky, over-cautious) demonstrate the framework detects opposite failure modes. Labeled corpus: 25 unsafe prompts tagged with taxonomy ids, 30 safe prompts non-overlapping with the lesson 83 benign set. Includes 15 unittest cases, quiz with 6 questions, skill output, ECE binning implementation, refusal phrase classifier.
1.1 KiB
1.1 KiB
name, description, version, phase, lesson, tags
| name | description | version | phase | lesson | tags | |||
|---|---|---|---|---|---|---|---|---|
| skill-refusal-evaluation | Two-sided refusal metrics with calibration and per-category breakdown over a labeled prompt corpus | 1.0.0 | 19 | 84 |
|
Refusal Evaluation
A labeled corpus of safe and unsafe prompts goes through one or more model policies. Outputs are classified as refusals or answers. The framework returns:
- under-refusal: answered prompts labeled unsafe / total unsafe
- over-refusal: refused prompts labeled safe / total safe
- accuracy: (correct refusals + correct answers) / total
- ECE: expected calibration error binned by stated confidence
- per-category under-refusal: joined against the lesson 82 taxonomy
Plugging in a real model
The mock LLM is a callable (prompt: str) -> str. Replace it with an HTTP wrapper that returns the model output and embeds a confidence tag (or modify parse_confidence to read whatever your provider exposes). Everything else stays the same.
Artifact
outputs/refusal_eval_report.json contains the per-policy metrics. Lesson 87 reads this report to set thresholds.