mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Extractive QA (RoBERTa-SQuAD2) + retrieval-augmented (RAG) + generative (closed-book LLM) as three architectures with different failure modes. Worked examples with SQuAD-style EM and F1 metrics demonstrated in code, showing how paraphrase punishes exact match. Teaches the retrieval-reader split, citation-based evaluation, and refusal calibration as the three production-grade metrics beyond raw answer accuracy. Ship artifact: qa-architect skill that refuses closed-book LLM answers for regulatory questions and flags missing retrieval-recall baselines. ~75 minutes. Prerequisites lesson 05/11 and 05/10.