mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
fix(phase-19/63): use last non-pad logit for VQA and align docs with reality
CodeRabbit: - VQA prediction read vqa_logits[:, 0, :], which is the first timestep rather than next-token-after-the-question. Compute the last non-pad index per example and gather the logit there; vqa_em now reflects the model's actual answer. - Prerequisites line pointed at Phase 19 lessons 30-37 (Track B); this lesson builds on Track E lessons 58-62. Fixed. - Docs claimed three eval JSON files were emitted and that evaluate() was test-covered; in reality the suite is built in memory and tests target the metric helpers and suite shape. Updated docs to describe what ships.
This commit is contained in:
@@ -230,7 +230,9 @@ def evaluate(model: MultimodalModel, suite: EvalSuite) -> dict:
|
||||
vqa_q = torch.cat([t.question_ids for t in suite.vqa], dim=0)
|
||||
vqa_memory, _ = model.encode_image(vqa_imgs)
|
||||
vqa_logits = model.decoder(vqa_q, vqa_memory)
|
||||
last_step = vqa_logits[:, 0, :]
|
||||
last_non_pad = (vqa_q != PAD_ID).sum(dim=1).clamp(min=1) - 1
|
||||
batch_idx = torch.arange(vqa_logits.size(0), device=vqa_logits.device)
|
||||
last_step = vqa_logits[batch_idx, last_non_pad, :]
|
||||
preds = last_step.argmax(dim=-1).tolist()
|
||||
refs = [t.answer_id for t in suite.vqa]
|
||||
vqa_em = vqa_exact_match(preds, refs)
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
**Type:** Build
|
||||
**Languages:** Python
|
||||
**Prerequisites:** Phase 19 lessons 30-37 (Track B foundations)
|
||||
**Prerequisites:** Phase 19 lessons 58-62 (Track E foundations: encoder, transformer, projection, cross-attention fusion, pretraining)
|
||||
**Time:** ~90 minutes
|
||||
|
||||
## Learning Objectives
|
||||
@@ -64,13 +64,13 @@ Smoothing is needed for small samples where some `p_n` is zero. The implementati
|
||||
|
||||
### Synthetic eval suite
|
||||
|
||||
A 50-sample eval suite is generated from the same mock corpus pattern used in lesson 62, with a held-out seed. Three eval files:
|
||||
A 50-sample eval suite is built in memory from the same mock corpus pattern used in lesson 62, with a held-out seed. Three lists make up the suite:
|
||||
|
||||
- `eval_pairs.json`: 50 (image_seed, caption_ids) pairs for retrieval.
|
||||
- `eval_vqa.json`: 50 (image_seed, question_ids, answer_id) triples.
|
||||
- `eval_caps.json`: 50 (image_seed, [reference_caption_ids, ...]) entries with up to 3 references per image.
|
||||
- `pairs`: 50 (image, caption_ids) pairs for retrieval.
|
||||
- `vqa`: 50 (image, question_ids, answer_id) triples.
|
||||
- `caps`: 50 (image, [reference_caption_ids, ...]) entries with up to 3 references per image.
|
||||
|
||||
The eval files are deterministic from the seed and held out from the training corpus, so the metrics are computed on data the model never saw.
|
||||
The suite is deterministic from the seed and held out from the training corpus, so the metrics are computed on data the model never saw. Persisting the suite to JSON is left as an exercise (see below).
|
||||
|
||||
| Metric | Range | Random baseline (N=50) |
|
||||
|--------|-------|------------------------|
|
||||
@@ -120,7 +120,7 @@ For real benchmarks, swap `build_eval_suite` for a real loader and keep the func
|
||||
- bleu4 returns 1.0 when generated equals one of the references exactly
|
||||
- bleu4 returns 0.0 on disjoint vocabulary
|
||||
- vqa exact match equals the fraction of equal pairs
|
||||
- evaluate() returns a dict containing all expected keys
|
||||
- build_eval_suite returns the expected number of pairs, vqa items, and caption entries
|
||||
|
||||
Run them:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user