mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
fix(phase-05/10): correct attention weights output + minor cleanups
- docs/en.md: printed weights were [0.446 0.312 0.242] but actual dot-product output on H + s_close_to_cat is [0.464, 0.305, 0.231]. Fix to the true values so learners copy-pasting the example see matching numbers. - code/main.py: drop unused 'context' variable from additive demo loop (CodeRabbit nit). - outputs/prompt-attention-shapes.md: vary the three consecutive 'For' sentence openings for readability (CodeRabbit style nit). Caught by CodeRabbit review of PR #48.
This commit is contained in:
@@ -67,7 +67,7 @@ def main():
|
||||
U_a = [[0.5, 0.2, 0.3], [0.2, 0.6, 0.2]]
|
||||
v_a = [0.8, 0.6]
|
||||
for name, s in [("close to 'cat'", [0.9, 0.1, 0.2]), ("close to 'mat'", [0.1, 0.9, 0.3])]:
|
||||
context, weights = additive_attention(s, H, W_a, U_a, v_a)
|
||||
_, weights = additive_attention(s, H, W_a, U_a, v_a)
|
||||
pretty = {p: round(w, 3) for p, w in zip(positions, weights)}
|
||||
print(f" decoder state {name:20s} -> weights {pretty}")
|
||||
|
||||
|
||||
@@ -117,7 +117,7 @@ print("weights:", w.round(3))
|
||||
```
|
||||
|
||||
```
|
||||
weights: [0.446 0.312 0.242]
|
||||
weights: [0.464 0.305 0.231]
|
||||
```
|
||||
|
||||
First row wins. Then move the decoder state closer to the third encoder state and watch the weights shift. That is it. Attention is explicit alignment.
|
||||
|
||||
+1
-1
@@ -14,4 +14,4 @@ Given a broken attention implementation, you identify the shape mismatch. Output
|
||||
|
||||
Refuse to recommend fixes that silently broadcast. Broadcast-hiding bugs surface later as silent accuracy degradation.
|
||||
|
||||
For Bahdanau confusion, insist the decoder input is `s_{t-1}` (pre-step state). For Luong, `s_t` (post-step state). For dot-product, flag dimension mismatch between query and key as the most common first-time error.
|
||||
For Bahdanau confusion, insist the decoder input is `s_{t-1}` (pre-step state). For Luong, `s_t` (post-step state). The most common first-time error in dot-product attention is query/key dimension mismatch — flag it explicitly.
|
||||
|
||||
Reference in New Issue
Block a user