fix(phase-05/10): correct attention weights output + minor cleanups

- docs/en.md: printed weights were [0.446 0.312 0.242] but actual
  dot-product output on H + s_close_to_cat is [0.464, 0.305, 0.231].
  Fix to the true values so learners copy-pasting the example see
  matching numbers.
- code/main.py: drop unused 'context' variable from additive demo
  loop (CodeRabbit nit).
- outputs/prompt-attention-shapes.md: vary the three consecutive 'For'
  sentence openings for readability (CodeRabbit style nit).

Caught by CodeRabbit review of PR #48.
This commit is contained in:
Rohit Ghumare
2026-04-22 18:01:12 +01:00
parent 1807fbd0a4
commit 17645a2991
3 changed files with 3 additions and 3 deletions
@@ -67,7 +67,7 @@ def main():
U_a = [[0.5, 0.2, 0.3], [0.2, 0.6, 0.2]] U_a = [[0.5, 0.2, 0.3], [0.2, 0.6, 0.2]]
v_a = [0.8, 0.6] v_a = [0.8, 0.6]
for name, s in [("close to 'cat'", [0.9, 0.1, 0.2]), ("close to 'mat'", [0.1, 0.9, 0.3])]: for name, s in [("close to 'cat'", [0.9, 0.1, 0.2]), ("close to 'mat'", [0.1, 0.9, 0.3])]:
context, weights = additive_attention(s, H, W_a, U_a, v_a) _, weights = additive_attention(s, H, W_a, U_a, v_a)
pretty = {p: round(w, 3) for p, w in zip(positions, weights)} pretty = {p: round(w, 3) for p, w in zip(positions, weights)}
print(f" decoder state {name:20s} -> weights {pretty}") print(f" decoder state {name:20s} -> weights {pretty}")
@@ -117,7 +117,7 @@ print("weights:", w.round(3))
``` ```
``` ```
weights: [0.446 0.312 0.242] weights: [0.464 0.305 0.231]
``` ```
First row wins. Then move the decoder state closer to the third encoder state and watch the weights shift. That is it. Attention is explicit alignment. First row wins. Then move the decoder state closer to the third encoder state and watch the weights shift. That is it. Attention is explicit alignment.
@@ -14,4 +14,4 @@ Given a broken attention implementation, you identify the shape mismatch. Output
Refuse to recommend fixes that silently broadcast. Broadcast-hiding bugs surface later as silent accuracy degradation. Refuse to recommend fixes that silently broadcast. Broadcast-hiding bugs surface later as silent accuracy degradation.
For Bahdanau confusion, insist the decoder input is `s_{t-1}` (pre-step state). For Luong, `s_t` (post-step state). For dot-product, flag dimension mismatch between query and key as the most common first-time error. For Bahdanau confusion, insist the decoder input is `s_{t-1}` (pre-step state). For Luong, `s_t` (post-step state). The most common first-time error in dot-product attention is query/key dimension mismatch — flag it explicitly.