diff --git a/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/code/main.py b/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/code/main.py index 85b87f670..0922c5086 100644 --- a/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/code/main.py +++ b/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/code/main.py @@ -67,7 +67,7 @@ def main(): U_a = [[0.5, 0.2, 0.3], [0.2, 0.6, 0.2]] v_a = [0.8, 0.6] for name, s in [("close to 'cat'", [0.9, 0.1, 0.2]), ("close to 'mat'", [0.1, 0.9, 0.3])]: - context, weights = additive_attention(s, H, W_a, U_a, v_a) + _, weights = additive_attention(s, H, W_a, U_a, v_a) pretty = {p: round(w, 3) for p, w in zip(positions, weights)} print(f" decoder state {name:20s} -> weights {pretty}") diff --git a/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/docs/en.md b/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/docs/en.md index e4cd6bb5c..bbdfc785d 100644 --- a/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/docs/en.md +++ b/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/docs/en.md @@ -117,7 +117,7 @@ print("weights:", w.round(3)) ``` ``` -weights: [0.446 0.312 0.242] +weights: [0.464 0.305 0.231] ``` First row wins. Then move the decoder state closer to the third encoder state and watch the weights shift. That is it. Attention is explicit alignment. diff --git a/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/outputs/prompt-attention-shapes.md b/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/outputs/prompt-attention-shapes.md index 252aefb5d..a1fd44e8b 100644 --- a/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/outputs/prompt-attention-shapes.md +++ b/phases/05-nlp-foundations-to-advanced/10-attention-mechanism/outputs/prompt-attention-shapes.md @@ -14,4 +14,4 @@ Given a broken attention implementation, you identify the shape mismatch. Output Refuse to recommend fixes that silently broadcast. Broadcast-hiding bugs surface later as silent accuracy degradation. -For Bahdanau confusion, insist the decoder input is `s_{t-1}` (pre-step state). For Luong, `s_t` (post-step state). For dot-product, flag dimension mismatch between query and key as the most common first-time error. +For Bahdanau confusion, insist the decoder input is `s_{t-1}` (pre-step state). For Luong, `s_t` (post-step state). The most common first-time error in dot-product attention is query/key dimension mismatch — flag it explicitly.