mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Critical (1): - Mini GPT training loop now actually trains (backward pass through FFN, LayerNorm, output projection -- loss drops 5.57 to 5.06) Major (11): - BCE gradient scaled by 1/n to match mean loss - AdamW weight decay applied before Adam step (decoupled correctly) - Dropout backward now applies mask and 1/(1-p) scaling - BatchNorm forward handles single-sample batches via running stats - LR schedule trajectory samples actual final step - Data parallel distributes remainder samples instead of dropping - Tensor parallel validates divisibility before sharding - Memory calculator uses actual optimizer param, not hardcoded Adam - Auto GPU heuristic verifies plan fits before returning - 70B GPU count fixed: 2 A100s for weights, 3+ with optimizer - Memory table: gradients column corrected to 2 bytes (FP16) Minor (6): - Length assertions on all paired-input functions - w2 initialized with its own fan_in/fan_out - warmup_cosine and one_cycle guarded against division by zero - embed_dim % num_heads validated at construction - ROADMAP footer updated to 68 complete - README badge updated to 68 complete