Files
Rohit Ghumare 564eba02c8 Fix all 18 CodeRabbit findings across 8 lessons
Critical (1):
- Mini GPT training loop now actually trains (backward pass through
  FFN, LayerNorm, output projection -- loss drops 5.57 to 5.06)

Major (11):
- BCE gradient scaled by 1/n to match mean loss
- AdamW weight decay applied before Adam step (decoupled correctly)
- Dropout backward now applies mask and 1/(1-p) scaling
- BatchNorm forward handles single-sample batches via running stats
- LR schedule trajectory samples actual final step
- Data parallel distributes remainder samples instead of dropping
- Tensor parallel validates divisibility before sharding
- Memory calculator uses actual optimizer param, not hardcoded Adam
- Auto GPU heuristic verifies plan fits before returning
- 70B GPU count fixed: 2 A100s for weights, 3+ with optimizer
- Memory table: gradients column corrected to 2 bytes (FP16)

Minor (6):
- Length assertions on all paired-input functions
- w2 initialized with its own fan_in/fan_out
- warmup_cosine and one_cycle guarded against division by zero
- embed_dim % num_heads validated at construction
- ROADMAP footer updated to 68 complete
- README badge updated to 68 complete
2026-03-30 18:20:52 +01:00
..