mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Hand-rolled MHA with split_heads/combine_heads operating on row-major Vec<f32> tiles, plus a GQA variant that shares KV heads across Q groups. Prints per-head weight matrices, MHA-vs-GQA KV cache ratio, and a 5K forward microbench. Stdlib only.