Files
Rohit Ghumare fd53ad179e feat(phase-05/04): GloVe, FastText, and subword embeddings
Three approaches side by side. GloVe as co-occurrence matrix
factorization with weighted log-loss. FastText as sum of character
n-grams (with OOV-tolerant lookup demonstrated). BPE as iterative
pair-merging that produces the subword vocabulary every modern LLM
tokenizer uses.

The teaching thread: Word2Vec had no OOV story. FastText fixed it for
words. BPE solved it for the whole vocabulary at a different level and
became the transformer-era default.

Ship artifact: tokenizer-picker skill with vocab-size targets (32k
English, 64-100k multilingual) and the reproducibility pitfall
(tokenizer-model mismatch is the single most common silent production
bug).

~45 minutes. Prerequisites lesson 05/03.
2026-04-22 15:45:57 +01:00
..