mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Three approaches side by side. GloVe as co-occurrence matrix factorization with weighted log-loss. FastText as sum of character n-grams (with OOV-tolerant lookup demonstrated). BPE as iterative pair-merging that produces the subword vocabulary every modern LLM tokenizer uses. The teaching thread: Word2Vec had no OOV story. FastText fixed it for words. BPE solved it for the whole vocabulary at a different level and became the transformer-era default. Ship artifact: tokenizer-picker skill with vocab-size targets (32k English, 64-100k multilingual) and the reproducibility pitfall (tokenizer-model mismatch is the single most common silent production bug). ~45 minutes. Prerequisites lesson 05/03.