Files
Rohit Ghumare 5d62a7b90e feat(phase-05/16): text generation before transformers — n-gram LMs
Bigram language model with Laplace and Kneser-Ney smoothing from
scratch in pure Python. Perplexity evaluation. Sampling. Demo output
shows KN beating Laplace by 2.3x on toy data and produces the classic
locally-plausible globally-incoherent n-gram sampling output.

Covers the full smoothing progression (Laplace -> Good-Turing ->
interpolation -> backoff -> absolute discounting -> Kneser-Ney -> MKN)
with the Kneser-Ney continuation-probability insight explained via the
classic 'San Francisco' example.

Bridges to neural LMs by naming the ~10x perplexity gap between KN
4-gram (~140 on Brown) and transformer LMs (~20).

Ship artifact: lm-baseline prompt for using n-gram LMs as a baseline
before shipping a neural model. Refuses to compare perplexity across
different tokenizations.

~45 minutes. Prerequisites lesson 05/01 and phase 2/14.
2026-04-22 17:35:23 +01:00
..