Files
ai-engineering-from-scratch/phases/05-nlp-foundations-to-advanced/14-information-retrieval-search
Rohit Ghumare 06a3429237 fix(phase-05/14): regex tokenize + flatten over squeeze + strict empty-corpus check
Three CodeRabbit findings for lesson 14.

1. docs/en.md tokenize: previous 'text.lower().split() + isalnum' left
   trailing punctuation glued to tokens (e.g., 'IPC.', '2007.'), which
   broke keyword matching. Switch to re.findall(r'[a-z0-9]+', ...) so
   'ipc' and '2007' fall out cleanly.
2. docs/en.md dense_search: .squeeze() collapses a single-doc corpus
   to a 0-D scalar, breaking the subsequent indexing. Use .flatten()
   so sims is always 1-D regardless of corpus size.
3. code/main.py BM25.__init__: earlier fix used max(n_docs, 1) to
   dodge divide-by-zero, which masked empty-corpus inputs. Replace
   with explicit ValueError so the failure is loud, not silent.
2026-04-22 19:37:59 +01:00
..