mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Three CodeRabbit findings for lesson 14. 1. docs/en.md tokenize: previous 'text.lower().split() + isalnum' left trailing punctuation glued to tokens (e.g., 'IPC.', '2007.'), which broke keyword matching. Switch to re.findall(r'[a-z0-9]+', ...) so 'ipc' and '2007' fall out cleanly. 2. docs/en.md dense_search: .squeeze() collapses a single-doc corpus to a 0-D scalar, breaking the subsequent indexing. Use .flatten() so sims is always 1-D regardless of corpus size. 3. code/main.py BM25.__init__: earlier fix used max(n_docs, 1) to dodge divide-by-zero, which masked empty-corpus inputs. Replace with explicit ValueError so the failure is loud, not silent.