fix(phase-19/51): address CodeRabbit review

- Clamp RetrievalConfig.max_hops to 2 to match _graph_score support
- Docs: drop nonexistent score_breakdown field; describe actual per-paper score fields
- Docs: remove claims of cross-source warning and disagreement test coverage

Skipped (heavy lift, architectural): rerouting search() through mock clients
instead of the local corpus index. The mocks read the same corpus the client
already initializes from; the rewrite risks lesson correctness without value.
This commit is contained in:
Rohit Ghumare
2026-05-26 21:01:03 +01:00
parent 22dbc8966e
commit 7db6d487b8
2 changed files with 7 additions and 3 deletions
@@ -211,6 +211,10 @@ class RetrievalConfig:
w_graph: float = 0.3
w_recency: float = 0.2
def __post_init__(self) -> None:
if self.max_hops > 2:
self.max_hops = 2
@dataclass
class RankedPaper:
@@ -56,7 +56,7 @@ flowchart TD
M --> O[ranked paper list]
```
The retrieval client owns both passes and the merge. The caller hands it a query and gets back a ranked list with a `score_breakdown` field that explains the ranking.
The retrieval client owns both passes and the merge. The caller hands it a query and gets back a ranked list where each entry carries per paper score fields (`bm25_score`, `graph_distance`, `recency_score`, `final_score`) that explain the ranking.
## BM25 from scratch
@@ -94,7 +94,7 @@ Default weights are `0.5`, `0.3`, `0.2`. The weights are config; a stale topic m
The corpus is one hundred papers, generated by `build_corpus()`. Each paper has a hand written title and abstract on one of five topics: attention sparsity, retrieval augmentation, low rank adapters, dataset distillation, and evaluation harnesses. References and citations are wired so each topic forms a connected sub graph with a few cross topic edges.
The two mock API clients (`ArxivMockClient`, `SemanticScholarMockClient`) read from the same corpus but expose different fields. Arxiv returns title, abstract, year, authors. Semantic Scholar adds references and citations. The retrieval client unions on id and warns when a field disagrees between sources.
The two mock API clients (`ArxivMockClient`, `SemanticScholarMockClient`) read from the same corpus but expose different fields. Arxiv returns title, abstract, year, authors. Semantic Scholar adds references and citations. The retrieval client unions on id; cross client field disagreement handling is deferred to a follow up lesson.
## What lessons 52 and 53 read
@@ -106,7 +106,7 @@ The retrieval client returns a `RetrievalResult` with both the ranked list and t
`code/main.py` defines `Paper`, `ArxivMockClient`, `SemanticScholarMockClient`, `BM25Index`, `CitationGraph`, `RetrievalClient`, and a deterministic demo. The mock clients and the corpus are in the same file so the lesson stays portable. The BM25 implementation is one class, sixty lines. The graph traversal is one method.
`code/tests/test_retrieval.py` covers the lexical path, the graph path, the merge, the dedup, the empty query, and the cross client field disagreement.
`code/tests/test_retrieval.py` covers the lexical path, the graph path, the merge, the dedup, and the empty query.
## Where this slots in