mirror of
https://github.com/colbymchenry/codegraph.git
synced 2026-10-02 01:37:32 +08:00
* feat(extraction): index Python docstrings, not just preceding comments (#1905) getPrecedingDocstring only walks preceding comment siblings, so a Python docstring — a bare string literal first in the body — never reached the docstring column, never entered nodes_fts, and was never shown by `codegraph node`. The identical sentence written as a leading `#` comment was both. For a Python codebase that is most of the prose there is. (#1905) Adds an optional getBodyDocstring() to LanguageExtractor, in the same shape as getSignature(), and routes every docstring call site through one docstringFor() helper that consults both sources. A node carrying a comment AND a docstring keeps both, joined: they are two things the author wrote about the same symbol and the column is free text. The Python implementation reads the grammar`s string_content rather than slicing quotes off the raw text, so r/u/b prefixes and both triple-quote forms work without a regex per case; f-strings are skipped because an interpolated string is code, not prose. Dedent follows PEP 257 (first line exempt). Scoped to Python deliberately. Julia (a string sibling before the def) and Elixir (@doc) fit the same hook and are left as follow-ups. * fix(extraction): index Python body docstrings across backends (#1905) Python extraction only consulted preceding comments, leaving body docstrings absent from search and rendered prose. Extend the contributor's hook to module and definition docstrings, and mirror it in the native kernel. Handle comments, concatenated literals, and indentation consistently while rejecting bytes, f-strings, and tuples. Verify persistence, exact-name ranking, MCP/CLI rendering, and native/WASM parity. Co-authored-by: Max Hsu <maxmilian@gmail.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Max Hsu <maxmilian@gmail.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>