mirror of
https://github.com/DeusData/codebase-memory-mcp.git
synced 2026-10-02 04:54:47 +08:00
A mysqldump-style .sql file is a few CREATE TABLEs followed by megabytes of `INSERT ... VALUES (..),(..),...` rows. Tree-sitter built a full tree for every row (~70 bytes of tree per source byte; 0.26 s/MB with MySQL \' escapes, which the grammar does not know and error-recovers from) until the 5 s parse budget ran out, and the whole file, tables included, was skipped as "parse timeout". The reporter's database/sql/world.sql was dropped this way. The rows contribute nothing to the graph. A linear, quote- and comment-aware pre-scan (internal/cbm/sql_values.c) finds INSERT/REPLACE ... VALUES lists and excludes every tuple after the first that consists only of literals: strings (standard '' and MySQL \' escapes, X'..'/B'..'/N'.., _charset introducers), numbers, NULL, TRUE, FALSE, DEFAULT. A tuple holding a subquery, a function call, an identifier or a quoted identifier stays in the parse, and so does one whose string spans a raw newline. The excluded bytes go to ts_parser_set_included_ranges as the complement, so byte offsets and line/column positions of all kept text are unchanged. The ranges are cleared right after the parse because the parser is thread-local and reused. This is a syntax rule, not a work cap: the same file always gives the same ranges, and the parse budget is unchanged. Measured on generated mysqldump-shaped dumps, release build, same machine, index worker's own numbers (base = origin/main): dump extraction pipeline peak RSS graph 8 MB std '' 1310 -> 23 ms 1431 -> 139 ms 625 -> 37 MB identical 8 MB MySQL \' 2712 -> 24 ms 2827 -> 146 ms 474 -> 37 MB identical 32 MB MySQL \' 5145 -> 82 ms 5467 -> 396 ms 1132 -> 62 MB timeout -> indexed 59 MB MySQL \' 5130 -> 138 ms 5667 -> 674 ms 1157 -> 89 MB timeout -> indexed "identical" = every stored node and edge, Module end line included. For the 32/59 MB dumps the base run skipped the file as "parse timeout" (Module + 3 Tables + their DEFINES/USAGE missing). The fixed graph matches the 8 MB graph except for the Module's end line. With standard escapes the extraction and the stored graph are identical to the full parse. With MySQL \' escapes the full parse's error recovery swallows rows next to a misread escape, a subquery row among them, and the cut parse keeps those rows. So it loses nothing the full parse found outside the dropped rows and recovers what the full parse lost. The only things the full parse extracted INSIDE the dropped rows were junk usages (DEFAULT and _binary read as identifiers, fragments of misread strings) that name nothing. Tests (parse_coverage, pipeline; a CBM_TEST_SQL_FULL_PARSE_ON seam turns the exclusion off for one file so the same source can be compared both ways): - scanner cases: comments, strings, dollar quotes, ON DUPLICATE KEY UPDATE, VALUES outside INSERT, unclosed tuples, exact kept positions - extraction equals the full parse on a small dump (standard escapes); nothing lost outside dropped rows (MySQL escapes) - the subquery and function-call rows are still parsed - the tree does not grow when the row count doubles (a count, not a clock) - a 25 MB MySQL dump under the production budget is indexed, not "parse timeout" (asserts the outcome, not the time) - stored nodes and edges of a pipeline index equal the full parse's Refs #1735 Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
41 lines
2.0 KiB
C
41 lines
2.0 KiB
C
/* sql_values.h — keep a SQL data dump's literal rows out of the parse (#1735).
|
|
*
|
|
* A mysqldump-style file is almost entirely `INSERT ... VALUES (..),(..),...`
|
|
* rows of plain literals. Those rows contribute nothing to the graph (no
|
|
* definition, call or usage lives in a literal), yet tree-sitter builds a full
|
|
* tree for every one of them: ~70 bytes of tree per source byte and seconds of
|
|
* parse per ten megabytes, until the per-file budget gives up and the whole file
|
|
* — its CREATE TABLE statements included — is skipped as "parse timeout".
|
|
*
|
|
* cbm_sql_values_kept_ranges() finds, in one linear quote- and comment-aware
|
|
* scan, every value tuple AFTER THE FIRST of an INSERT/REPLACE ... VALUES list
|
|
* that consists only of literals (strings, numbers, NULL, TRUE, FALSE, DEFAULT),
|
|
* and returns the complement as tree-sitter included ranges. The parser then
|
|
* sees `INSERT INTO t VALUES (first row);` — a complete statement — while every
|
|
* byte offset and line/column of the kept text is unchanged, so every node the
|
|
* extractors read still points at the real source.
|
|
*
|
|
* A syntax rule, not a work cap: the same file always yields the same ranges,
|
|
* whatever its size. A tuple holding anything that could reference code — a
|
|
* subquery, a function call, an identifier, a quoted identifier — is kept. */
|
|
#ifndef CBM_SQL_VALUES_H
|
|
#define CBM_SQL_VALUES_H
|
|
|
|
#include "tree_sitter/api.h"
|
|
#include <stdbool.h>
|
|
#include <stdint.h>
|
|
|
|
typedef struct {
|
|
TSRange *items; /* ascending, non-overlapping; owned (cbm_sql_kept_ranges_free) */
|
|
uint32_t count;
|
|
} CBMSqlKeptRanges;
|
|
|
|
/* Compute the included ranges for `src`. Returns true and fills `out` only when
|
|
* at least one literal tuple was excluded; returns false (out zeroed) when the
|
|
* whole file must be parsed as-is or on allocation failure. */
|
|
bool cbm_sql_values_kept_ranges(const char *src, uint32_t len, CBMSqlKeptRanges *out);
|
|
|
|
void cbm_sql_kept_ranges_free(CBMSqlKeptRanges *ranges);
|
|
|
|
#endif /* CBM_SQL_VALUES_H */
|