Files
Martin Vogel c0a06b3d0a fix(sql): keep literal INSERT rows out of the SQL parse (#1735)
A mysqldump-style .sql file is a few CREATE TABLEs followed by megabytes of
`INSERT ... VALUES (..),(..),...` rows. Tree-sitter built a full tree for
every row (~70 bytes of tree per source byte; 0.26 s/MB with MySQL \'
escapes, which the grammar does not know and error-recovers from) until the
5 s parse budget ran out, and the whole file, tables included, was skipped
as "parse timeout". The reporter's database/sql/world.sql was dropped this
way. The rows contribute nothing to the graph.

A linear, quote- and comment-aware pre-scan (internal/cbm/sql_values.c)
finds INSERT/REPLACE ... VALUES lists and excludes every tuple after the
first that consists only of literals: strings (standard '' and MySQL \'
escapes, X'..'/B'..'/N'.., _charset introducers), numbers, NULL, TRUE,
FALSE, DEFAULT. A tuple holding a subquery, a function call, an identifier
or a quoted identifier stays in the parse, and so does one whose string
spans a raw newline. The excluded bytes go to ts_parser_set_included_ranges
as the complement, so byte offsets and line/column positions of all kept
text are unchanged. The ranges are cleared right after the parse because
the parser is thread-local and reused. This is a syntax rule, not a work
cap: the same file always gives the same ranges, and the parse budget is
unchanged.

Measured on generated mysqldump-shaped dumps, release build, same machine,
index worker's own numbers (base = origin/main):

  dump            extraction   pipeline    peak RSS   graph
  8 MB std ''     1310 ->  23 ms  1431 -> 139 ms   625 -> 37 MB   identical
  8 MB MySQL \'   2712 ->  24 ms  2827 -> 146 ms   474 -> 37 MB   identical
  32 MB MySQL \'  5145 ->  82 ms  5467 -> 396 ms  1132 -> 62 MB   timeout -> indexed
  59 MB MySQL \'  5130 -> 138 ms  5667 -> 674 ms  1157 -> 89 MB   timeout -> indexed

"identical" = every stored node and edge, Module end line included. For the
32/59 MB dumps the base run skipped the file as "parse timeout" (Module +
3 Tables + their DEFINES/USAGE missing). The fixed graph matches the 8 MB
graph except for the Module's end line.

With standard escapes the extraction and the stored graph are identical
to the full parse. With MySQL \' escapes the full parse's error recovery
swallows rows next to a misread escape, a subquery row among them, and the
cut parse keeps those rows. So it loses nothing the full parse found
outside the dropped rows and recovers what the full parse lost. The only
things the full parse extracted INSIDE the dropped rows were junk usages
(DEFAULT and _binary read as identifiers, fragments of misread strings)
that name nothing.

Tests (parse_coverage, pipeline; a CBM_TEST_SQL_FULL_PARSE_ON seam turns the
exclusion off for one file so the same source can be compared both ways):
- scanner cases: comments, strings, dollar quotes, ON DUPLICATE KEY
  UPDATE, VALUES outside INSERT, unclosed tuples, exact kept positions
- extraction equals the full parse on a small dump (standard escapes);
  nothing lost outside dropped rows (MySQL escapes)
- the subquery and function-call rows are still parsed
- the tree does not grow when the row count doubles (a count, not a clock)
- a 25 MB MySQL dump under the production budget is indexed, not
  "parse timeout" (asserts the outcome, not the time)
- stored nodes and edges of a pipeline index equal the full parse's

Refs #1735

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
2026-09-25 19:15:43 +02:00

41 lines
2.0 KiB
C

/* sql_values.h — keep a SQL data dump's literal rows out of the parse (#1735).
*
* A mysqldump-style file is almost entirely `INSERT ... VALUES (..),(..),...`
* rows of plain literals. Those rows contribute nothing to the graph (no
* definition, call or usage lives in a literal), yet tree-sitter builds a full
* tree for every one of them: ~70 bytes of tree per source byte and seconds of
* parse per ten megabytes, until the per-file budget gives up and the whole file
* — its CREATE TABLE statements included — is skipped as "parse timeout".
*
* cbm_sql_values_kept_ranges() finds, in one linear quote- and comment-aware
* scan, every value tuple AFTER THE FIRST of an INSERT/REPLACE ... VALUES list
* that consists only of literals (strings, numbers, NULL, TRUE, FALSE, DEFAULT),
* and returns the complement as tree-sitter included ranges. The parser then
* sees `INSERT INTO t VALUES (first row);` — a complete statement — while every
* byte offset and line/column of the kept text is unchanged, so every node the
* extractors read still points at the real source.
*
* A syntax rule, not a work cap: the same file always yields the same ranges,
* whatever its size. A tuple holding anything that could reference code — a
* subquery, a function call, an identifier, a quoted identifier — is kept. */
#ifndef CBM_SQL_VALUES_H
#define CBM_SQL_VALUES_H
#include "tree_sitter/api.h"
#include <stdbool.h>
#include <stdint.h>
typedef struct {
TSRange *items; /* ascending, non-overlapping; owned (cbm_sql_kept_ranges_free) */
uint32_t count;
} CBMSqlKeptRanges;
/* Compute the included ranges for `src`. Returns true and fills `out` only when
* at least one literal tuple was excluded; returns false (out zeroed) when the
* whole file must be parsed as-is or on allocation failure. */
bool cbm_sql_values_kept_ranges(const char *src, uint32_t len, CBMSqlKeptRanges *out);
void cbm_sql_kept_ranges_free(CBMSqlKeptRanges *ranges);
#endif /* CBM_SQL_VALUES_H */