Files
Colby Mchenry 4efc6c70e2 fix(scale): kernel-scale hardening — OOM-safe pass skipping + watchdog-safe index recreate (#1323)
Two hazards found by running today's full stack against the Linux kernel
(70,129 files) in the cg1212 repro container:

1. The parallel-synthesis fallback retried a worker-failed pass on the MAIN
   thread. At multi-million-node scale a worker failure is usually a memory
   ceiling, so the retry would OOM the process and take the whole index with
   it. Above 1.5M nodes a failed pass is now skipped with a clear stderr
   message (its synthesized edges are absent; the index completes). Below
   that, the main-thread retry stays — small-scale worker crashes are
   transient and the retry keeps coverage.

2. endBulkEdgeLoad rebuilt all four edge indexes in one synchronous span —
   measured 79s at kernel scale, past the #850 liveness watchdog's 60s
   stall window. A daemon-triggered re-index would have been SIGKILLed right
   after doing the work. Now async with an event-loop yield between builds,
   keeping each stall to a single index (~20s at kernel scale).

Validation: full Linux kernel index to completion in the repro container —
2,048,674 nodes / 6,405,964 edges, EXIT 0, zero passes skipped, on a 2-CPU
VM (worst case: pool disabled, sequential resolution + synthesis) in ~27min.
Phase walls: parse 6.0m, resolution 19.5m (incl. synthesis 6.3m, recreate
79s), maintenance 74s. Suite green (2444).

Also adds docs/design/native-extraction-kernel.md — the spike-validated
design for the native extraction kernel (Rust parse+walk over dubbo's Java:
202ms rayon / 1.07s single-thread vs 4.7s for the current wasm pipeline).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 19:09:01 -05:00

4.2 KiB
Raw Permalink Blame History

Native extraction kernel — design + spike results

Status: spike validated 2026-07-16; project approved, not yet started. Owner context: the last structural lever for fresh-index wall clock after the 2026-07-16 arc (#1305, #1320, #1321, #1322) exhausted Node-side scheduling.

Why

Post-arc, the fresh-index profile on dubbo (4,402 Java files) is: parse-loop ~4.7s, resolution ~5.5s (persist-bound), synthesis ~0.9s, total ~11.1s vs codebase-memory-mcp v0.9.0 at 7.1s. Two levers were measured dead:

  • RAM-backed DB (parse-loop 6.9s on a ramdisk vs 4.64.8s on SSD, n=2 interleaved): fast-init synchronous=OFF already writes at page-cache speed. The parse phase is CPU-bound.
  • TreeCursor rewrite (earlier arc): web-tree-sitter's traversal is not the cost; the floor is per-node JS↔WASM marshaling — every node.kind, .childForFieldName, .text crosses the boundary.

The only remaining parse lever is doing the walk on the native side and crossing the boundary once per file instead of once per node.

Spike (2026-07-16)

Minimal Rust binary (tree-sitter 0.25 + tree-sitter-java, TreeCursor walk touching every node's kind/range + name-field text, emitting flat (kind_id, start, end, name_len) rows — the extraction access pattern). Dubbo's 4,048 .java files, 17MB, 3.59M AST nodes, Apple M3 Pro:

wall
Current pipeline parse-loop (7 wasm workers, incl. extraction + store dispatch) 4,700ms
Rust parse+walk, rayon 202ms
Rust parse+walk, single thread 1,067ms

One native thread beats the whole 7-worker wasm pool 4.4×; at equal parallelism the walk is ~14× faster. Even charging the kernel for the extraction logic it must still perform, parse-loop 4.7s → ~1.01.5s is realistic, putting dubbo ≈ 7.58s total (parity with cbm).

Architecture

  • Crate: codegraph-kernel, napi-rs, links tree-sitter's C library and vendored grammars natively. Input: (filePath, content, language). Output: flat typed buffers (nodes, edges, unresolved refs) — one boundary crossing per file.
  • Per-language logic: migrate extractors to tree-sitter query files (.scm) executed by a generic Rust emitter; bespoke TS logic that queries can't express (macro salvage, dialect sniffing, content-gated .h detection) stays as TS pre/post passes over the returned buffers.
  • Distribution: prebuilt .node per platform through the existing release-bundle pipeline (same per-platform packages as the Node runtime). The wasm path remains as the universal fallback — same crate compiled to wasm keeps one implementation.
  • Rollout: per-language, funnel languages first (TS/JS → Java → Python → Go). A language ships only when its equivalence gate passes.

Equivalence gate (per language)

Byte-identity against hand-written extractors is NOT expected (bespoke logic ports approximately). The gate is:

  1. Node/edge/ref counts within ±0.5% on 3 real repos (small/medium/large), with every diff category eyeballed.
  2. The retrieval invariants hold: explore-flow connects the language's canonical flows end-to-end (docs/design/dynamic-dispatch-coverage-playbook.md), agent A/B shows no regression per the standard methodology.
  3. Fresh-index wall improves on the language's repos; no regression on a control repo of a non-migrated language.

Non-goals

  • Porting resolution, synthesis, frameworks, MCP, or the installer — they are pool-parallel and not marshal-bound. The measured native advantage there is ~1.4× CPU, not worth the correctness moat (2,444 tests, byte-identical determinism, years of invariants).
  • A single static binary (distribution polish, orthogonal to speed).

Risks

  • ABI drift between vendored native grammars and the wasm fallback grammars (keep both built from the same grammar source revs; CI asserts).
  • .scm expressiveness ceilings — budget for a per-language "escape hatch" callback in the emitter before declaring a language blocked.
  • napi-rs threading vs the parse-pool: the kernel replaces the wasm workers' parse+extract; the pool orchestration (file-order commit, retry, recycle) stays in TS and drives the kernel synchronously per file.