Two hazards found by running today's full stack against the Linux kernel (70,129 files) in the cg1212 repro container: 1. The parallel-synthesis fallback retried a worker-failed pass on the MAIN thread. At multi-million-node scale a worker failure is usually a memory ceiling, so the retry would OOM the process and take the whole index with it. Above 1.5M nodes a failed pass is now skipped with a clear stderr message (its synthesized edges are absent; the index completes). Below that, the main-thread retry stays — small-scale worker crashes are transient and the retry keeps coverage. 2. endBulkEdgeLoad rebuilt all four edge indexes in one synchronous span — measured 79s at kernel scale, past the #850 liveness watchdog's 60s stall window. A daemon-triggered re-index would have been SIGKILLed right after doing the work. Now async with an event-loop yield between builds, keeping each stall to a single index (~20s at kernel scale). Validation: full Linux kernel index to completion in the repro container — 2,048,674 nodes / 6,405,964 edges, EXIT 0, zero passes skipped, on a 2-CPU VM (worst case: pool disabled, sequential resolution + synthesis) in ~27min. Phase walls: parse 6.0m, resolution 19.5m (incl. synthesis 6.3m, recreate 79s), maintenance 74s. Suite green (2444). Also adds docs/design/native-extraction-kernel.md — the spike-validated design for the native extraction kernel (Rust parse+walk over dubbo's Java: 202ms rayon / 1.07s single-thread vs 4.7s for the current wasm pipeline). Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
4.2 KiB
Native extraction kernel — design + spike results
Status: spike validated 2026-07-16; project approved, not yet started. Owner context: the last structural lever for fresh-index wall clock after the 2026-07-16 arc (#1305, #1320, #1321, #1322) exhausted Node-side scheduling.
Why
Post-arc, the fresh-index profile on dubbo (4,402 Java files) is: parse-loop ~4.7s, resolution ~5.5s (persist-bound), synthesis ~0.9s, total ~11.1s vs codebase-memory-mcp v0.9.0 at 7.1s. Two levers were measured dead:
- RAM-backed DB (parse-loop 6.9s on a ramdisk vs 4.6–4.8s on SSD, n=2
interleaved): fast-init
synchronous=OFFalready writes at page-cache speed. The parse phase is CPU-bound. - TreeCursor rewrite (earlier arc): web-tree-sitter's traversal is not the
cost; the floor is per-node JS↔WASM marshaling — every
node.kind,.childForFieldName,.textcrosses the boundary.
The only remaining parse lever is doing the walk on the native side and crossing the boundary once per file instead of once per node.
Spike (2026-07-16)
Minimal Rust binary (tree-sitter 0.25 + tree-sitter-java, TreeCursor walk
touching every node's kind/range + name-field text, emitting flat
(kind_id, start, end, name_len) rows — the extraction access pattern).
Dubbo's 4,048 .java files, 17MB, 3.59M AST nodes, Apple M3 Pro:
| wall | |
|---|---|
| Current pipeline parse-loop (7 wasm workers, incl. extraction + store dispatch) | 4,700ms |
| Rust parse+walk, rayon | 202ms |
| Rust parse+walk, single thread | 1,067ms |
One native thread beats the whole 7-worker wasm pool 4.4×; at equal parallelism the walk is ~14× faster. Even charging the kernel for the extraction logic it must still perform, parse-loop 4.7s → ~1.0–1.5s is realistic, putting dubbo ≈ 7.5–8s total (parity with cbm).
Architecture
- Crate:
codegraph-kernel, napi-rs, links tree-sitter's C library and vendored grammars natively. Input:(filePath, content, language). Output: flat typed buffers (nodes, edges, unresolved refs) — one boundary crossing per file. - Per-language logic: migrate extractors to tree-sitter query files
(
.scm) executed by a generic Rust emitter; bespoke TS logic that queries can't express (macro salvage, dialect sniffing, content-gated.hdetection) stays as TS pre/post passes over the returned buffers. - Distribution: prebuilt
.nodeper platform through the existing release-bundle pipeline (same per-platform packages as the Node runtime). The wasm path remains as the universal fallback — same crate compiled to wasm keeps one implementation. - Rollout: per-language, funnel languages first (TS/JS → Java → Python → Go). A language ships only when its equivalence gate passes.
Equivalence gate (per language)
Byte-identity against hand-written extractors is NOT expected (bespoke logic ports approximately). The gate is:
- Node/edge/ref counts within ±0.5% on 3 real repos (small/medium/large), with every diff category eyeballed.
- The retrieval invariants hold: explore-flow connects the language's
canonical flows end-to-end (
docs/design/dynamic-dispatch-coverage-playbook.md), agent A/B shows no regression per the standard methodology. - Fresh-index wall improves on the language's repos; no regression on a control repo of a non-migrated language.
Non-goals
- Porting resolution, synthesis, frameworks, MCP, or the installer — they are pool-parallel and not marshal-bound. The measured native advantage there is ~1.4× CPU, not worth the correctness moat (2,444 tests, byte-identical determinism, years of invariants).
- A single static binary (distribution polish, orthogonal to speed).
Risks
- ABI drift between vendored native grammars and the wasm fallback grammars (keep both built from the same grammar source revs; CI asserts).
.scmexpressiveness ceilings — budget for a per-language "escape hatch" callback in the emitter before declaring a language blocked.- napi-rs threading vs the parse-pool: the kernel replaces the wasm workers' parse+extract; the pool orchestration (file-order commit, retry, recycle) stays in TS and drives the kernel synchronously per file.