feature/CG-3
806 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
52b194a6be |
merge main into CG-3: keep the envelope view alongside occupancy
CG-3 branched from main before CG-1 landed and rewrote parse-run.mjs wholesale into an exported parseSession(), which dropped CG-1's --envelope/--answer reporting entirely. That view is the instrument the CG-1/CG-22 allocation gate measures bar 2 with, and it is in that benchmark's documented reproduce steps, so it cannot be lost to the merge. Resolution takes CG-3's rewrite as the structure and ports the envelope feature into it: parseSession now collects codegraph_explore response text in call order, formatEnvelope renders the per-file share, and the CLI parses --envelope/--answer ahead of the positional filter so a glob is never mistaken for a log path. The glob sentinel stays written as a \u0000 escape, never a literal NUL byte -- a raw one makes git treat the whole script as binary, exactly as the comment there warns. Verified: --selftest 18/18, and a synthetic explore transcript reports the expected per-file shares and answer-set total. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d333017a8 |
test(agent-eval): count blocked CLI attempts apart from real contamination (CG-7)
The hook denies the invocation, so a denied attempt puts no codegraph output in the window and must not disqualify the run -- only a call that actually returned content does. Attempts are still reported, since an agent hunting for the CLI is worth seeing. |
||
|
|
e35d4861e0 |
test(agent-eval): block the codegraph CLI outright — hiding it from PATH was not enough (CG-7)
An agent denied `codegraph` on PATH ran `find / -maxdepth 4 -iname "*codegraph*"`, found the binary, and invoked it by ABSOLUTE PATH — 12 times in one without-arm run. So block the invocation itself with a PreToolUse hook on Bash, written into the run's output dir as an artifact alongside the MCP configs rather than as a repo file. The pattern matches command positions only, so looking is still allowed and only using is denied: `grep codegraph src/`, `ls .codegraph` and `which codegraph` pass through, while `codegraph explore`, `/abs/path/codegraph …`, `cd x && codegraph …` and `VAR=1 codegraph …` are refused. run-all.sh proves both directions at startup and refuses to run if either fails. parse-run.mjs's detector uses the same rule, so prevention and detection cannot drift — and it no longer false-positives on the corpus path, which contains the word codegraph. Verified end-to-end: the without-arm now probes with `ls .codegraph; which codegraph`, finds nothing usable, and falls back to Read/Bash. |
||
|
|
d3c01ce8ed |
test(agent-eval): stop the arms reaching codegraph through the shell (CG-7)
The without-arm had no MCP server but still had Bash, and the target repo carries the .codegraph/ index the with-arm needs. Agents found it: 14 of 15 without-arm runs in a 7-repo pass ran `codegraph explore` through Bash, one of them via `ls .codegraph && codegraph explore ...`. That arm was measuring codegraph-over-CLI, not codegraph-absent, so every without-arm number it produced was wrong. It bit the with-arm too -- output arriving through Bash is attributed to Bash, understating what codegraph itself occupies (1 of 15 runs). Both arms now run on a PATH where the CLI is hidden, so the MCP server is the only way to reach codegraph and stays the single variable. The binary shares a directory with tools the run needs -- claude itself sits next to it -- so the directory is substituted in place by one of symlinks to every entry except codegraph, preserving PATH order and precedence. The run aborts if claude or node did not survive the substitution. Prevention alone would fail silently the next time the CLI lands somewhere new, so parse-run.mjs flags any Bash command naming codegraph and parse-bench-readme drops contaminated without-arm runs from the aggregate (CG_INCLUDE_CONTAMINATED=1 keeps them). CG_ARMS re-runs one arm without redoing the other. |
||
|
|
257a7b7327 | test(agent-eval): RUN_FROM, to extend a pass without redoing finished runs (CG-7) | ||
|
|
af3ce390da | test(agent-eval): drop the duplicated fixed-overhead line (CG-7) | ||
|
|
77845c747d |
test(agent-eval): show every run's residual and tool mix, not just the median (CG-7)
A median over 2-3 runs hides swings big enough to flip a repo's sign. On vscode the without-arm ranged 40k to 67k and the with-arm 59k to 65k across two runs; the deciding variable is the tool mix, since a with-arm run that reads files ON TOP of calling explore pays for both. |
||
|
|
b93c8d2b6c |
test(agent-eval): self-test the occupancy math, and fix ratio calibration under shedding (CG-7)
parse-run.mjs --selftest runs the math over synthetic transcripts with known answers: attribution, message.id dedupe, compact_boundary, FIFO micro-compaction, and multi-turn stitching. It found a real bug. A gap where the window also SHED content has a delta far below what was added, which reads as absurdly dense text and dragged the whole run's ratio with it -- a shed gap in the fixture pushed 2.5 chars/tok to 4.4 and left the wrong result resident. Shedding can only push a gap's ratio up, so the calibration now takes the lower median as its centre, drops gaps well above it, and pools the rest. Runs that never shed are unaffected (gin and vscode re-measure identically). Also drafts docs/benchmarks/residual-context-occupancy.md -- method, error bar, and the limitations this metric does not settle. Baseline numbers to follow. |
||
|
|
4d5f8d371a |
test(agent-eval): report the occupancy metric's own error bar (CG-7)
On a gap that is >=95% one tool result, the measured context delta IS that result's token count, so the spread between it and the run-level ratio is the attribution error. Median over such gaps: +/-1-2% on real runs. |
||
|
|
9b4df2133b |
test(agent-eval): price codegraph's fixed context cost alongside its residual (CG-7)
The first request's prompt is system + tool schemas + the question, before any tool has answered, so differencing the arms' ctxBase prices what codegraph occupies whether or not the agent ever calls it. Measured on gin: +775 tokens, small because the tool is deferred -- only its name is in the initial listing. |
||
|
|
4080b7501e |
test(agent-eval): measure residual context occupancy, over multi-turn sessions (CG-7)
The A/B arms reported cost, tokens, time and tool counts for one headless question. They could not report what issue #1500 actually measured: how much of the context window a tool's responses still occupy once the question is answered, which every later turn is then charged for. parse-run.mjs now measures that. Tokens are measured, not estimated: for each assistant request, input + cache_read + cache_creation is the exact token count of its whole prompt, so consecutive requests differ by exactly what was appended between them. That delta is priced against the characters in the gap, calibrated on gaps that are >=80% tool result. Explore output lands near 2.3 chars/token, so the usual bytes/4 estimate would have under-counted it by ~40%. Content also leaves the window, so residual is tracked apart from contributed: a compact_boundary clears the resident set, and a mid-run context drop is micro-compaction, which sheds the oldest tool results first and is applied FIFO. run-all.sh takes "Q1||Q2||Q3" and runs them as one resumed session, one segment file per turn; parse-run.mjs stitches the segments back together. bench-readme.sh now runs each README repo as a three-turn session (CG_TURNS=1 restores the single-question form). parse-bench-readme.mjs reports the arms' retrieval residual side by side -- codegraph's responses against the without-arm's Read/Grep/Bash -- in absolute tokens, share of context, and share of window, and says so explicitly when the rows it aggregated were single-turn. Two transcript traps are handled and documented at the call site: Claude Code emits one assistant event per content block, all carrying the same usage (summing per event double-counts every turn with both thinking and a tool_use), and the streamed output_tokens is a partial snapshot. Occupancy lives in parse-run.mjs and is imported by the aggregator rather than extracted to a module -- a new scripts/agent-eval/*.mjs scores into the self-query fixture's own corpus and moves its numbers. |
||
|
|
c65d56ceba |
docs: CG-22 — the epic's gate, re-run at CG-15's exact setup (#1500)
CG-21 fixed the unspent-reservation defect and re-ran the A/B itself. CG-22 is
the gate proper: CG-15's setup, unchanged, measured independently of the task
that wrote the fix. RUNS=3, both arms codegraph-on, sonnet/high,
CODEGRAPH_NO_PROMPT_HOOK=1 on both, baseline pinned to
|
||
|
|
abee46c5e4 |
docs: CG-21 A/B — the gate passes, all four bars (#1500)
Re-runs CG-15's agent A/B on the fixed build: same harness, same three prompts, same baseline ref, n=6 per arm on express and excalidraw. Read = 0 in all 15 new-arm runs. The express regression that routed the defect to CG-21 does not reproduce in 6 attempts, and the baseline now reads in 4 of 6 while the new arm reads in none (median 24.5s -> 21.5s), so the control beats the arm it previously lost to. client-go holds 92.7-96.2% answer share against a baseline run at 53.8%. Excalidraw's new arm is ~8s slower at the median and that is recorded as NOT attributable to the build rather than waved through: explore's own latency is 374ms vs 372ms on the same query and index, the deterministic responses differ by +2% with one byte-identical, and the unchanged main build's own median moved 34s -> 26.5s between the two sessions — the same magnitude as the gap. Bars were not re-baselined; they are CG-15's four, applied to a larger sample. The CG-15 section is kept intact and marked superseded, because its root-cause analysis is the record of why the fix looks like it does. |
||
|
|
fca7d87047 |
fix(explore): a funded whole-file buy must also fit the render ceiling (CG-21)
Found reviewing the CG-21 fix rather than by a failing test, and it is the same defect inverted. A whole render that overruns `renderCeiling` is skipped ENTIRELY (the branch refuses to slice a file mid-method), so a buy that is approved by the funding pool but refused by the ceiling trades a clustered section for NO section. Only reachable on the 24K tiers. The funding line is `reservedTotal + 0.15 * envelope` — ~27.2K when a medium repo saturates — while `renderCeiling` is `min(1.5 * envelope, 25000) - 600` = 24.4K, so funding can approve ~2.8K the ceiling then refuses. At 13K the line is ~14.4K against an 18.9K ceiling and the two cannot cross, which is why the small-tier fixtures cannot see it. Failing the test in `buysWhole` drops through to the cluster path, which is bounded by `headroom` and always renders something. The GRACE arm is left alone deliberately: a file within a sliver of its reservation that still does not fit is genuinely at the end of a full response, and that predates this epic. Verified inert on the three A/B repos, so the agent A/B measured the same behaviour: excalidraw byte-identical on 3 queries, client-go byte-identical on 3 queries, express reproducer unchanged at 15,984. Full suite green (2,867 passed; one unrelated fs.watch timing flake that passes 30/30 in isolation). |
||
|
|
51cd053d85 |
docs: CG-21 — spending the reservation (design record + CHANGELOG precision)
Records both levers, the funding-pool design (and the per-file version that dropped payslip_builder.go), the resolved memory-budget.ts exception, and the two hermetic fixtures with their mutation matrix. The CHANGELOG clause 'no longer trimmed while a smaller, weakly-related one is included whole' was imprecise after CG-21: the smaller file often IS still included whole now, when its share nearly covers it. Reworded to say what the fix actually guarantees. |
||
|
|
fa7fb8d127 |
fix(explore): spend the reservation instead of dropping it (CG-21, #1500)
A file whose proportional reservation lands below its own size stopped rendering whole, and the fallback cluster render could leave most of that reservation unspent — the bytes were neither delivered nor redistributed. Found by CG-15's agent A/B on the express control: `lib/utils.js`, the top-ranked file, was reserved 3,870 chars and spent 583. The whole-file grace bound (reservation + a sliver) sat just under the file's 5,293 bytes, so the whole render was declined and three matched symbols became a stub. The source envelope fell 13,849 -> 9,241 against an UNCHANGED budget, and the agent Read the file back four times in 1 run of 3. Two levers, per the task's candidate fixes: - WHOLE_FILE_BUY_FRACTION: a reservation that already covers 60% of a file buys the whole file. Funded from ONE shared overshoot pool sized at 15% of the envelope, spent in rank order. Per-file funding is the version that fails, and it fails the same way the bug does — the merit test is a ratio, so several files qualify at once and N independent overshoots push the last section past the render ceiling. Measured on the payroll fixture: three files bought whole and `payslip_builder.go` was dropped entirely. A dropped section is strictly worse than a clustered one. - Reservation carry-forward: what a file cannot spend goes to the next file down, bounded by MAX_SHARE. Tracked as two running totals rather than a `spent` variable threaded through the render loop's dozen exit paths, so no path can forget to account, and symmetric — a buy that overshoots suppresses slack until a later under-spend covers it. Express reproducer: `lib/utils.js` 583 -> 6,268 whole, envelope 9,241 -> 14,505 on the same 13,000 budget. The `memory-budget.ts` exception CG-14 documented is RESOLVED rather than re-justified: it ships whole again at 5,672 (27.3%) while `src/mcp/tools.ts` rises to 52.6% — so the answer file wins the envelope AND no previously-unclipped file is clipped, which is CG-12's own acceptance criterion finally holding. Two hermetic fixtures added, one per lever, because nothing in the suite had this shape — which is how it shipped. Both mutation-tested: removing the buy arm reddens 3, removing the carry-forward reddens 2, and removing the funding guard reddens 4 (including payroll's dropped `payslip_builder.go`). Their `fixture shape` blocks are load-bearing: the gates pass vacuously if a target ever drifts inside the grace bound, so the window is asserted directly. Full suite green (2,868 passed); both #1500 regression fixtures pass. |
||
|
|
8077b83eb3 |
docs: consolidate the #1500 CHANGELOG entries into user-facing shape (CG-16)
Three separate engineer-shaped entries (CG-5 generated detection, CG-10 scoring, CG-12 allocation) become two user-facing bullets under Fixes, in the shape #1500 actually reported: explore concentrates on the code that answers the question, and a generated CRUD/protobuf layer no longer crowds out the hand-written code beside it. Per the CHANGELOG rules: strips the benchmark counts and percentages (the client-go 2,001-file count, the quarter-to-four-fifths envelope shift) and the internal symbol names, keeps the re-index note, and credits the reporter. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7103c7f2f |
docs: CG-15 agent A/B of the #1500 allocation change — gate fails on the control
Three repos, both arms codegraph-on, sonnet/high, 3 runs per arm. PASS on the two medium repos: client-go (the reporter's Go shape, 2,001 of 2,454 files generated) and excalidraw hold Read 0 in every run of both arms, excalidraw goes 34s -> 24s at the median with one fewer explore call, and the generated clientsets/informers that took 10.5%% of a baseline envelope appear in no new run. FAIL on express, the small control, in 1 run of 3: 4 Reads and 52s against a baseline that read once. Not agent variance — replaying that run's query deterministically, lib/utils.js goes from 6,380 bytes whole to a 583-byte cluster stub and the envelope shrinks 13.8K -> 9.2K against an unchanged 13,000 budget. The diagnostic shows the allocator was right and the render loop was not: utils.js is the top-ranked file, was reserved 3,870 chars, and spent 583. The whole-file bound (allowance + grace = 4,450) lands just under the file's 5,293 bytes, so the whole-file render is declined and the unspent reservation is dropped rather than redistributed. Bar 1 is the hard gate, so per CG-15's acceptance rule the design goes back to CG-12 — the budget is not to be widened to compensate. Two candidate fixes are written up in the design doc. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
edce18f586 |
test(agent-eval): restore the engine on INT/TERM too (CG-15)
Killing ab-new-vs-baseline.sh mid-baseline-arm left the engine checked out at the baseline ref with the post-baseline files deleted, so every later build in the working tree was silently the OLD code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
48a2309b92 |
test(agent-eval): RUNS knob + explore envelope-share view for new-vs-baseline A/B (CG-15)
ab-new-vs-baseline.sh now builds and indexes once per arm and runs the task RUNS times (default 1), so the >=2-runs-per-arm rule costs one build instead of N. Both arms run with CODEGRAPH_NO_PROMPT_HOOK=1 — the machine's ambient front-load hook resolves to whatever is in dist/, a second uncontrolled channel that confounds the tool-call counts — and point explore's CG-4 diagnostic at a per-arm sidecar. parse-run.mjs gains --envelope/--answer: the per-file share of the explore source envelope, parsed from the rendered markdown so it works on ANY build. The CG-4 sidecar only exists post-CG-4, so it cannot measure the baseline arm; this is the only view that measures both arms the same way. Folded into parse-run.mjs rather than added as a new script on purpose: a new file named after explore's budget scores into the self-query fixture's own corpus and moved its answer share 59.9%% -> 47.9%%. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d9206d2d0 |
test(explore): lock down proportional byte allocation (CG-14, #1500)
Coverage for the CG-12 allocator, built around "would this go red if the lever were removed" rather than line coverage — every way this regresses is silent, ending in an agent falling back to Read. Unit (`explore-proportional-allocation.test.ts`, 18 -> 38): calibration pins, envelope safety across every tier and 30 candidate shapes, the cliff boundary, spine weighting/trim survival, the diffuse control, and the degenerate inputs — identical scores, a lone file, a runaway top scorer, zero results, maxFiles 0, a non-finite score. End-to-end (`explore-allocation-e2e.test.ts`, new): CG-6's second regression fixture as a deterministic synthetic mirror — a large relevant file, a small helper that used to win by shipping whole, and an incidental `explore`/`BUDGET` collision — asserting per-file budget share, not file presence. Plus degenerate result sets and a survey-style diffuse control through the real render loop. The live self-query arm stays in probe-allocation.mjs, where drift is a number to re-baseline rather than a red suite. Reverting the render loop to the pre-CG-12 rules reproduces #1500 on the mirror exactly and takes 5 e2e + 2 payroll gates red: file score pre-CG-12 CG-12 src/mcp/allocator.ts 77.5 4,843 (39.7%) 9,335 (80.1%) src/util/budget-math.ts 36.0 6,079 (49.8%) 1,037 ( 8.9%) Two defects the invariants surfaced, both fixed in tools.ts: - rounded shares could sum past `pool`, so "reservations fit the envelope" was approximate rather than exact; both terms now floor - a non-finite score made every share Infinity/Infinity, handing the render loop a NaN allowance; `weightOf` now fails safe to 0 Also adds a hard-ceiling gate to the payroll fixture — at 19.3K against a 19.5K ceiling it is the only fixture that stresses the ~25K inline cap — and exports EXPLORE_ALLOCATION so invariant tests read the constants while one test pins the literals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5f7f5f59df |
feat(mcp): score-proportional byte allocation for explore, with a relative cliff (CG-12, #1500)
The explore envelope used to follow FILE SIZE, not relevance. Every admitted
file was capped at the same flat `maxCharsPerFile`, while the whole-file rule
handed anything under `maxCharsPerFile * 3` its entire contents — a 3x swing
decided by how big a file happened to be:
- self-query: `memory-budget.ts` (score 18) shipped whole and took 51.2% of
the response; `src/mcp/tools.ts` (score 41, 4x the graph mass, 3x the term
hits — it holds the allocator itself) was clipped at 3,800 and got 32.9%.
- #1500 Go fixture: two generated CRUD files shipped whole at ~4.5K each AND
consumed two of the tier's four file slots, so `BuildPayslip` — the
hand-written "calculate" half of the question — ranked #6 and never
rendered at all.
`allocateExploreBudget` now reserves each ranked file a share of the envelope
before anything renders, so the render loop spends a reservation instead of
racing for whatever the files above it left:
- weight = score x worth x (spine ? 2 : 1), where `worth` is `rankPenalty`
applied a SECOND time — ranking answers "is this file about the query",
allocation answers "will these bytes teach the agent anything", and
generated CRUD can legitimately rank while its bytes stay boilerplate;
- a relative cliff at 15% of the top weight (capped at SCORE_FLOOR_MAX, so a
god-file can't silence peers the score floor just admitted) gives a file
ZERO source — path, symbols and line numbers only — and crucially frees its
`maxFiles` slot for a file that earns its bytes;
- every admitted file gets MIN_CHARS, then the remainder splits by weight:
the floor keeps a diffuse survey question returning a spread, the remainder
concentrates a precise one;
- the flat per-file cap is retired as the primary guard, leaving a 70%-of-
envelope safety valve.
Two changes were needed to make the reservation bite: an oversize cluster now
shrinks by whole MEMBER symbol ranges (a single-cluster god-file previously
took ~40% more than allotted, and the file below it was dropped for lack of
room), and the arrival-order budget stops are gone — they cut files by the
order they were reached rather than by merit.
Measured: payroll-go answer group 25.6% -> 78.7%, generated 57.4% -> 0%, and
`func (s *Service) BuildPayslip` now delivered; self-query `tools.ts` 18.5% ->
60.6%, past the epic's >50% bar. Controls hold: cobra/gin diffuse survey
queries keep their file spread (3->3, 3->4), express's middleware query is
byte-identical, and gin's flow query moves its top file from the thin `ginS`
singleton wrapper to `routergroup.go`.
One documented exception to "no previously-unclipped file becomes clipped":
`memory-budget.ts` was unclipped-whole at 5,672 and now clusters within its
3.1K reservation. That is the epic's own diagnosis of the bug — it scored 18
against 58 and was taking the larger slice purely for being small.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a3898cdc70 |
feat(mcp): relevance scoring overhaul for explore — kill incidental name-collision matches (CG-10, #1500)
Explore's per-file relevance awarded +50/+10/+3/+1 by match class and admitted
anything scoring >= 3. Neither half held up: the tier said HOW a symbol reached
us, never whether the match was evidence, and an absolute floor admits noise on
any repo where the top file scores 50+. Three scripts/agent-eval/*.mjs harnesses
took 63% of this repo's own "how does explore allocate its output budget" answer
on nothing but an unused `const explore` and a `const BUDGET`.
Four levers:
- KIND WEIGHT (RELEVANCE_KIND_WEIGHT): callables and types 1.0, members ~0.5,
variable/constant/parameter 0.15-0.35. A weak-kind symbol with no usage edge
anywhere in the graph (`contains` excluded — nesting is not usage) drops to
0.08. Only weak kinds in the top two tiers pay for the DB probe; the subgraph's
own edges answer most cases free. No measurable latency change (210 vs 211
ms/call, n=12 interleaved).
- PERIPHERAL CAP: nodes >=2 hops from any match accumulate into a bucket capped
at 5. Uncapped they added a flat +1 each, so a file grew more relevant by being
bigger — parse-session.mjs reached 22 off one constant plus twelve unrelated
symbols.
- RANK PENALTY: generated files x0.3, low-value x0.5, applied to the score AND
the graph mass. Score alone would not have fixed #1500 — the generated CRUD
carries MORE graph mass than the hand-written use-case, and graph mass outranks
score in the comparator. Self-normalizing, never a hard exclusion.
- RELATIVE FLOOR: clamp(topScore * 0.2, 1, 10). Capped at one full-strength
direct match so concentration elsewhere can never exclude one (without it a
named-seed-heavy file pushed the floor to 21 and dropped a file the agent had
named by class name). Backfills to 3 candidates when it would leave fewer, and
drops the evidence requirement rather than return nothing at all.
excludeLowValueFiles was dead config — declared per tier, read nowhere; the
test/spec exclusion has been unconditional for a while. Removed. The real gap was
the detector: `isLowValue` anchored on a leading `/`, so a repo-ROOT `test/` dir
(express, cobra, most of npm and Go) never matched — express's routing question
spent 59% of its envelope on three test files. Anchored at `^` too, and the
filter now runs before the floor and judges "are there other candidates?" on the
whole gather.
Measured before/after on the same indexes (baseline
|
||
|
|
bd86ad2061 |
test(explore): #1500 regression fixtures for budget allocation (CG-6)
Two permanent fixtures pinning the failure mode from issue #1500 — explore spending its byte envelope on files that merely name-collide with the query. BOTH FAIL TODAY, by design: they document the bug and become the pass gate for CG-10 (scoring) + CG-12 (proportional allocation). __tests__/fixtures/payroll-go/ — a synthetic Go service mirroring the reporter's shape: generated FKIT CRUD beside a hand-written payroll use-case, entered from an HTTP route. Half the generated tree carries ORDINARY names detectable only by their `// Code generated ... DO NOT EDIT.` header (the #1500 case, and end-to-end cover for CG-5); `payrollpb/*.pb.go` covers the path-detectable channel. BuildPayslip, Upsert and Store each exist twice, generated and hand-written. cycle.go sits above the whole-file window so it clips; the generated files sit below it so they ship whole. Asking "how does payroll cycle create and calculate payslips?" — naming none of the answering symbols — the generated CRUD delivers 57.4% of the envelope against the hand-written layer's 25.6%, all of the latter domain types. cycle.go is allocated the single largest slice (30.6%) and delivers ZERO: the hard ceiling drops its whole section. runPayrollCycleAll, the hand-written BuildPayslip and the real Upsert never reach the agent. The second fixture is this repo, "how does explore allocate its output budget across files", where scripts/agent-eval/*.mjs take 71.8% against tools.ts's 18.5% despite scoring 4.6x lower. It reads the live index, so its assertions are relative rather than fixed percentages. - scripts/agent-eval/probe-allocation.mjs — per-file budget-share probe, driving the CG-4 diagnostic through a JSONL sidecar so it measures the shipping allocator. Fixture entries are hermetic (copy + re-index per run, verified byte-identical across runs); exits 1 while any assertion fails. - scripts/agent-eval/allocation-fixtures.json — both fixtures declared, with the 2026-08-03 baselines. - __tests__/explore-allocation-1500.test.ts — fixture-shape assertions green today; the allocation assertions held as `it.fails` so the suite stays green while the bug is open and goes RED the moment it is fixed. Also documented and deliberately left unfixed: runPayrollCycleAll's `s.store.Upsert` edge resolves to the GENERATED Store.Upsert, not the hand-written one — same-name method resolution across two packages picks the wrong receiver. It is upstream of the allocation bug, so it belongs with CG-10's scoring work. Refs #1500 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
16e17495f4 |
feat(extraction): content-based generated-file detection (CG-5, #1500)
`isGeneratedFile` was path-only, but Go's own convention is a CONTENT marker (`// Code generated by <tool>. DO NOT EDIT.`), not a filename one. A Go monorepo with generated CRUD in ordinarily-named files sitting beside hand-written use-cases was therefore invisible to every generated-file down-rank in the codebase — that is #1500. Measured on kubernetes/client-go (2,453 Go files): the canonical banner appears in 2,001 of them, the path check flags 0, the new content check flags exactly those 2,001 — no false positives, no misses. Design: decide at INDEX time (content is already in memory for parsing), persist on `files.generated`, read from the DB. Explore never reads file headers per request. - `hasGeneratedHeader(content)` recognizes the standard banners — Go's, protoc's, `@generated`, `<auto-generated>`, Thrift, OpenAPI Generator, FlatBuffers, bindgen, ANTLR. Precision-first and fenced three ways: an 8KB/60-line header window, a comment-line requirement (leader or open block comment), and markers tight enough that prose can't trip them. A generator's own source, holding the banner as a string constant in its body, is not flagged; neither is this module itself (pinned by test). - `isGeneratedFile(path)` is unchanged — cheap, sync, still the fallback. - Schema v9 adds `files.generated` + a PARTIAL index. DDL only, no backfill: the flag derives from content the migration cannot see, so rows stay 0 until a re-index and every reader unions the flag with the path check — an un-migrated index keeps pre-#1500 behavior rather than regressing. Re-index required; noted in the CHANGELOG. - `generatedPredicateFor(paths)` gives ranking a bounded probe + O(1) lookups. Bounded, not cached: no invalidation, so a ranking call can never serve a verdict the last sync already replaced. Wired into explore ranking, findSymbolMatches, findAllSymbols, search (MCP + CLI), the context formatter, and the dominant-file/route-file hygiene filters. Cost (acceptance bar was no measurable index-time regression): a single unanchored `/generat/i` test over the header rejects ~every hand-written file before any line splitting. 4.6 µs/file on client-go (worst case — 82% generated). End-to-end `codegraph init` on client-go, n=3 alternating arms: 5.73s median with detection vs 5.76s path-only baseline; the arms cross over between runs, so the difference is inside run-to-run noise. Scope note: generated status remains a stable TIEBREAK at equal score, exactly where it was. Making it a strong negative signal is CG-10, which this unblocks by making the signal correct and available. Two pre-existing tests hard-coded schema version 8; both now track CURRENT_SCHEMA_VERSION (or the migration table) so future migrations don't require editing them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b37f191f5a |
feat(mcp): per-file allocation diagnostic for explore (CG-4)
How codegraph_explore divides its byte envelope among files was unobservable — you could read a response and guess, but not say "this file took 16% and that one took 20%." Nothing else in the budget- allocation epic is measurable without that. CODEGRAPH_EXPLORE_DEBUG now emits one report per explore call (stderr table, stderr JSON, or a JSONL sidecar path). Per file: relevance score, graph mass, term hits, ranking flags, render mode, bytes allocated vs delivered, both shares, and whether it was clipped — plus why a ranked candidate never rendered. Totals cover envelope vs maxOutputChars vs the hard ceiling, the source/meta split, the selection funnel, and the score floor and relevance-gate thresholds applied. Allocated and delivered are reported separately on purpose: they diverge exactly when the 25K ceiling truncates, and conflating them is how a dropped trailing file goes unnoticed. Off by default and byte-identical when off — it ships in the product binary, and a diagnostic that perturbs the response by one byte would invalidate every A/B taken with it on. ExploreDiagnostics.start() returns null unless the env var is set, so every call site is a `diag?.` no-op. Baseline recorded in docs/design/explore-budget-allocation.md: on this repo, src/mcp/tools.ts gets 15.8% of the envelope while three weakly- relevant agent-eval scripts take 61% between them — despite tools.ts carrying 5.4x the score and 2.6x the graph mass of any of them. Small files ship whole; the large answer file is clipped at maxCharsPerFile. Rank ordering is correct and buys nothing. The loop also allocated 23,193 chars against an 18,000 budget. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
49c11fc2e0 |
Self-hosted telemetry on Cloudflare D1 + password-gated admin dashboard (CG-7) (#1497)
* feat(telemetry): D1 schema + migrations for raw events and daily rollups First step of replacing PostHog with self-hosted telemetry on Cloudflare D1. Creates the codegraph-telemetry database binding and the initial migration; no worker code paths change yet (the ingest write path and the nightly rollup cron land next). Schema is raw events plus daily rollups: `events` holds one row per sanitized event with the envelope broken out into columns and event-specific props as JSON; `daily_machines`, `daily_event_counts` and `daily_dim_counts` are the nightly rollups the dashboard reads; `machine_first_seen` and `machine_days` carry the retention cohorts and are never purged. One generic dimension table covers every bar and pie, so a new breakdown is a cron change rather than a migration. The migration is commented as an audit surface, like the rest of this worker — every column, and which dashboard chart each rollup table serves. Three judgment calls worth flagging, all documented in the file: - `events` gets `(day, event)` instead of the separate `(day)` and `(event, day)` indexes. D1 bills a row write per index touched, so a third index on the hot table costs ~97k writes/day, and `(day, event)` is a covering index for plain day-range scans anyway (verified with EXPLAIN QUERY PLAN). - `daily_event_counts` and `daily_dim_counts` carry a `machines` column, and `machine_days` a `prod` flag. The "users by ..." panels and the production-user count are distinct-machine numbers, not event counts, and they are unrecoverable once raw events are purged. - No CHECK constraint on `event`: the worker's allowlist is the source of truth and the write path is fail-silent, so a rejected INSERT would lose data quietly instead of erroring loudly. Volume note in the migration footer: ~30M row writes/month against the 50M included on Workers Paid. Storage is the tighter constraint — raw events grow ~74 MB/day, so retention should start at 90 days (~6.7 GB) rather than 180, which would exceed D1's 10 GB per-database cap. * feat(telemetry): admin dashboard worker — scaffold + shared-password auth New Cloudflare Worker at telemetry-dashboard/, sibling of telemetry-worker/ and bound read-only to the same D1 database. Serves a static frontend plus a JSON API behind a shared password, on stats.getcodegraph.com. Auth is the simplest thing that is actually safe for exactly two users: one password in a secret, compared in constant time over SHA-256 digests, and an HMAC-signed cookie (HttpOnly; Secure; SameSite=Lax; Path=/) with a one-year expiry so you sign in once per browser. The cookie is a signed assertion, not a lookup key — no session store. Its payload carries a fingerprint of the password it was minted against, so rotating ADMIN_PASSWORD signs everyone out. Login attempts are capped at 5/min per IP via a ratelimit binding. Everything is deny-by-default: assets.run_worker_first routes every request through the worker before the static-asset server sees it, so the dashboard HTML, its JS, its CSS and the chart library are all behind the session check. The login page is rendered inline by the worker rather than served from public/, which leaves no "is this file public?" judgement calls in the asset directory. Unauthenticated pages 302 to /login, unauthenticated /api/* gets 401. A missing secret fails closed rather than opening the dashboard. scripts/smoke-auth.sh is the regression net — 54 assertions against a throwaway `wrangler dev` covering the gate, cookie flags and persistence, forged/flipped/ truncated cookies, open-redirect refusal, brute-force capping, and password rotation invalidating live sessions. Refs CG-11. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(telemetry-dashboard): simplify the chart-library probe in the shell Refs CG-11. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(telemetry): nightly rollup cron + raw-event retention purge (CG-10) Adds a scheduled() handler to the ingest worker that recomputes daily_event_counts / daily_dim_counts / daily_machines for the just-completed UTC day plus a 2-day overlap (late-arriving offline buffers), then purges raw events past the retention window. Rollup writes are idempotent upserts, so a re-run never double-counts. Also adds an ADMIN_TOKEN-guarded POST /admin/rollup?day=YYYY-MM-DD for backfill/repair, and drops the PostHog forwarding path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(telemetry): dashboard charts — SQL API over D1 + the Chart.js views (CG-12, CG-13) Replaces the scaffold page with the dashboard proper: 19 panels covering every view of the PostHog dashboard this retires, driven by one filter row. src/api.ts is the read API CG-12 specified: /api/{meta,summary,timeseries, breakdown,activation,retention}, all range-scoped, all parameterized against a closed set of dims and metrics, all shaped labels[] + datasets[] so the frontend does no arithmetic. Rollups answer everything except the activation funnel, which needs raw events and says where they start. The frontend splits into a DOM-free panel registry (public/panels.js) and the page that mounts it (public/app.js), so the render check can drive the same registry the browser rendered from. Panels fail alone, refetch dims rather than flashing, and every chart carries a table twin. Two numbers are labelled rather than rounded off: range-wide "users" per dimension is machine-days (the rollups cannot give distinct machines, and per-day counts are taken as the largest single-event count so one machine's install + index + usage is not counted three times), and recent activation and retention cohorts are marked as still-converting instead of drawn as a cliff. Both colour scales were run through the data-viz validator against the panel surface, not picked by eye; the results are recorded in public/theme.js. Verification, all against the committed fixture (12 machines over 10 days, every expected number worked out by hand from the events, not recorded from a run): scripts/smoke-api.sh 98 assertions scripts/render-check.mjs 79 assertions — real Chromium over CDP, no new deps scripts/smoke-auth.sh 54 assertions (unchanged, still green) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(telemetry): cutover runbook + the end-to-end gate that de-risks it (CG-14) The account-level steps of the PostHog cutover are the maintainer's to run, so this lands the runbook they follow and the check that has to pass first. The runbook (telemetry-worker/README.md) walks the six steps in the order that keeps them reversible: Workers Paid → migrate → deploy → watch 24h → verify the first rollup and the dashboard → only then delete POSTHOG_KEY and cancel the subscription. Step 3 records the outgoing version id because `wrangler rollback` is the escape hatch for the whole verification window, and that window is precisely why the PostHog key is deleted last rather than first. The new gate (scripts/smoke-cutover.sh, `npm run smoke:cutover`) covers the one seam nothing else did. Both workers declare the same D1 database_id, so pointing them at a single --persist-to directory runs the real chain: a client batch → the ingest worker → D1 → the nightly rollup → the dashboard API reading the numbers back. Every other suite stops at one link — smoke-ingest at the events table, smoke-rollup at hand-checked SQL, smoke-api at a hand-written fixture that the cron never touched. That left the dimension names the rollup WRITES versus the ones the dashboard READS agreeing by convention across two branches, where a mismatch is silent: no error, no failed request, just a panel reading zero forever. 61 assertions, all 13 dimensions, and three deliberate traps — a ci machine that is active but not a production user, usage_rollup counts that must be summed rather than tallied, and an uninstall's `targets` that must not leak into the install-scoped breakdown. Writing it caught that the activation funnel's denominator is first-seen machines, not install events (deliberate — a reinstall must not re-enter the funnel), so the suite now pins that distinction rather than assuming it. Also rewords the last PostHog reference in dashboard code: a comment justifying the 14-day retention curve by pointing at a dashboard step 6 deletes. The reasoning now stands on its own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(telemetry): tell the truth about where events are stored (CG-15) The telemetry docs are a privacy contract, and they still described a managed analytics store that no longer receives anything. Replace that with what actually happens now — events land in our own D1 database on Cloudflare, the endpoint makes no outbound requests, raw events are purged after 90 days and only anonymous daily rollups outlive them. This strengthens the guarantee rather than restating it: there is no second party to share with. - TELEMETRY.md: new "Where it is stored" section; the never-collected IP bullet no longer leans on a vendor-side setting to hold. - docs/design/telemetry.md: ingest section rewritten around D1 + the nightly rollup/retention cron; volume math redone on Workers Paid and the D1 quota (storage, not writes, is what sets the 90-day window); new section documenting the dashboard worker and cross-linking it. - Fixed three drifts from the worker allowlist the sweep surfaced: schema_version was still 1, client_name/client_version was still marked "plumbing to add" though session.ts passes it today, and the legacy sqlite_backend field the worker still accepts was undocumented. - telemetry-worker/README.md: step 6 claimed a repo-wide grep came back clean, which this runbook itself falsifies. Added step 7 — deleting the runbook is what makes that grep true, and is the completion check. - smoke-cutover.sh: the vendor guarantee is now asserted by class (no analytics-ingest endpoint referenced) rather than by one vendor's name, so it keeps working once the name is gone. Verified it still catches a planted forwarding URL. 61/61 pass. Retention is documented as 90 days, not the 180 in the task notes: 180 days of raw events exceeds D1's 10 GB per-database cap, and the code purges at 90. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: untrack local Kommandr issue DB and ignore its sqlite artifacts Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f6ac7b36e6 |
fix(mcp): blast radius follows caller chains before claiming no test coverage (#1475) (#1494)
The "no covering tests found" flag only inspected a symbol's direct callers, so helpers exercised transitively by tests (logDebug runs 1,471x under npm test) were reported untested — wrong for ~40% of flagged symbols per the issue's measurement. The check now BFSes up the caller graph (3 hops, 64-lookup budget per entry) and reports indirect coverage as "tested via callers: <files>". When nothing is found it claims only what was measured — "no tests found within 3 caller hops", or the weaker "no test calls this directly" if the budget ran out — and drops the warning glyph. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
38580e0b04 |
fix(python): bare class references produce references edges to classes (#1478) (#1493)
Python's class-as-value idioms (return SomeClass, x = SomeClass, registry dicts, classes passed as arguments) produced no references edges, so callers/impact on a Django/DRF serializer missed the views that consume it. Three gates dropped them: - return_statement was never dispatched by PYTHON_SPEC (kernel mirrored) - the extraction gate (definedHere) collected function/method names only - resolution accepted function/method targets only (matchFunctionRef + the function_ref import fast path) Capture return_statement for Python (single expression; tuple returns not descended), admit same-file CLASS names to the gate, and accept class targets for Python bare identifiers — scoped to Python so the TS/JS KIND FILTER contract is untouched. The docopt false-positive mechanism behind the function-only rule (lowercase locals vs same-named methods) doesn't transfer: methods stay excluded for bare ids, and the same-file/import gate + unique-or-drop rules still apply. Probed on django-rest-framework (~250 files): 559 new references→class edges, 10/10 sampled genuine (serializer_class = AuthTokenSerializer, the ModelSerializer field-mapping registry, aliases, ctor args, isinstance). EXTRACTION_VERSION 24 → 25. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f2a5df34de |
fix(mcp): never serve a mis-sliced symbol body from a file that drifted from its index (#1474) (#1492)
codegraph_node / codegraph_explore read CURRENT bytes but slice them at INDEXED line ranges; after an un-synced edit that slice can be a DIFFERENT symbol's code served under the requested name — isError: false, introduced by the 'verbatim … do not Read' guarantee. The watcher-based pending (#403) and degraded (#876) banners cannot cover a project reached via projectPath: cross-project instances have no watcher, by construction. Freshness is now verified at the point of emission from data the index already stores: one stat per rendered file (size + floored mtime, the sync fast path's own test), sha256 content-hash compare only on stat mismatch (so a touch/identical rewrite never false-positives), memoized briefly per handler. On drift: - codegraph_node: small files ship WHOLE and CURRENT (Read-parity, still no Read needed); large ones omit the body with an explicit notice steering to the tool's file-read mode or Read. Location/signature stay, flagged as possibly shifted. - codegraph_explore: the whole-file render (already correct by construction) is kept and flagged; adaptive/skeleton/cluster slicing is disabled for drifted files — a too-big drifted file is omitted with a notice instead. The verbatim/do-not-Read header gains a per-file exception, and a trailing note flags shifted line references (flow, blast radius, symbol lists). The guarantee itself is preserved: everything actually rendered is still byte-accurate — drifted files ship whole or not at all, never as a possibly-wrong slice. A re-sync of the target project restores normal output (covered by test). Adds __setLoadCodeGraphForTests (same seam pattern as __setFsWatchForTests) so in-process tests can exercise a genuine cross-project open, which vitest's transform cannot service through the lazy require. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
02c0e2c935 |
fix(db): stop watchdog-killed sessions from leaking the SQLite WAL without bound (#1431) (#1490)
A SIGKILL'd process (the #850 liveness watchdog, OOM, a crash) leaves its WAL on disk; the next session appends to the same file; and nothing ever truncated it — PASSIVE checkpoints fold frames but keep the file at its high-water mark, and the one shrinking path (a clean last-connection close) is exactly what a killed-daemon world never takes. Observed at 25.6 GB on a 5.46 GB DB, growing until the disk filled. - journal_size_limit on every connection: resetting checkpoints now clip the WAL back to the cap instead of leaving it at its high-water mark. - healOversizedWal() fired from every DatabaseConnection.open: off-thread PASSIVE fold + TRUNCATE when the leftover WAL exceeds the cap (64 MB, CODEGRAPH_WAL_HEAL_MB to override). Single-flight per connection with bounded retries — concurrent passes defeat each other (each checkpoint sees the other as a busy reader). - Daemon/direct MCP watchdogs now pass progressPaths (DB + WAL), extending the #1231 slow-disk deferral to the long-lived server so a healthy daemon mid slow statement isn't SIGKILL'd — fewer kills, fewer leaked WALs. - codegraph status shows WAL size (human + JSON) and warns when it dwarfs the DB; daemon.log lines and the watchdog kill notice now carry ISO timestamps so kills can be placed in time. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0682137a42 |
fix(installer): write the Claude prompt hook as codegraph.cmd on Windows (#1466) (#1489)
The standalone bundle's bin dir exposes only codegraph.cmd, and Claude Code executes UserPromptSubmit hooks through Git Bash, which applies no PATHEXT — so the bare `codegraph prompt-hook` the installer wrote was "command not found" (exit 127) on every prompt. Write the platform-correct spelling, recognize both spellings on uninstall/opt-out, and self-heal an installer-written entry from the other platform in place on install/upgrade re-runs (npx/hand-edited variants stay untouched). Reproduced and validated on the Windows VM: bare form exits 127 under Git Bash on a standalone-only PATH, codegraph.cmd exits 0; full installer suite (165 tests, including the new migration coverage) green on Windows + macOS. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
572d22bfbe |
fix(installer): Codex TOML block finder preserves trailing array-of-tables siblings (#1351) (#1370)
The Codex installer's `findNextTableHeader` skipped `[[array-of-tables]]` headers instead of treating them as a block boundary, so any `[[...]]` block after `[mcp_servers.codegraph]` in ~/.codex/config.toml was silently deleted on install/upgrade/uninstall. Now treats both `[...]` and `[[...]]` as boundaries, with a small line lexer so header-shaped text inside multiline strings/arrays isn't mistaken for a boundary. Adds round-trip regression coverage (install → reinstall → uninstall) + CHANGELOG entry. Fixes #1351. Supersedes #624. Thanks @KtzeAbyss. |
||
|
|
ea72e1b190 |
docs(changelog): promote [Unreleased] into [1.5.0]
[skip ci] Auto-generated by Release workflow.v1.5.0 |
||
|
|
9b1fd6dbe8 |
release: sync package-lock.json to 1.5.0
[skip ci] Auto-generated by Release workflow. |
||
|
|
a6682c6a07 |
ci(release): kernel builds required + full walker-parity gate (#1401)
Three R1-era assumptions retired now that the kernel is the release's headline rather than an optional speedup: 1. The kernel matrix drops continue-on-error (fail-fast stays false so every platform leg reports). A Rust toolchain failure now blocks the release instead of silently shipping wasm-only bundles under a Rust-engine banner. First real risk it guards: the vendored-grammar-C languages (kotlin/lua/scala/dart, incl. scala's 35MB parser.c) have never compiled on these runners — no release has run since the kernel merged. 2. The release-job gate expands from the two R1 suites to ALL __tests__/kernel-*.test.ts (14 files today: contract, grammar-source parity, and every language's walker byte-parity suite) — the glob keeps it current as languages land. 3. A missing linux-x64 prebuild at the gate is now a hard failure (the matrix guarantees it; absence means a wiring bug), and the artifact download step loses its best-effort flag for the same reason. No packaging changes needed: build-bundle.sh already stages lib/kernel/codegraph-kernel.node per target and pack-npm.sh repacks bundles verbatim into the per-platform npm packages. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7d4a3d0f2d |
docs(release): README polish + v1.5.0 (#1400)
- Hero: larger theme-aware standalone Rust logo (new assets/rust-logo{,-dark}.svg
— gear only, no tile card; <picture> swaps by GitHub theme), tagline text
trimmed to 'Kernel powered by Rust'
- 'Built for speed' section: removed the floated language-tile logo (its baked-in
paper card rendered as an odd box on dark theme and pushed the text)
- Removed the '1.0 Released!' banner and the Star History section (+ its
Contents entry)
- package.json → 1.5.0 for the Rust-engine release
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0f1096e238 |
docs: Opus 4.8 benchmark re-validation + release-notes headline (#1399)
README Benchmark Results re-run 2026-07-21 on the current build (Rust kernel + this cycle's resolution overhaul), Claude Opus 4.8, 7 repos, median of 4 runs/arm: 89% fewer tool calls, 60% cheaper, 69% fewer tokens, 20% faster on average, file reads 0-vs-1..24 on ALL seven repos. Per-repo floor effects reported honestly (excalidraw/alamofire wall, okhttp cost wash). Cost note updated to match the measured data. Changelog [Unreleased] headline now co-leads with near-instant sync. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f8e6f0066c |
chore: gitignore target-linux/ cross-build cache (two cache files slipped into #1397) (#1398)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c74e8b05e0 |
perf(sync): adaptive quick-fire debounce + scoped watcher sync — save-to-graph well under a second at any scale (#1397)
Two changes to the watcher path (the always-on daemon every agent session uses), which previously paid a flat 2s debounce plus a full-tree scan-diff on every save even though the OS events name the exact files: 1. Adaptive debounce: a pending set of ≤2 files fires after a 300ms quiet window; ≥3 keeps the full configured window so agent multi-file bursts coalesce exactly as before. Re-arming preserves trailing-edge semantics; a user-set CODEGRAPH_WATCH_DEBOUNCE_MS remains the authoritative upper bound (quick window never exceeds it, floor 100ms). 2. Scoped sync: watcher-triggered syncs pass their pending paths, and the reconciler stats exactly those — per-path logic identical to the full walk (stat pre-filter, hash confirm, the #1240 removal/resurrection flow) — skipping the O(repo) scan and tracked-load. Strict fallbacks keep the full scan-diff as ground truth: directory removals (#1285 — the events can't name the children), empty pending sets (retry paths), and >500-file storms (branch checkouts, which also self-heal anything event coalescing dropped). filesChecked counts examined PATHS so a deletion-only scoped sync can't mimic the #449 lock-unavailable signature. Measured (warm in-process, the daemon path): dubbo one-file sync work 512→335ms, Swift compiler (27k files) 884→385ms — save-to-fresh-graph ≈0.6-0.7s end-to-end including the quick debounce, from ~2.5-6s perceived before. Gates: scoped-vs-full dumps byte-identical on dubbo AND the Swift compiler; watcher suite 30/30 (3 new: scoped pass-through, dir-removal fallback, quick-fire timing); sync suite 34/34 (4 new scoped-parity cases incl. delete-resurrection and the lock signature); full suite 2,696 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
157c8e735d |
perf(resolution): generation-tagged supertype memo + method owner index — Swift compiler 185→98s, byte-identical (#1395)
The swiftc =2/nm:mc-* attribution located the wall: the getSupertypes conformance walk ran 971,200 times (565s of combined worker time, 581µs each) — every resolveMethodOnType miss re-queried implements/extends edges for every same-named type node, recursing depth-4 through Swift's protocol landscape with no memoization, and post-inference resolveMethodOnType averaged 1,912µs per call. Fix 1 — generation-tagged getSupertypes memo. Supertype edges GROW during the resolution loop (batch k persists its edges BEFORE batch k+1 fans out — the #1320 ordering), so a plain cache would freeze an early batch's emptier answer. Within a batch the edge state is fixed by that same ordering, so memo entries carry a generation that advances at every batch entry point (resolveBatchYielding / resolveListForAdmission — covering the sequential loop, pool workers, sync admission, and the conformance pass); a stale-gen entry recomputes. Behavior-identical to no memo at every point in time; walk invocation counts match the unmemoized run exactly (971,200 / 24,336 / 76,415). Fix 2 — per-(language, method-name) owner index in getMethodMatches: candidates bucket once by their qualifiedName's last two segments (exactly the span the match predicate tests), so a (type, method) query is a map lookup instead of an O(candidates) scan per methodMatchCache miss. ObjC selectors and multi-segment typeNames keep the legacy linear path. Also ships nm:mc-rmot / nm:rmot-supers =2 attribution rows. swiftc: settle 100.2→31.3s, resolveMethodOnType 1,912→202µs, wall 183.5→97.8s (was 185s at the head-to-head; cbm's same-box number is 119.1s). Gates: swiftc old-vs-new dump byte-identical (1,837,235 rows), swiftc pooled-vs-CODEGRAPH_NO_PARALLEL_RESOLVE=1 identical (the generation-semantics risk surface), dubbo old-vs-new identical (49k Java instance-method hits share both paths), Alamofire identical; suite 2,689 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
974e6c8b95 |
perf(resolution): incremental receiver-inference scan memo + compiled-pattern memo — kong −8% more (−23% cumulative), byte-identical (#1392)
The kong/tokio matcher-chain residue attributed (nm:mc-* sub-stage rows, shipped here too): matchMethodCall's cost is ~entirely inferLocalReceiverType — 61µs per miss on kong, 99% miss rate (39k `self:` calls hunting a local declaration Lua never writes), re-scanning the same scope lines for every ref. Two pure memos, both semantics-preserving by construction: - Compiled-pattern memo: localReceiverTypePatterns/phpPropertyTypePatterns built 2-4 fresh RegExp objects per call; patterns are a pure function of (language, receiver) and non-global, so instances are shared via a FIFO-capped map (no per-get mutation — the §7a.6 LRU-churn lesson). - Incremental scan memo: refs for the same (file, scope, receiver) arrive in ~ascending line order and the backward declaration scan is a pure function of immutable file lines — a per-context watermark scans each line once per key (query(c) = highest match in [start..c]; monotonic calls extend the watermark over (hi..c]; non-monotonic calls fall back to the plain bounded scan). componentScoped (CFML/PHP whole-file sweep) is keyed out. States drop with the context's file caches via clearNameMatcherMemos, wired into ReferenceResolver.clearCaches. kong mc-infer misses 61→20µs (2.4s→0.8s combined); fresh index 3.43 → 3.03-3.20s (−8%; 4.07 → 3.14 cumulative with #1391). tokio unchanged (tight scopes). Gates: dubbo (49k Java instance-method HITS ride this scan), kong, tokio, Fusion dumps all byte-identical; suite 2,689 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
abb0a916f7 |
perf(resolution): per-context basename index for Lua/Luau require resolution — kong fresh index −16% (#1391)
The full-README competitor matrix ranked lua/kong as the largest legitimate fresh-index gap (2.71×). Stage attribution (RESOLVE_PROFILE=2) pinned it: resolveLuaRequire ran getAllFiles().filter(endsWith) FOUR times per require ref — ~7.5k string suffix scans each, measured at ~0.9ms/ref, hit or miss (2.7s combined over kong's 3k requires). Replace the per-ref full-list scans with a per-context basename → file-paths index (the cobolCopybookIndexes pattern). Buckets preserve getAllFiles() iteration order, so each suffix's candidate filter yields exactly the array the full scan produced — identical matches, identical stable sort, identical winner, dump-proven. kong: 4.07-4.27 → 3.40-3.63s (−16%). Gates: kong old-vs-new dump byte-identical (157,650 rows), kong pooled-vs-sequential identical, Fusion (luau instance-path requires) old-vs-new identical; suite 2,689 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Also registers cobolCopybookIndexes in clearImportResolverMemos — it was never dropped on cache clears, so post-sync copybook lookups could serve a stale file list; cache-drop is the always-safe direction. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1aa4de6eaa |
perf(resolution): adaptive pool engagement — projected-settle bar replaces the fixed 150k-ref gate for mid-run boot; tokio −23% (#1390)
The 9-language competitor matrix exposed tokio as the worst fresh-index gap: 77% of its wall was resolution running SEQUENTIALLY — 56k Rust refs sit under the fixed 150k pool gate while costing 36µs each (9× Go's 4µs/ref on prometheus). A ref-count gate can't see per-ref cost. After each sequential batch the loop now projects the remaining sequential settle from the measured rate and boots the pool mid-run when it clears 400ms. The switch rides machinery that already existed: pool boot is async and fan-out engages only when ready, admission order is mode-independent, and the #1320 edges-before-fanout invariant holds at every batch boundary regardless of when the pool arrives. Up-front engagement at >=150k refs is unchanged; 2-core/low-memory hosts still decline inside tryCreate's sizing; CODEGRAPH_NO_PARALLEL_RESOLVE still disables; downgrade permanence is preserved (one engage attempt per run). Measured (n=3, interleaved, caffeinated): tokio 3.06-3.12 → 2.40-2.57s (resolution 2,443→~1,330ms); express (tiny control) unchanged with zero engagements; dubbo unchanged (ref-count path). Gates: tokio + excalidraw adaptive-vs-sequential dumps byte-identical (87,302 / 89,903 rows), dubbo dump identical to the session baseline, suite 2,689 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Known follow-up: sampling the rate mid-first- batch would close the remaining ~0.3s to the forced-engage ceiling. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
082ea65f3a |
perf(synthesis): provably-empty pass gates + prefilters — render/expo/rn/mybatis stop scanning repos they can't match; iface memo (#1389)
Store-arc round 2 (#1388 follow-up). The synthesis pool barrier on dubbo carried ~1.4s of passes that provably could not emit an edge for the project: reactRenderEdges fanned out over every class before checking for a render method (now: one indexed name lookup bounds candidates — not a language gate, Java Litho-style render+setState still matches); expo/rn cross-platform pairing streamed every method row without the languages their edges require (now registry-gated: expo needs swift AND kotlin file-languages, rn needs a JS-family caller for isBridge); mybatis built its full java-method index before discovering there were no mapper-XML methods (now collects the XML side first). ifaceEdges — real work — stops re-fetching a hub interface's methods once per implementer and skips supertype-less classes before any per-class lookup. dubbo warm wall 8.49-8.79 → 8.14-8.24s (n=3/arm, caffeinated); barrier 784→435ms; the full removed pass work lands on low-core envelopes where synthesis runs sequentially. Dumps byte-identical: dubbo old-vs-new, pooled-vs-sequential, kernel-vs-wasm (441,270 rows) + excalidraw JSX-live control (89,903 rows, 46 react-render edges reproduced). Suite 2,689 ×2 with CODEGRAPH_KERNEL_EXPECT=1. Also ships the diagnostics that located the round (zero cost when off): CODEGRAPH_RESOLVE_PROFILE=2 attributes per-ref time to resolveOne's strategies (stage:*) and the name-matcher's sub-matchers (nm:*); CODEGRAPH_SYNTH_TIMINGS now prints the store worker's decode-vs-SQL split. Killed by measurement, recorded in the PR: import-failure negative cache (both-outcome names exist — static imports resolve via instance-method on jvm-miss), jvm-miss early return (1,939 later-strategy edges), jsxEdges language gate (Java generics text produces jsx edges), and §4d buffer→bind on Spring repos (extract() hook forces the decoded path — kernel=0 bundles measured). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
27c3c55436 |
perf(resolution): darwin-honest memory budget — vm_stat-based availability unstrangles the resolver pool on macOS (#1388)
Post-R7b store-arc round 1, found by the dubbo warm-wall decomposition (the cbm bar): resolution's loop-stage profile showed settle=3.0s — the main thread idling on TWO resolver workers on an 11-core Mac. Pool sizing logged `size=2 (budget=1068MB)`: memoryBudgetBytes() falls back to os.freemem() when uncontained, and macOS keeps RAM deliberately full of reclaimable cache, so freemem reads ~1GB on a mostly-idle 64GB machine. The memory term then capped the pool at 2 where the CPU term allowed 6 — the macOS sibling of §7a.1's os.cpus() cpuset-blindness (that round fixed the CPU term; this fixes the memory term). Fix: darwinMemoryAvailable() reads /usr/bin/vm_stat once per sizing call and reports free + inactive + speculative + purgeable pages — what Activity Monitor calls available, the same reclaimable-inclusive convention the Linux branch already uses by crediting inactive_file back. Parse failure → null → freemem fallback; Linux/cgroup and Windows paths untouched. Measured (dubbo 4,402 files, warm, caffeinated, n=3 each): pool now self-sizes to 6 (budget 5.7-6.3GB) — wall 8.62-8.83s vs 9.67-10.87s baseline, resolution phase 6.9→5.3s, loop settle 3.0→1.9s. Matches the CODEGRAPH_RESOLVE_WORKERS=6 probe exactly (probe-before-build). Dumps byte-identical pool-6 vs sequential (441,270 lines). Second consumer unblocked: the cFnPtr LRU cache cap no longer spuriously degrades to 128 on Macs (its full-cache tier is worth ~60s at kernel scale). Suite: resolver-pool-sizing gains a darwin-gated reclaimable-pages test + an off-darwin null pin; full suite 2,689 green ×2 with CODEGRAPH_KERNEL_EXPECT=1. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3c1f30ab48 |
docs(kernel): mark R7b complete — 20 languages default-routed, Linux leg validated (#1387)
Flips the migration plan's R7b milestone to done: eleven languages across four same-day batches (rust #1371; csharp/ruby/php #1378-#1380; swift/ kotlin #1381-#1382; r/lua+luau/scala/dart #1383-#1386), batch 4 going 4-for-4 first-run parity (12-of-13 arc-wide). Also records the batch-4 upfront grammar-probe method and the dart wasm byte-copy vendor. Validation note: the full suite (2,688 tests) also ran green on linux-arm64 in a fresh rust:1-bookworm + node 22 container with the kernel built from scratch and CODEGRAPH_KERNEL_EXPECT=1 — the Linux leg for all 11 post-R7a walkers. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d1b75a1a27 |
feat(kernel): R7b Dart walker — dart module, vendored-grammar-C d4d8f3e + wasm byte-copy vendor, dart default-routed (#1386)
R7b batch 4 #4 — the FINAL R7b language (docs/design/dart-kernel-port-checklist.md is the authoritative quirk list). The fourth vendored-grammar-C language, with a twist: production dart resolved its wasm from tree-sitter-wasms, whose dart dependency is an UNPINNED github:UserNobody14/tree-sitter-dart — a routine dependency update would have silently changed dart's grammar. This PR byte-copies the shipping 0.1.13 artifact into src/extraction/wasm/ (VENDORED_WASM_LANGS += dart) and compiles the same-commit (d4d8f3e337d8) parser.c/scanner.c in the kernel — table identity proven by the kernel-grammar-parity row. crates.io tree-sitter-dart is the nielsenko fork (different lineage) — rejected. The center of gravity is THE SIBLING-BODY DOUBLE-WALK, reproduced bug-for-bug: dart attaches every function/method body as a NEXT SIBLING of its signature, and the TS walkers consume each body TWICE — once via resolveBody (attributed to the function/method) and once via the enclosing generic walk (attributed to the file/class). Duplicate local-function nodes with the SAME id under different parents, duplicated calls/instantiates refs, and file/class-attributed fn-ref twins all emit in the exact observed interleave (a dedicated fixture pins the duplicate-id rows; the bloc kind-census spot-check pins the counts). Also preserved (probe-pinned): the extractBareCall selector matrix (the first callTypes=[] language — cascades completely invisible, `?.` encodes like `.`, the `ConfigT.load()` calls+references double emission with no callee-of-call skip, capitalized-chain `Foo.create().run` re-encode, const-object callee names); the constructor hooks (unnamed ctor skipped, named ctors/factories renamed to the CTOR name with the class as returnType, `@override (T) m()` record-misparse rescued by class-name validation); operator methods minting `method "<anonymous>"`; static_final_declaration constants via the visitNode hook while instance fields mint NOTHING; the prefixed-return-type prefix bug (`other.OtherClass f()` → returnType `other`); enum `with` mixins silent vs `implements` working; anonymous extensions named after the ON type; deferred imports invisible; named-argument callbacks NOT fn-ref-captured (the Flutter `onPressed:` idiom — future accuracy PR, TS-side first); `async*`/`sync*` NOT async; value-refs with the LIVE dart sibling-body pull and the `$X`-vs-`${X}` interpolation asymmetry; dartdoc kept in all three comment forms with the annotation-broken chain. Gates: parity sweeps first-run 0-diff on shelf/bloc/flutter — 5,815 clean files byte-parity, deferrals 10/21/1341 ≈ the survey's 10/21/~1340 (both-arm grammar reality: empty object patterns — the sealed-class idiom — and unnamed `library;` dominate; --max-deferral 0.3); full-init dumps byte-identical ×3 (shelf 7,959 / bloc 40,026 / flutter 1,855,319 dump lines); bloc per-kind node census identical across arms (the double-walk duplicate rows survive the store identically); kernel-dart-parity suite (7 fixtures + in-memory CRLF variants + double-walk duplicate-id pin + generated-file skip pin + two defer pins); full suite 2,688 green ×2 with CODEGRAPH_KERNEL_EXPECT=1. DEFAULT_ROUTED += dart (20 langs — R7b COMPLETE). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bdd687b49f |
feat(kernel): R7b Scala walker — scala module, vendored-grammar-C master@0aca5d0a6f, scala default-routed (#1385)
R7b batch 4 #3 (docs/design/scala-kernel-port-checklist.md is the authoritative quirk list). The third vendored-grammar-C language and the biggest grammar in the tree (35MB parser.c): the vendored wasm is tree-sitter/tree-sitter-scala master@0aca5d0a6f — a post-v0.26.0 generation sync that is not a release (the 0.26.0 crate is 30 states BEHIND, so a crate pin would be a silent downgrade). NO wasm change: production has parsed with this exact revision since #91 — the kernel-grammar-parity row (ABI 15, 26,650 states, 32 fields, id-by-id tables) is the whole alignment proof. Preserved bug-for-bug (all probe-pinned): the leak-through asymmetries — extension methods mint NO nodes (first def's body calls leak to the enclosing scope, later defs invisible, and the braced form resolves its body field to the `{` TOKEN via first-match-wins field lookup → whole extension invisible); anonymous `new T { … }` template_body members leak to the enclosing scope (findAnonymousClassBody misses template_body); the bodied-vs-bodiless class asymmetry (bodiless headers walk class_parameters → default-value calls emit FROM the class; bodied ones never see them) — plus first-segment import names (`import com.example.C` → `com`), the val/var hook keyed on the enclosing-definition NODE TYPE (object vals → constants/value-ref targets, class/trait/enum/given vals → fields) with consumed initializers, every def routed through extractMethod with the top-level function fallback, nested defs in bodies minting NOTHING (the inverse of kotlin) while body-local classes extract fully, curried signatures keeping only the FIRST parameter list (type params win the `parameters` field), enum cases positioned at the CASE node with invisible params/extends tails, extends with-chains via scalaBaseTypeName, `@deprecated(args)` decorates, the #750 capitalized-chain re-encode (`WidgetS.create().render`), literal-receiver silence, static-member reads AND writes, infix invisibility, `derives` silence, scaladoc retention with the CRLF `\r` pin, full value-reference machinery (shadow prune, last-wins same-name targets, `$X`/`${X}` interpolation reads), and SCALA_SPEC fn-refs (bare ids + postfix eta unwrap + varinit, var-init non-capture). Gates: parity sweeps first-run 0-diff on os-lib/cats/scala3-compiler-src/ scala3-library-src — 1,935 clean files byte-parity, deferrals 0/15/57/116 matching the survey's predictions exactly (scala-3's PHANTOM hasError files — flag-true, zero ERROR nodes, capture-checking `^` — defer on the FLAG); full-init dumps byte-identical ×3 (os-lib, cats, scala3 whole-repo 950,889 dump lines); kernel-scala-parity suite (9 fixtures + 9 in-memory CRLF variants incl. Scala-3 indentation through the external scanner + phantom/real-error defer pins + first-segment/namespace/value-ref pins); full suite 2,669 green ×3 with CODEGRAPH_KERNEL_EXPECT=1 (kernel-scaffold's stays-wasm example moved scala → pascal). DEFAULT_ROUTED += scala (19 langs). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e32135171e |
feat(kernel): R7b Lua+Luau walker — one lua module, vendored-grammar-C lua v0.4.1, tree-sitter-luau 1.2.0 pin, both default-routed (#1384)
R7b batch 4 #2 (docs/design/lua-luau-kernel-port-checklist.md is the authoritative quirk list). ONE walker for both dialects (ccpp precedent) — the differences are exactly four: luau's type_definition aliases, the `export `-slice isExported hook, the return-type signature suffix, and the grammar handle. Grammar prep is kernel-side only, no wasm change: lua is the SECOND vendored-grammar-C language (the vendored wasm is the v0.4.1 tag, a revision not on crates.io — tag artifacts compiled via build.rs, shas pinned); luau is a plain crate pin =1.2.0 whose tarball is sha-identical to the tag (the swift tag≠crate divergence does not recur). Grammar-parity rows replace the bump gate entirely. Preserved bug-for-bug (all probe-pinned): the require/visitNode-hook ASYMMETRIES (top-level requires — including inside top-level if/for/while — mint import nodes while the identical body-level statement emits `calls "require"`; top-level `local x = foo()` initializers are invisible while global `x = foo()` calls emit), the BFS string-win inside require args (`require(script:WaitForChild("Kid"))` → import Kid) and Roblox instance paths, receiver-QN methods (`M.sub.deep::chained`, `_G::installed`, stack-QN nested globals like `render::leakedGlobal`), the raw-text callee world (colon forms with `self` never stripped, bracket callees, newline-glued chains byte-verbatim, the `(handler)` paren-conversion), LUA_SPEC function-as-value capture with the `M.cb = cb` param-storage skip and first-occurrence dedupe, LuaDoc `---` keeping a leading `- ` plus `--!strict` joining docstring chains (block-comment docstrings keep interior CRLF bytes), variable nodes at the IDENTIFIER with positional value pairing, duplicate same-(kind,name,line) ids, and the lua↔luau isExported wire divergence (lua functions: flag absent; luau functions: present-false; methods: absent in both; variables: present-false in both; `export type`: true). Gates: parity sweeps first-run 0-diff on kong/lazy.nvim/lua-resty-core (lua) + lune/Fusion (luau) — 1,734 clean files byte-parity, deferrals 1/0/0/3/8 matching the survey's both-arm predictions exactly (kong's 1 = a deliberately invalid fixture; luau's = grammar-inherent generic type packs and default type params); full-init dumps byte-identical kernel-vs-wasm ×4 (kong 157,650 dump lines); kernel-lua-parity suite (both torture fixtures + in-memory CRLF variants + glue-chain, duplicate-id, and cross-dialect defer pins + kernel-arm wire-flag pins); full suite 2,647 green ×2 with CODEGRAPH_KERNEL_EXPECT=1. DEFAULT_ROUTED += lua, luau (18 langs). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |