From 5dd4db68cdf4f112ffbf1e688295132bb3ad4db2 Mon Sep 17 00:00:00 2001 From: Colby McHenry Date: Wed, 5 Aug 2026 13:06:51 -0500 Subject: [PATCH] =?UTF-8?q?docs:=20the=20occupancy=20baseline=20says=20our?= =?UTF-8?q?=20residual=20is=20higher=20=E2=80=94=20write=20that=20down=20(?= =?UTF-8?q?CG-13)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fills the empty RESULTS placeholder with the 2026-08-05 campaign (bgjob-6d357cd2: 7 repos x 2 arms x 4 runs x 3 turns, 137 min). The finding is not the flattering one. Retrieval residual is 82% HIGHER with codegraph and share-of-context 27% higher, on all seven repos — vscode 67k resident against 18k. At the same time six of seven without-arms *process* more total tokens (gin 660k vs 290k) while leaving less behind. Both are true: one dense verbatim payload stays resident where many small Read/Grep results evict. This corroborates issue #1500 on our own harness; the aggregator used to print it as "-82% lower with codegraph" until the sign bug at 520ed9d. Also: - States the regime everywhere. This ran claude-sonnet-5 / 3-turn; the README's table is Opus 4.8 / single-question. Measured 24/23/20/84 against the published 60/69/20/89 — model and turn count, not contamination. Records the two inversions honestly (vscode processes 98% more tokens, django costs 17% more) and that 4 of 28 with-arm sessions still touched Read. - Corrects the "Settled" section, which claimed a 7-repo baseline existed before one did, and adds the unclaimed Opus rerun. - Records the contamination gate: 0 CLI calls returned output in 56 sessions, but 29 attempts were blocked — 26 of 28 without-arm sessions tried. no-cli-shim.sh is load-bearing, not precautionary. - Records the secondary readings as absolute, not before/after: 86.7% allocation efficiency pooled over 110 calls, read-of-a-file-we- returned 2%, explore-again 73% and ambiguous by construction. README.md is deliberately untouched — restating its numbers from sonnet 3-turn data would be wrong. A proposed README paragraph is drafted at the end of the benchmark doc for the maintainer to accept or reject. Co-Authored-By: Claude Opus 5 --- .../benchmarks/agent-eval-feedback-metrics.md | 5 +- docs/benchmarks/residual-context-occupancy.md | 217 +++++++++++++++++- 2 files changed, 217 insertions(+), 5 deletions(-) diff --git a/docs/benchmarks/agent-eval-feedback-metrics.md b/docs/benchmarks/agent-eval-feedback-metrics.md index 797804d..c762498 100644 --- a/docs/benchmarks/agent-eval-feedback-metrics.md +++ b/docs/benchmarks/agent-eval-feedback-metrics.md @@ -54,7 +54,10 @@ table still renders against whichever arm's logs are already in `$AGENT_EVAL_OUT **A campaign — `bench-readme.sh`.** The 7 README repos, three turns each, `RUNS` per arm, through `run-all.sh` — so every run in a campaign carries all -three metrics. Aggregate with `parse-bench-readme.mjs`. +three metrics. Aggregate with `parse-bench-readme.mjs`. One has been run: +[the 2026-08-05 baseline](residual-context-occupancy.md#baseline-the-7-readme-repos) +(sonnet, 3 turns, 4 runs/arm) — read its regime box before comparing anything to +it, and note that it is **not** the regime the README's table was published in. **A log you already have.** `parse-run.mjs [run.tN.jsonl …]` prints the three blocks for any stream-json log; `--brief` drops the numbered call diff --git a/docs/benchmarks/residual-context-occupancy.md b/docs/benchmarks/residual-context-occupancy.md index afb6d8d..6570ecc 100644 --- a/docs/benchmarks/residual-context-occupancy.md +++ b/docs/benchmarks/residual-context-occupancy.md @@ -168,18 +168,185 @@ without-arm may have been using codegraph.** ## Baseline: the 7 README repos - +### The regime — read this before any number below + +| | This campaign | README's published table | +|---|---|---| +| Model | **`claude-sonnet-5`** | **Claude Opus 4.8** | +| Session shape | **3 turns** (README question + 2 in-flow follow-ups) | **1 question** | +| Runs | 4 per arm × 7 repos = **56 sessions** | 4 per arm, median | +| Ran | 2026-08-05, 137 min (`bgjob-6d357cd2`), raw under `/tmp/ab-readme` | 2026-07-21 | + +**These two are not comparable, and the difference is model + turn count — not +contamination.** Sonnet is the deliberate floor model for this harness +(`CLAUDE.md`: an affordance that lands on Sonnet generalizes up; one that only +works on Opus does not generalize down). Three turns is what makes occupancy +chargeable at all. Both choices move the efficiency numbers, so the throughput +row below reads *lower* than the README's and neither figure invalidates the +other. Settling whether the published Opus figures still hold needs a matched +**Opus 4.8, single-question** rerun; that is deliberately out of scope here. + +Reproduce with: + +```bash +CORPUS=/tmp/codegraph-corpus scripts/agent-eval/bench-readme.sh # RUNS=4 CG_TURNS=3 +node scripts/agent-eval/parse-bench-readme.mjs /tmp/ab-readme +``` + +### The finding: codegraph's residual is 82% HIGHER, on all seven repos + +``` +repo turns W→WO final ctx W→WO residual W→WO % of ctx W→WO % of window W→WO +vscode 12.5/44 113k→53k 67k→18k (+276%) 59.7%→36.0% 33.7%→9.0% +excalidraw 9/32 87k→57k 43k→25k (+71%) 49.5%→47.6% 21.5%→12.5% +django 7.5/16.5 60k→51k 18k→10k (+71%) 29.3%→19.9% 8.8%→5.1% +tokio 9/33.5 87k→64k 45k→31k (+45%) 52.1%→50.5% 22.7%→15.7% +okhttp 6/14 61k→59k 20k→16k (+27%) 33.2%→27.6% 10.1%→8.0% +gin 6/15.5 56k→49k 15k→8k (+79%) 26.3%→16.7% 7.3%→4.1% +alamofire 10/31 76k→65k 34k→32k (+7%) 44.7%→50.6% 16.9%→15.8% + +AVERAGE: retrieval residual 82% HIGHER with codegraph · share-of-context 27% HIGHER +``` + +W = codegraph's responses still resident. WO = Read + Grep/Glob + Bash results +still resident. `turns` is median assistant turns per session. + +**Seven of seven.** There is no repo where codegraph leaves less behind. On +vscode it leaves **67k tokens resident against the without-arm's 18k** — a third +of a 200k window, gone before turn 4 starts. The only near-tie is Alamofire +(+7%), and it is a tie because that arm's share-of-context is actually *lower* +(44.7% vs 50.6%), not because the residual is small. + +**Both things are true at once.** On six of seven repos the without-arm +*processes* far more total tokens than the with-arm — gin 660k vs 290k, okhttp +704k vs 302k — while leaving *less* behind. Throughput and stock are different +quantities and they point opposite ways here: + +- codegraph front-loads **one large verbatim payload** (2 explore calls on gin, + each tens of thousands of dense source characters) and that payload **stays + resident** for every turn after it; +- Read/Grep/Bash churn **many small results** (gin: ~6 reads + ~5 bash per run), + most of which are re-derivation the agent then discards, and which evict. + +**This corroborates issue [#1500](https://github.com/colbymchenry/codegraph/issues/1500) +on our own harness.** The reporter's complaint was exactly this axis, and until +this campaign we had no measurement that could see it. Note for anyone reading +git history: the aggregator originally printed this as "-82% *lower* with +codegraph" — a sign bug, fixed at `520ed9d`. The honest number is the entire +point of the metric; do not soften it. + +**Fixed overhead.** codegraph's tool schema + MCP instructions cost **+546 tok** +of context before any tool is called (median with-arm `ctxBase` minus median +without-arm `ctxBase`, averaged over repos). Paid whether or not the agent ever +calls codegraph. Small — the residual, not the schema, is where the context goes. + +### Throughput in the same campaign (sonnet · 3 turns) + +Reported for completeness and because the occupancy finding only means anything +read against it. **These are not the README's numbers and must not be quoted as +such.** + +``` +repo time W→WO tools W→WO tokens W→WO (saved) cost W→WO (saved) +vscode 2m 59s→1m 59s 8→60 949k→478k (-98%) $1.21→$1.62 (25%) +excalidraw 1m 45s→2m 1s 5→44 557k→869k (36%) $0.78→$1.03 (24%) +django 1m 4s→1m 30s 3→13 366k→686k (47%) $0.52→$0.45 (-17%) +tokio 1m 47s→4m 35s 5→43 574k→793k (28%) $0.67→$1.25 (47%) +okhttp 49s→1m 25s 2→11 302k→704k (57%) $0.35→$0.48 (27%) +gin 1m 6s→1m 43s 2→12 290k→660k (56%) $0.38→$0.43 (12%) +alamofire 1m 35s→1m 47s 6→29 545k→870k (37%) $0.67→$1.30 (49%) + +AVERAGE saved: cost 24% · tokens 23% · time 20% · tool calls 84% +``` + +| | this campaign (sonnet, 3-turn) | README (Opus 4.8, 1-question) | +|---|---|---| +| cost saved | **24%** | 60% | +| tokens saved | **23%** | 69% | +| time saved | **20%** | 20% | +| tool calls saved | **84%** | 89% | + +Tool-call reduction and wall-clock survive the regime change almost intact; the +cost and token savings roughly halve. Two repos invert outright — **vscode +processes 98% *more* tokens with codegraph** (8 calls of dense source against a +without-arm that mostly greps), and **django costs 17% more**. Neither is hidden +here. The with-arm is also not read-free in this regime: 4 of 28 with-arm +sessions still touched Read (vscode run4 `rd5 bs7`, tokio run2 `rd3 bs2`, +django run4 `rd1`, alamofire run2 `rd1`), against the README's "zero file reads +on all seven repos" under Opus. + +### Contamination gate: clean, and the channel is real + +**0 CLI calls returned output in any of the 56 sessions.** The aggregate is +uncontaminated and no run was dropped. + +But **29 attempts were blocked** — 26 in the without-arm (in **26 of its 28 +sessions**) and 3 in the with-arm. Ninety-three percent of without-arm sessions +tried to reach codegraph through Bash and were stopped by the sanitized PATH + +PreToolUse hook (`no-cli-shim.sh`). That is not a hypothetical channel the +harness guards out of caution; it is the agent's *default* move once it notices +`.codegraph/` in the tree. **`no-cli-shim.sh` is load-bearing** — without it this +campaign would have been codegraph-over-CLI vs codegraph-over-MCP, exactly as the +earlier 14-of-15 pass was (see the section above). Check the contamination row +before believing any number from this harness. + +### Secondary readings — absolute, not before/after + +There is **no baseline-build arm in this campaign** — every number below is the +current build's absolute reading on these questions. Allocation efficiency in +particular is *relative* (attribution is by citation): it compares builds on the +same question and says nothing on its own about waste. For a before/after +allocation A/B see [`explore-allocation-ab-1500.md`](explore-allocation-ab-1500.md). + +``` +repo calls again read-ret read-miss grep MOVED ON alloc eff envelope +vscode 26 21 81% 0 0% 1 4% 1 4% 3 12% 63.2% 442k +excalidraw 18 14 78% 0 0% 0 0% 0 0% 4 22% 94.7% 338k +django 11 6 55% 1 9% 0 0% 0 0% 4 36% 96.2% 186k +tokio 16 12 75% 0 0% 0 0% 1 6% 3 19% 92.9% 343k +okhttp 8 4 50% 0 0% 0 0% 0 0% 4 50% 97.9% 151k +gin 8 4 50% 0 0% 0 0% 0 0% 4 50% 99.0% 116k +alamofire 23 19 83% 1 4% 0 0% 0 0% 3 13% 89.0% 306k + +POOLED (110 answered explore calls): + explore again 73% · Read a file we returned 2% · Read a file we did NOT return 1% + · Grep/Glob 2% · moved on / answered 23% +POOLED allocation efficiency: 86.7% over 110 calls / 1.9M chars +``` + +- **Allocation efficiency 86.7%** pooled. vscode is the outlier at 63.2% — the + largest envelope (442k chars) and the lowest citation share, which is where an + allocation change would show up first. +- **`Read a file we returned` = 2%** (2 of 110). Right file, wrong bytes is + nearly absent; the allocation misses this metric was built to catch are not + what is driving vscode's number. +- **Recall misses** are 3% total (1 read-miss, 2 grep). +- **`explore again` = 73%** and is **ambiguous by construction** — it is + indistinguishable between "the first call was insufficient" and "the agent is + working through a 3-turn session and this is turn 2's first call." In a 3-turn + regime that ambiguity is much larger than it was single-turn; treat the + high-`again` repos (alamofire 83%, vscode 81%) as unresolved, not as failures. --- ## What this settles, and what it does not -**Settled.** The metric exists, it is measured rather than estimated, it runs over -multi-turn sessions — the regime where occupancy is actually charged — and there -is a baseline across the 7 README repos to compare future changes against. +**Settled.** The metric exists, it is measured rather than estimated, and it runs +over multi-turn sessions — the regime where occupancy is actually charged. As of +2026-08-05 there is a baseline across the 7 README repos (above) to compare +future changes against, and it says codegraph's residual is **higher**, on every +repo. (Before that campaign this section claimed such a baseline existed when it +did not; it does now, and it is one regime — `claude-sonnet-5`, 3 turns — not a +general result.) **Not settled, and deliberately not claimed:** +- **The README's efficiency figures.** The baseline above ran sonnet / 3 turns; + the README published Opus 4.8 / single-question. The gap between 24/23/20/84 + and 60/69/20/89 is regime, not regression, and this campaign cannot tell you + which way the published numbers have moved. That needs a matched **Opus 4.8, + single-question** rerun. Out of scope here, and `README.md` was deliberately + left untouched. - **A different host.** The reporter was in Cursor. We measure Claude Code. Window size, system prompt, and compaction policy all differ, so the *share* numbers do not transfer host to host; the ratio between the arms is the part @@ -203,3 +370,45 @@ is a baseline across the 7 README repos to compare future changes against. *enough* — that is [explore sufficiency](explore-sufficiency.md), which every run now prints alongside this block — nor about how much of the returned bytes the answer actually used (CG-9). + +--- + +## Proposed README wording — for the maintainer, not applied + +`README.md` is **deliberately untouched by this work.** Its benchmark table is +Opus 4.8 / single-question and nothing measured here can restate it. What follows +is a *proposal*: the occupancy finding as an honest counterweight to the +efficiency table, phrased so it does not depend on the sonnet-vs-Opus regime for +its claim. Accept, reject, or rewrite — this is not a pending edit. + +Suggested placement: immediately after the "A note on cost" paragraph (README +line ~195), as a second `>` note under the same table. + +> **A note on context.** The efficiency table above measures *throughput* — +> tokens processed, tools called, dollars spent to reach one answer. It does not +> measure what is still sitting in the window afterward, and on that axis +> CodeGraph costs more, not less. Across the same seven repos in multi-turn +> sessions, CodeGraph's responses leave **~80% more retrieval context resident** +> at the end of a session than the file-reading agent's do — on VS Code, 67k +> tokens against 18k. The mechanism is the same one that makes it fast: +> CodeGraph returns one dense, verbatim payload that answers the question and +> then stays in the window, where a grep-and-read agent churns many small results +> that get evicted. Fewer tokens *processed* and a larger persistent *footprint* +> are both real. If you are running long sessions in a small window, budget for +> it. Measured, per-repo: +> [`docs/benchmarks/residual-context-occupancy.md`](docs/benchmarks/residual-context-occupancy.md). + +Three notes on the drafting, if it gets edited: + +1. **No percentages from this campaign's efficiency table appear in it.** "~80% + more resident" is the occupancy ratio between arms, which is the part that + travels across models and hosts; the 24/23/20/84 throughput figures are + sonnet-3-turn-specific and must not go near the README. +2. **It concedes the point rather than framing it as a feature.** That is + deliberate — `CLAUDE.md`'s "honesty in the product is load-bearing" applies to + the README before it applies to a product screen, and a reader who hits #1500 + in their own session and finds the README silent on it trusts nothing else in + the table. +3. **The share numbers (33.7% of a 200k window) are Claude Code's** and do not + transfer to another host, so the draft quotes absolute tokens and the arm + ratio only.