docs: the occupancy baseline says our residual is higher — write that down (CG-13)

Fills the empty RESULTS placeholder with the 2026-08-05 campaign
(bgjob-6d357cd2: 7 repos x 2 arms x 4 runs x 3 turns, 137 min).

The finding is not the flattering one. Retrieval residual is 82% HIGHER
with codegraph and share-of-context 27% higher, on all seven repos —
vscode 67k resident against 18k. At the same time six of seven
without-arms *process* more total tokens (gin 660k vs 290k) while
leaving less behind. Both are true: one dense verbatim payload stays
resident where many small Read/Grep results evict. This corroborates
issue #1500 on our own harness; the aggregator used to print it as
"-82% lower with codegraph" until the sign bug at 520ed9d.

Also:

- States the regime everywhere. This ran claude-sonnet-5 / 3-turn; the
  README's table is Opus 4.8 / single-question. Measured 24/23/20/84
  against the published 60/69/20/89 — model and turn count, not
  contamination. Records the two inversions honestly (vscode processes
  98% more tokens, django costs 17% more) and that 4 of 28 with-arm
  sessions still touched Read.
- Corrects the "Settled" section, which claimed a 7-repo baseline
  existed before one did, and adds the unclaimed Opus rerun.
- Records the contamination gate: 0 CLI calls returned output in 56
  sessions, but 29 attempts were blocked — 26 of 28 without-arm
  sessions tried. no-cli-shim.sh is load-bearing, not precautionary.
- Records the secondary readings as absolute, not before/after: 86.7%
  allocation efficiency pooled over 110 calls, read-of-a-file-we-
  returned 2%, explore-again 73% and ambiguous by construction.

README.md is deliberately untouched — restating its numbers from sonnet
3-turn data would be wrong. A proposed README paragraph is drafted at
the end of the benchmark doc for the maintainer to accept or reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Colby McHenry
2026-08-05 13:06:51 -05:00
parent 520ed9d933
commit 5dd4db68cd
2 changed files with 217 additions and 5 deletions
@@ -54,7 +54,10 @@ table still renders against whichever arm's logs are already in `$AGENT_EVAL_OUT
**A campaign — `bench-readme.sh`.** The 7 README repos, three turns each,
`RUNS` per arm, through `run-all.sh` — so every run in a campaign carries all
three metrics. Aggregate with `parse-bench-readme.mjs`.
three metrics. Aggregate with `parse-bench-readme.mjs`. One has been run:
[the 2026-08-05 baseline](residual-context-occupancy.md#baseline-the-7-readme-repos)
(sonnet, 3 turns, 4 runs/arm) — read its regime box before comparing anything to
it, and note that it is **not** the regime the README's table was published in.
**A log you already have.** `parse-run.mjs <run.jsonl> [run.tN.jsonl …]` prints
the three blocks for any stream-json log; `--brief` drops the numbered call
+213 -4
View File
@@ -168,18 +168,185 @@ without-arm may have been using codegraph.**
## Baseline: the 7 README repos
<!-- RESULTS -->
### The regime — read this before any number below
| | This campaign | README's published table |
|---|---|---|
| Model | **`claude-sonnet-5`** | **Claude Opus 4.8** |
| Session shape | **3 turns** (README question + 2 in-flow follow-ups) | **1 question** |
| Runs | 4 per arm × 7 repos = **56 sessions** | 4 per arm, median |
| Ran | 2026-08-05, 137 min (`bgjob-6d357cd2`), raw under `/tmp/ab-readme` | 2026-07-21 |
**These two are not comparable, and the difference is model + turn count — not
contamination.** Sonnet is the deliberate floor model for this harness
(`CLAUDE.md`: an affordance that lands on Sonnet generalizes up; one that only
works on Opus does not generalize down). Three turns is what makes occupancy
chargeable at all. Both choices move the efficiency numbers, so the throughput
row below reads *lower* than the README's and neither figure invalidates the
other. Settling whether the published Opus figures still hold needs a matched
**Opus 4.8, single-question** rerun; that is deliberately out of scope here.
Reproduce with:
```bash
CORPUS=/tmp/codegraph-corpus scripts/agent-eval/bench-readme.sh # RUNS=4 CG_TURNS=3
node scripts/agent-eval/parse-bench-readme.mjs /tmp/ab-readme
```
### The finding: codegraph's residual is 82% HIGHER, on all seven repos
```
repo turns W→WO final ctx W→WO residual W→WO % of ctx W→WO % of window W→WO
vscode 12.5/44 113k→53k 67k→18k (+276%) 59.7%→36.0% 33.7%→9.0%
excalidraw 9/32 87k→57k 43k→25k (+71%) 49.5%→47.6% 21.5%→12.5%
django 7.5/16.5 60k→51k 18k→10k (+71%) 29.3%→19.9% 8.8%→5.1%
tokio 9/33.5 87k→64k 45k→31k (+45%) 52.1%→50.5% 22.7%→15.7%
okhttp 6/14 61k→59k 20k→16k (+27%) 33.2%→27.6% 10.1%→8.0%
gin 6/15.5 56k→49k 15k→8k (+79%) 26.3%→16.7% 7.3%→4.1%
alamofire 10/31 76k→65k 34k→32k (+7%) 44.7%→50.6% 16.9%→15.8%
AVERAGE: retrieval residual 82% HIGHER with codegraph · share-of-context 27% HIGHER
```
W = codegraph's responses still resident. WO = Read + Grep/Glob + Bash results
still resident. `turns` is median assistant turns per session.
**Seven of seven.** There is no repo where codegraph leaves less behind. On
vscode it leaves **67k tokens resident against the without-arm's 18k** — a third
of a 200k window, gone before turn 4 starts. The only near-tie is Alamofire
(+7%), and it is a tie because that arm's share-of-context is actually *lower*
(44.7% vs 50.6%), not because the residual is small.
**Both things are true at once.** On six of seven repos the without-arm
*processes* far more total tokens than the with-arm — gin 660k vs 290k, okhttp
704k vs 302k — while leaving *less* behind. Throughput and stock are different
quantities and they point opposite ways here:
- codegraph front-loads **one large verbatim payload** (2 explore calls on gin,
each tens of thousands of dense source characters) and that payload **stays
resident** for every turn after it;
- Read/Grep/Bash churn **many small results** (gin: ~6 reads + ~5 bash per run),
most of which are re-derivation the agent then discards, and which evict.
**This corroborates issue [#1500](https://github.com/colbymchenry/codegraph/issues/1500)
on our own harness.** The reporter's complaint was exactly this axis, and until
this campaign we had no measurement that could see it. Note for anyone reading
git history: the aggregator originally printed this as "-82% *lower* with
codegraph" — a sign bug, fixed at `520ed9d`. The honest number is the entire
point of the metric; do not soften it.
**Fixed overhead.** codegraph's tool schema + MCP instructions cost **+546 tok**
of context before any tool is called (median with-arm `ctxBase` minus median
without-arm `ctxBase`, averaged over repos). Paid whether or not the agent ever
calls codegraph. Small — the residual, not the schema, is where the context goes.
### Throughput in the same campaign (sonnet · 3 turns)
Reported for completeness and because the occupancy finding only means anything
read against it. **These are not the README's numbers and must not be quoted as
such.**
```
repo time W→WO tools W→WO tokens W→WO (saved) cost W→WO (saved)
vscode 2m 59s→1m 59s 8→60 949k→478k (-98%) $1.21→$1.62 (25%)
excalidraw 1m 45s→2m 1s 5→44 557k→869k (36%) $0.78→$1.03 (24%)
django 1m 4s→1m 30s 3→13 366k→686k (47%) $0.52→$0.45 (-17%)
tokio 1m 47s→4m 35s 5→43 574k→793k (28%) $0.67→$1.25 (47%)
okhttp 49s→1m 25s 2→11 302k→704k (57%) $0.35→$0.48 (27%)
gin 1m 6s→1m 43s 2→12 290k→660k (56%) $0.38→$0.43 (12%)
alamofire 1m 35s→1m 47s 6→29 545k→870k (37%) $0.67→$1.30 (49%)
AVERAGE saved: cost 24% · tokens 23% · time 20% · tool calls 84%
```
| | this campaign (sonnet, 3-turn) | README (Opus 4.8, 1-question) |
|---|---|---|
| cost saved | **24%** | 60% |
| tokens saved | **23%** | 69% |
| time saved | **20%** | 20% |
| tool calls saved | **84%** | 89% |
Tool-call reduction and wall-clock survive the regime change almost intact; the
cost and token savings roughly halve. Two repos invert outright — **vscode
processes 98% *more* tokens with codegraph** (8 calls of dense source against a
without-arm that mostly greps), and **django costs 17% more**. Neither is hidden
here. The with-arm is also not read-free in this regime: 4 of 28 with-arm
sessions still touched Read (vscode run4 `rd5 bs7`, tokio run2 `rd3 bs2`,
django run4 `rd1`, alamofire run2 `rd1`), against the README's "zero file reads
on all seven repos" under Opus.
### Contamination gate: clean, and the channel is real
**0 CLI calls returned output in any of the 56 sessions.** The aggregate is
uncontaminated and no run was dropped.
But **29 attempts were blocked** — 26 in the without-arm (in **26 of its 28
sessions**) and 3 in the with-arm. Ninety-three percent of without-arm sessions
tried to reach codegraph through Bash and were stopped by the sanitized PATH +
PreToolUse hook (`no-cli-shim.sh`). That is not a hypothetical channel the
harness guards out of caution; it is the agent's *default* move once it notices
`.codegraph/` in the tree. **`no-cli-shim.sh` is load-bearing** — without it this
campaign would have been codegraph-over-CLI vs codegraph-over-MCP, exactly as the
earlier 14-of-15 pass was (see the section above). Check the contamination row
before believing any number from this harness.
### Secondary readings — absolute, not before/after
There is **no baseline-build arm in this campaign** — every number below is the
current build's absolute reading on these questions. Allocation efficiency in
particular is *relative* (attribution is by citation): it compares builds on the
same question and says nothing on its own about waste. For a before/after
allocation A/B see [`explore-allocation-ab-1500.md`](explore-allocation-ab-1500.md).
```
repo calls again read-ret read-miss grep MOVED ON alloc eff envelope
vscode 26 21 81% 0 0% 1 4% 1 4% 3 12% 63.2% 442k
excalidraw 18 14 78% 0 0% 0 0% 0 0% 4 22% 94.7% 338k
django 11 6 55% 1 9% 0 0% 0 0% 4 36% 96.2% 186k
tokio 16 12 75% 0 0% 0 0% 1 6% 3 19% 92.9% 343k
okhttp 8 4 50% 0 0% 0 0% 0 0% 4 50% 97.9% 151k
gin 8 4 50% 0 0% 0 0% 0 0% 4 50% 99.0% 116k
alamofire 23 19 83% 1 4% 0 0% 0 0% 3 13% 89.0% 306k
POOLED (110 answered explore calls):
explore again 73% · Read a file we returned 2% · Read a file we did NOT return 1%
· Grep/Glob 2% · moved on / answered 23%
POOLED allocation efficiency: 86.7% over 110 calls / 1.9M chars
```
- **Allocation efficiency 86.7%** pooled. vscode is the outlier at 63.2% — the
largest envelope (442k chars) and the lowest citation share, which is where an
allocation change would show up first.
- **`Read a file we returned` = 2%** (2 of 110). Right file, wrong bytes is
nearly absent; the allocation misses this metric was built to catch are not
what is driving vscode's number.
- **Recall misses** are 3% total (1 read-miss, 2 grep).
- **`explore again` = 73%** and is **ambiguous by construction** — it is
indistinguishable between "the first call was insufficient" and "the agent is
working through a 3-turn session and this is turn 2's first call." In a 3-turn
regime that ambiguity is much larger than it was single-turn; treat the
high-`again` repos (alamofire 83%, vscode 81%) as unresolved, not as failures.
---
## What this settles, and what it does not
**Settled.** The metric exists, it is measured rather than estimated, it runs over
multi-turn sessions — the regime where occupancy is actually charged — and there
is a baseline across the 7 README repos to compare future changes against.
**Settled.** The metric exists, it is measured rather than estimated, and it runs
over multi-turn sessions — the regime where occupancy is actually charged. As of
2026-08-05 there is a baseline across the 7 README repos (above) to compare
future changes against, and it says codegraph's residual is **higher**, on every
repo. (Before that campaign this section claimed such a baseline existed when it
did not; it does now, and it is one regime — `claude-sonnet-5`, 3 turns — not a
general result.)
**Not settled, and deliberately not claimed:**
- **The README's efficiency figures.** The baseline above ran sonnet / 3 turns;
the README published Opus 4.8 / single-question. The gap between 24/23/20/84
and 60/69/20/89 is regime, not regression, and this campaign cannot tell you
which way the published numbers have moved. That needs a matched **Opus 4.8,
single-question** rerun. Out of scope here, and `README.md` was deliberately
left untouched.
- **A different host.** The reporter was in Cursor. We measure Claude Code.
Window size, system prompt, and compaction policy all differ, so the *share*
numbers do not transfer host to host; the ratio between the arms is the part
@@ -203,3 +370,45 @@ is a baseline across the 7 README repos to compare future changes against.
*enough* — that is [explore sufficiency](explore-sufficiency.md), which every
run now prints alongside this block — nor about how much of the returned bytes
the answer actually used (CG-9).
---
## Proposed README wording — for the maintainer, not applied
`README.md` is **deliberately untouched by this work.** Its benchmark table is
Opus 4.8 / single-question and nothing measured here can restate it. What follows
is a *proposal*: the occupancy finding as an honest counterweight to the
efficiency table, phrased so it does not depend on the sonnet-vs-Opus regime for
its claim. Accept, reject, or rewrite — this is not a pending edit.
Suggested placement: immediately after the "A note on cost" paragraph (README
line ~195), as a second `>` note under the same table.
> **A note on context.** The efficiency table above measures *throughput* —
> tokens processed, tools called, dollars spent to reach one answer. It does not
> measure what is still sitting in the window afterward, and on that axis
> CodeGraph costs more, not less. Across the same seven repos in multi-turn
> sessions, CodeGraph's responses leave **~80% more retrieval context resident**
> at the end of a session than the file-reading agent's do — on VS Code, 67k
> tokens against 18k. The mechanism is the same one that makes it fast:
> CodeGraph returns one dense, verbatim payload that answers the question and
> then stays in the window, where a grep-and-read agent churns many small results
> that get evicted. Fewer tokens *processed* and a larger persistent *footprint*
> are both real. If you are running long sessions in a small window, budget for
> it. Measured, per-repo:
> [`docs/benchmarks/residual-context-occupancy.md`](docs/benchmarks/residual-context-occupancy.md).
Three notes on the drafting, if it gets edited:
1. **No percentages from this campaign's efficiency table appear in it.** "~80%
more resident" is the occupancy ratio between arms, which is the part that
travels across models and hosts; the 24/23/20/84 throughput figures are
sonnet-3-turn-specific and must not go near the README.
2. **It concedes the point rather than framing it as a feature.** That is
deliberate — `CLAUDE.md`'s "honesty in the product is load-bearing" applies to
the README before it applies to a product screen, and a reader who hits #1500
in their own session and finds the README silent on it trusts nothing else in
the table.
3. **The share numbers (33.7% of a 200k window) are Claude Code's** and do not
transfer to another host, so the draft quotes absolute tokens and the arm
ratio only.