docs: the occupancy baseline says our residual is higher — write that down (CG-13)
Fills the empty RESULTS placeholder with the 2026-08-05 campaign
(bgjob-6d357cd2: 7 repos x 2 arms x 4 runs x 3 turns, 137 min).
The finding is not the flattering one. Retrieval residual is 82% HIGHER
with codegraph and share-of-context 27% higher, on all seven repos —
vscode 67k resident against 18k. At the same time six of seven
without-arms *process* more total tokens (gin 660k vs 290k) while
leaving less behind. Both are true: one dense verbatim payload stays
resident where many small Read/Grep results evict. This corroborates
issue #1500 on our own harness; the aggregator used to print it as
"-82% lower with codegraph" until the sign bug at 520ed9d.
Also:
- States the regime everywhere. This ran claude-sonnet-5 / 3-turn; the
README's table is Opus 4.8 / single-question. Measured 24/23/20/84
against the published 60/69/20/89 — model and turn count, not
contamination. Records the two inversions honestly (vscode processes
98% more tokens, django costs 17% more) and that 4 of 28 with-arm
sessions still touched Read.
- Corrects the "Settled" section, which claimed a 7-repo baseline
existed before one did, and adds the unclaimed Opus rerun.
- Records the contamination gate: 0 CLI calls returned output in 56
sessions, but 29 attempts were blocked — 26 of 28 without-arm
sessions tried. no-cli-shim.sh is load-bearing, not precautionary.
- Records the secondary readings as absolute, not before/after: 86.7%
allocation efficiency pooled over 110 calls, read-of-a-file-we-
returned 2%, explore-again 73% and ambiguous by construction.
README.md is deliberately untouched — restating its numbers from sonnet
3-turn data would be wrong. A proposed README paragraph is drafted at
the end of the benchmark doc for the maintainer to accept or reject.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -54,7 +54,10 @@ table still renders against whichever arm's logs are already in `$AGENT_EVAL_OUT
|
||||
|
||||
**A campaign — `bench-readme.sh`.** The 7 README repos, three turns each,
|
||||
`RUNS` per arm, through `run-all.sh` — so every run in a campaign carries all
|
||||
three metrics. Aggregate with `parse-bench-readme.mjs`.
|
||||
three metrics. Aggregate with `parse-bench-readme.mjs`. One has been run:
|
||||
[the 2026-08-05 baseline](residual-context-occupancy.md#baseline-the-7-readme-repos)
|
||||
(sonnet, 3 turns, 4 runs/arm) — read its regime box before comparing anything to
|
||||
it, and note that it is **not** the regime the README's table was published in.
|
||||
|
||||
**A log you already have.** `parse-run.mjs <run.jsonl> [run.tN.jsonl …]` prints
|
||||
the three blocks for any stream-json log; `--brief` drops the numbered call
|
||||
|
||||
@@ -168,18 +168,185 @@ without-arm may have been using codegraph.**
|
||||
|
||||
## Baseline: the 7 README repos
|
||||
|
||||
<!-- RESULTS -->
|
||||
### The regime — read this before any number below
|
||||
|
||||
| | This campaign | README's published table |
|
||||
|---|---|---|
|
||||
| Model | **`claude-sonnet-5`** | **Claude Opus 4.8** |
|
||||
| Session shape | **3 turns** (README question + 2 in-flow follow-ups) | **1 question** |
|
||||
| Runs | 4 per arm × 7 repos = **56 sessions** | 4 per arm, median |
|
||||
| Ran | 2026-08-05, 137 min (`bgjob-6d357cd2`), raw under `/tmp/ab-readme` | 2026-07-21 |
|
||||
|
||||
**These two are not comparable, and the difference is model + turn count — not
|
||||
contamination.** Sonnet is the deliberate floor model for this harness
|
||||
(`CLAUDE.md`: an affordance that lands on Sonnet generalizes up; one that only
|
||||
works on Opus does not generalize down). Three turns is what makes occupancy
|
||||
chargeable at all. Both choices move the efficiency numbers, so the throughput
|
||||
row below reads *lower* than the README's and neither figure invalidates the
|
||||
other. Settling whether the published Opus figures still hold needs a matched
|
||||
**Opus 4.8, single-question** rerun; that is deliberately out of scope here.
|
||||
|
||||
Reproduce with:
|
||||
|
||||
```bash
|
||||
CORPUS=/tmp/codegraph-corpus scripts/agent-eval/bench-readme.sh # RUNS=4 CG_TURNS=3
|
||||
node scripts/agent-eval/parse-bench-readme.mjs /tmp/ab-readme
|
||||
```
|
||||
|
||||
### The finding: codegraph's residual is 82% HIGHER, on all seven repos
|
||||
|
||||
```
|
||||
repo turns W→WO final ctx W→WO residual W→WO % of ctx W→WO % of window W→WO
|
||||
vscode 12.5/44 113k→53k 67k→18k (+276%) 59.7%→36.0% 33.7%→9.0%
|
||||
excalidraw 9/32 87k→57k 43k→25k (+71%) 49.5%→47.6% 21.5%→12.5%
|
||||
django 7.5/16.5 60k→51k 18k→10k (+71%) 29.3%→19.9% 8.8%→5.1%
|
||||
tokio 9/33.5 87k→64k 45k→31k (+45%) 52.1%→50.5% 22.7%→15.7%
|
||||
okhttp 6/14 61k→59k 20k→16k (+27%) 33.2%→27.6% 10.1%→8.0%
|
||||
gin 6/15.5 56k→49k 15k→8k (+79%) 26.3%→16.7% 7.3%→4.1%
|
||||
alamofire 10/31 76k→65k 34k→32k (+7%) 44.7%→50.6% 16.9%→15.8%
|
||||
|
||||
AVERAGE: retrieval residual 82% HIGHER with codegraph · share-of-context 27% HIGHER
|
||||
```
|
||||
|
||||
W = codegraph's responses still resident. WO = Read + Grep/Glob + Bash results
|
||||
still resident. `turns` is median assistant turns per session.
|
||||
|
||||
**Seven of seven.** There is no repo where codegraph leaves less behind. On
|
||||
vscode it leaves **67k tokens resident against the without-arm's 18k** — a third
|
||||
of a 200k window, gone before turn 4 starts. The only near-tie is Alamofire
|
||||
(+7%), and it is a tie because that arm's share-of-context is actually *lower*
|
||||
(44.7% vs 50.6%), not because the residual is small.
|
||||
|
||||
**Both things are true at once.** On six of seven repos the without-arm
|
||||
*processes* far more total tokens than the with-arm — gin 660k vs 290k, okhttp
|
||||
704k vs 302k — while leaving *less* behind. Throughput and stock are different
|
||||
quantities and they point opposite ways here:
|
||||
|
||||
- codegraph front-loads **one large verbatim payload** (2 explore calls on gin,
|
||||
each tens of thousands of dense source characters) and that payload **stays
|
||||
resident** for every turn after it;
|
||||
- Read/Grep/Bash churn **many small results** (gin: ~6 reads + ~5 bash per run),
|
||||
most of which are re-derivation the agent then discards, and which evict.
|
||||
|
||||
**This corroborates issue [#1500](https://github.com/colbymchenry/codegraph/issues/1500)
|
||||
on our own harness.** The reporter's complaint was exactly this axis, and until
|
||||
this campaign we had no measurement that could see it. Note for anyone reading
|
||||
git history: the aggregator originally printed this as "-82% *lower* with
|
||||
codegraph" — a sign bug, fixed at `520ed9d`. The honest number is the entire
|
||||
point of the metric; do not soften it.
|
||||
|
||||
**Fixed overhead.** codegraph's tool schema + MCP instructions cost **+546 tok**
|
||||
of context before any tool is called (median with-arm `ctxBase` minus median
|
||||
without-arm `ctxBase`, averaged over repos). Paid whether or not the agent ever
|
||||
calls codegraph. Small — the residual, not the schema, is where the context goes.
|
||||
|
||||
### Throughput in the same campaign (sonnet · 3 turns)
|
||||
|
||||
Reported for completeness and because the occupancy finding only means anything
|
||||
read against it. **These are not the README's numbers and must not be quoted as
|
||||
such.**
|
||||
|
||||
```
|
||||
repo time W→WO tools W→WO tokens W→WO (saved) cost W→WO (saved)
|
||||
vscode 2m 59s→1m 59s 8→60 949k→478k (-98%) $1.21→$1.62 (25%)
|
||||
excalidraw 1m 45s→2m 1s 5→44 557k→869k (36%) $0.78→$1.03 (24%)
|
||||
django 1m 4s→1m 30s 3→13 366k→686k (47%) $0.52→$0.45 (-17%)
|
||||
tokio 1m 47s→4m 35s 5→43 574k→793k (28%) $0.67→$1.25 (47%)
|
||||
okhttp 49s→1m 25s 2→11 302k→704k (57%) $0.35→$0.48 (27%)
|
||||
gin 1m 6s→1m 43s 2→12 290k→660k (56%) $0.38→$0.43 (12%)
|
||||
alamofire 1m 35s→1m 47s 6→29 545k→870k (37%) $0.67→$1.30 (49%)
|
||||
|
||||
AVERAGE saved: cost 24% · tokens 23% · time 20% · tool calls 84%
|
||||
```
|
||||
|
||||
| | this campaign (sonnet, 3-turn) | README (Opus 4.8, 1-question) |
|
||||
|---|---|---|
|
||||
| cost saved | **24%** | 60% |
|
||||
| tokens saved | **23%** | 69% |
|
||||
| time saved | **20%** | 20% |
|
||||
| tool calls saved | **84%** | 89% |
|
||||
|
||||
Tool-call reduction and wall-clock survive the regime change almost intact; the
|
||||
cost and token savings roughly halve. Two repos invert outright — **vscode
|
||||
processes 98% *more* tokens with codegraph** (8 calls of dense source against a
|
||||
without-arm that mostly greps), and **django costs 17% more**. Neither is hidden
|
||||
here. The with-arm is also not read-free in this regime: 4 of 28 with-arm
|
||||
sessions still touched Read (vscode run4 `rd5 bs7`, tokio run2 `rd3 bs2`,
|
||||
django run4 `rd1`, alamofire run2 `rd1`), against the README's "zero file reads
|
||||
on all seven repos" under Opus.
|
||||
|
||||
### Contamination gate: clean, and the channel is real
|
||||
|
||||
**0 CLI calls returned output in any of the 56 sessions.** The aggregate is
|
||||
uncontaminated and no run was dropped.
|
||||
|
||||
But **29 attempts were blocked** — 26 in the without-arm (in **26 of its 28
|
||||
sessions**) and 3 in the with-arm. Ninety-three percent of without-arm sessions
|
||||
tried to reach codegraph through Bash and were stopped by the sanitized PATH +
|
||||
PreToolUse hook (`no-cli-shim.sh`). That is not a hypothetical channel the
|
||||
harness guards out of caution; it is the agent's *default* move once it notices
|
||||
`.codegraph/` in the tree. **`no-cli-shim.sh` is load-bearing** — without it this
|
||||
campaign would have been codegraph-over-CLI vs codegraph-over-MCP, exactly as the
|
||||
earlier 14-of-15 pass was (see the section above). Check the contamination row
|
||||
before believing any number from this harness.
|
||||
|
||||
### Secondary readings — absolute, not before/after
|
||||
|
||||
There is **no baseline-build arm in this campaign** — every number below is the
|
||||
current build's absolute reading on these questions. Allocation efficiency in
|
||||
particular is *relative* (attribution is by citation): it compares builds on the
|
||||
same question and says nothing on its own about waste. For a before/after
|
||||
allocation A/B see [`explore-allocation-ab-1500.md`](explore-allocation-ab-1500.md).
|
||||
|
||||
```
|
||||
repo calls again read-ret read-miss grep MOVED ON alloc eff envelope
|
||||
vscode 26 21 81% 0 0% 1 4% 1 4% 3 12% 63.2% 442k
|
||||
excalidraw 18 14 78% 0 0% 0 0% 0 0% 4 22% 94.7% 338k
|
||||
django 11 6 55% 1 9% 0 0% 0 0% 4 36% 96.2% 186k
|
||||
tokio 16 12 75% 0 0% 0 0% 1 6% 3 19% 92.9% 343k
|
||||
okhttp 8 4 50% 0 0% 0 0% 0 0% 4 50% 97.9% 151k
|
||||
gin 8 4 50% 0 0% 0 0% 0 0% 4 50% 99.0% 116k
|
||||
alamofire 23 19 83% 1 4% 0 0% 0 0% 3 13% 89.0% 306k
|
||||
|
||||
POOLED (110 answered explore calls):
|
||||
explore again 73% · Read a file we returned 2% · Read a file we did NOT return 1%
|
||||
· Grep/Glob 2% · moved on / answered 23%
|
||||
POOLED allocation efficiency: 86.7% over 110 calls / 1.9M chars
|
||||
```
|
||||
|
||||
- **Allocation efficiency 86.7%** pooled. vscode is the outlier at 63.2% — the
|
||||
largest envelope (442k chars) and the lowest citation share, which is where an
|
||||
allocation change would show up first.
|
||||
- **`Read a file we returned` = 2%** (2 of 110). Right file, wrong bytes is
|
||||
nearly absent; the allocation misses this metric was built to catch are not
|
||||
what is driving vscode's number.
|
||||
- **Recall misses** are 3% total (1 read-miss, 2 grep).
|
||||
- **`explore again` = 73%** and is **ambiguous by construction** — it is
|
||||
indistinguishable between "the first call was insufficient" and "the agent is
|
||||
working through a 3-turn session and this is turn 2's first call." In a 3-turn
|
||||
regime that ambiguity is much larger than it was single-turn; treat the
|
||||
high-`again` repos (alamofire 83%, vscode 81%) as unresolved, not as failures.
|
||||
|
||||
---
|
||||
|
||||
## What this settles, and what it does not
|
||||
|
||||
**Settled.** The metric exists, it is measured rather than estimated, it runs over
|
||||
multi-turn sessions — the regime where occupancy is actually charged — and there
|
||||
is a baseline across the 7 README repos to compare future changes against.
|
||||
**Settled.** The metric exists, it is measured rather than estimated, and it runs
|
||||
over multi-turn sessions — the regime where occupancy is actually charged. As of
|
||||
2026-08-05 there is a baseline across the 7 README repos (above) to compare
|
||||
future changes against, and it says codegraph's residual is **higher**, on every
|
||||
repo. (Before that campaign this section claimed such a baseline existed when it
|
||||
did not; it does now, and it is one regime — `claude-sonnet-5`, 3 turns — not a
|
||||
general result.)
|
||||
|
||||
**Not settled, and deliberately not claimed:**
|
||||
|
||||
- **The README's efficiency figures.** The baseline above ran sonnet / 3 turns;
|
||||
the README published Opus 4.8 / single-question. The gap between 24/23/20/84
|
||||
and 60/69/20/89 is regime, not regression, and this campaign cannot tell you
|
||||
which way the published numbers have moved. That needs a matched **Opus 4.8,
|
||||
single-question** rerun. Out of scope here, and `README.md` was deliberately
|
||||
left untouched.
|
||||
- **A different host.** The reporter was in Cursor. We measure Claude Code.
|
||||
Window size, system prompt, and compaction policy all differ, so the *share*
|
||||
numbers do not transfer host to host; the ratio between the arms is the part
|
||||
@@ -203,3 +370,45 @@ is a baseline across the 7 README repos to compare future changes against.
|
||||
*enough* — that is [explore sufficiency](explore-sufficiency.md), which every
|
||||
run now prints alongside this block — nor about how much of the returned bytes
|
||||
the answer actually used (CG-9).
|
||||
|
||||
---
|
||||
|
||||
## Proposed README wording — for the maintainer, not applied
|
||||
|
||||
`README.md` is **deliberately untouched by this work.** Its benchmark table is
|
||||
Opus 4.8 / single-question and nothing measured here can restate it. What follows
|
||||
is a *proposal*: the occupancy finding as an honest counterweight to the
|
||||
efficiency table, phrased so it does not depend on the sonnet-vs-Opus regime for
|
||||
its claim. Accept, reject, or rewrite — this is not a pending edit.
|
||||
|
||||
Suggested placement: immediately after the "A note on cost" paragraph (README
|
||||
line ~195), as a second `>` note under the same table.
|
||||
|
||||
> **A note on context.** The efficiency table above measures *throughput* —
|
||||
> tokens processed, tools called, dollars spent to reach one answer. It does not
|
||||
> measure what is still sitting in the window afterward, and on that axis
|
||||
> CodeGraph costs more, not less. Across the same seven repos in multi-turn
|
||||
> sessions, CodeGraph's responses leave **~80% more retrieval context resident**
|
||||
> at the end of a session than the file-reading agent's do — on VS Code, 67k
|
||||
> tokens against 18k. The mechanism is the same one that makes it fast:
|
||||
> CodeGraph returns one dense, verbatim payload that answers the question and
|
||||
> then stays in the window, where a grep-and-read agent churns many small results
|
||||
> that get evicted. Fewer tokens *processed* and a larger persistent *footprint*
|
||||
> are both real. If you are running long sessions in a small window, budget for
|
||||
> it. Measured, per-repo:
|
||||
> [`docs/benchmarks/residual-context-occupancy.md`](docs/benchmarks/residual-context-occupancy.md).
|
||||
|
||||
Three notes on the drafting, if it gets edited:
|
||||
|
||||
1. **No percentages from this campaign's efficiency table appear in it.** "~80%
|
||||
more resident" is the occupancy ratio between arms, which is the part that
|
||||
travels across models and hosts; the 24/23/20/84 throughput figures are
|
||||
sonnet-3-turn-specific and must not go near the README.
|
||||
2. **It concedes the point rather than framing it as a feature.** That is
|
||||
deliberate — `CLAUDE.md`'s "honesty in the product is load-bearing" applies to
|
||||
the README before it applies to a product screen, and a reader who hits #1500
|
||||
in their own session and finds the README silent on it trusts nothing else in
|
||||
the table.
|
||||
3. **The share numbers (33.7% of a 200k window) are Claude Code's** and do not
|
||||
transfer to another host, so the draft quotes absolute tokens and the arm
|
||||
ratio only.
|
||||
|
||||
Reference in New Issue
Block a user