* feat(eval): pluggable benchmark harness with in-house coding-agent corpus Adds eval/ tree (outside files field so npm tarball stays thin) with Adapter interface, three reference adapters (grep / vector / agentmemory-hybrid), two benchmarks (LongMemEval _s public, coding-agent-life-v1 in-house 15 sessions), scoring (P@K, R@K, hit, top-gold-rank), NDJSON output, sandbox script. coding-agent-life-v1 published scorecard at docs/benchmarks/2026-05-20-coding-agent-life-v1.md: agentmemory-hybrid R@5=0.967 P@5=0.578 (100% hit) vs grep R@5=0.967 P@5=0.267. 2.2x better precision on identical input, sandbox-reproducible. Adapter contract: init(sessions, config) -> State; query(q, state, k) -> RankedDoc[] npm scripts: npm run eval:coding-life (no download, no API key for grep) npm run eval:longmemeval (needs OPENAI key + 278MB download) eval/scripts/sandbox.sh boots clean agentmemory + iii-engine on ports 3411/3412 with isolated data dir; tears down on exit. README headline updated. 1072/1072 tests pass + 5 new eval tests. * fix(eval): address review findings on benchmark harness - agentmemory adapter: prefer row.sessionId before observationToSession lookup - vector adapter: validate embedBatch response (length, indexes, non-empty rows) - coding-life: positive-int guard on --k; wrap query loop in try/finally so teardown runs - longmemeval: positive-int guards on --k/--limit/--stratify; per-question try/finally - load: throw on haystack_session_ids vs haystack_sessions length mismatch - score: P@K denominator is k (requested cutoff) not topK.length - sandbox.sh: guard rm -rf with non-empty + /tmp/ prefix check - README: drop unsafe rm "$(which iii)"; instruct ~/.local/bin + PATH instead; add language tag to repo-layout fenced block - sessions.json: fix "two-phase" -> "three-phase" wording mismatch
agentmemory-evals
Public benchmarks for agentmemory's hybrid memory stack (BM25 + embeddings + consolidation + graph).
Two families, both reproducible:
- LongMemEval — public 500-question retrieval benchmark over multi-session chat
- coding-agent-life-v1 — in-house corpus of 15 fictional Claude Code sessions for a Rust CLI project (
shipctl), with 15 hand-graded queries covering bug fixes, refactors, preferences, and multi-session causal reasoning
Adapters
| Adapter | Backend | API key needed |
|---|---|---|
grep |
Tokenized substring match | none |
vector |
OpenAI text-embedding-3-small + cosine |
OPENAI_API_KEY |
agentmemory |
Running agentmemory server, smart-search endpoint | none (auth optional via AGENTMEMORY_SECRET) |
Sandbox first
Running the agentmemory adapter against your real ~/.agentmemory directory pollutes the eval with pre-existing memories AND pollutes your real store with eval test data. Always sandbox.
eval/scripts/sandbox.sh spins up a clean agentmemory + iii-engine on ports 3411/3412 with state in /tmp/agentmemory-eval-sandbox/, exports AGENTMEMORY_BASE_URL, and tears down on exit.
source eval/scripts/sandbox.sh
npm run eval:coding-life -- --adapters grep,agentmemory
Requires iii v0.11.2 on PATH (agentmemory pin). If you already have a different version installed, install the pinned build into ~/.local/bin and make sure that directory comes first on PATH:
mkdir -p ~/.local/bin
curl -fsSL https://github.com/iii-hq/iii/releases/download/iii/v0.11.2/iii-aarch64-apple-darwin.tar.gz | tar -xz -C ~/.local/bin
export PATH="$HOME/.local/bin:$PATH" # add to ~/.zshrc or ~/.bashrc for persistence
Quickstart
coding-agent-life-v1 (in-house, no download)
# grep baseline, no sandbox needed
npm run eval:coding-life -- --adapters grep
# add agentmemory + vector (sandbox + OpenAI key)
source eval/scripts/sandbox.sh
OPENAI_API_KEY=sk-... npm run eval:coding-life -- --adapters grep,vector,agentmemory
LongMemEval _s (public, 278MB download)
mkdir -p ~/datasets/longmemeval
curl -Lo ~/datasets/longmemeval/longmemeval_s.json \
https://huggingface.co/datasets/xiaowu0162/longmemeval/resolve/main/longmemeval_s
source eval/scripts/sandbox.sh
# Stratified sample of 10 per type (fast iteration, ~$0.20 OpenAI cost)
OPENAI_API_KEY=sk-... LONGMEMEVAL_PATH=~/datasets/longmemeval/longmemeval_s.json \
npm run eval:longmemeval -- --stratify 10
# Full 500 questions × 3 adapters (~$2 OpenAI cost)
OPENAI_API_KEY=sk-... LONGMEMEVAL_PATH=~/datasets/longmemeval/longmemeval_s.json \
npm run eval:longmemeval
Repo layout
eval/
├── README.md
├── runner/
│ ├── types.ts Adapter, Question, RankedDoc, ScoreRow
│ ├── score.ts P@K, R@K, aggregation
│ ├── load.ts LongMemEval JSON → Question[]
│ ├── adapters/
│ │ ├── grep.ts tokenized substring baseline
│ │ ├── vector.ts OpenAI embeddings + cosine
│ │ └── agentmemory.ts POST /agentmemory/{remember,smart-search}
│ ├── longmemeval.ts public benchmark runner
│ └── coding-life.ts in-house benchmark runner
└── data/
└── coding-agent-life-v1/
├── sessions.json 15 fictional sessions (~6KB)
└── queries.json 15 queries with gold session IDs
Reports land in eval/reports/<bench>/ (gitignored): scores.ndjson + summary.json.
Published scorecards land in docs/benchmarks/YYYY-MM-DD-<bench>.md.
Writing a new adapter
- Implement
Adapter<State>fromeval/runner/types.ts:import type { Adapter } from "../types.js"; export const myAdapter: Adapter<MyState> = { name: "my-adapter", async init(sessions, config) { /* index */ return state; }, async query(q, state, k) { /* search */ return ranked; }, }; - Register in
eval/runner/{longmemeval,coding-life}.tsADAPTERSmap. - Run against
coding-agent-life-v1to sanity-check before committing OpenAI spend on LongMemEval.
Why a benchmark for agentmemory
agentmemory ships BM25 + embeddings + consolidation + graph retrieval. Numbers from those layers should be measured against grep/vector baselines so the value of each layer is provable.
The in-house corpus is small on purpose (15 sessions) — covers single-session, multi-session, preference, and temporal question types without taking 15 minutes to run. LongMemEval gives the public-comparison axis.