Files
Rohit Ghumare 857f71e3c6 feat: v0.6.0 advanced retrieval with real-world benchmarks (#76)
* feat: add 9 orchestration modules for v0.5.0

Add actions, frontier, leases, routines, signals, checkpoints,
flow-compress, mesh, and branch-aware modules with full MCP tools,
REST endpoints, and 170 new tests. Includes SSRF protection,
race-condition-safe keyed mutex locking, SHA-256 fingerprinting,
and fixes for 29 CodeRabbit review findings.

- 9 new source files (src/functions/*)
- 8 new test files (170 tests, total 386)
- 10 new MCP tools, 23 new REST endpoints
- 8 new KV scopes, new types for orchestration
- Version bump to 0.5.0, README and viewer updated

* feat: add sentinels, sketches, crystallize, diagnostics, facets modules

5 new modules inspired by beads patterns but with original naming and
iii-engine real-time streaming (no polling):

- sentinels: event-driven condition watchers (webhook, timer, threshold,
  pattern, approval) that auto-unblock gated actions via SSE
- sketches: ephemeral action graphs with auto-expiry, promote or discard
- crystallize: LLM-powered compaction of completed action chains into
  compact crystal digests with key outcomes and lessons
- diagnostics: self-diagnosis across 8 categories (actions, leases,
  sentinels, sketches, signals, sessions, memories, mesh) with auto-heal
- facets: multi-dimensional tagging (dimension:value) with AND/OR queries

- 5 new source files, 5 new test files (132 tests, total 518)
- 9 new MCP tools (total 37), 21 new REST endpoints (total 93)
- 4 new KV scopes (total 33), new types for all modules
- README stats and function table updated

* fix: address code review findings across v0.5.0 modules

- actions: validate edges before persisting, set blocked status for requires deps
- leases: use mem:action lock key, reject blocked actions, check expiry on release
- checkpoints: validate linkedActionIds exist, check requires edges in unblock
- mesh: add SSRF validation on peer registration
- routines: remove invalid "failed" action status check
- export-import: add v0.5.0 scope export/import (actions, sentinels, sketches, etc)
- mcp/server: validate CSV inputs are strings before splitting
- schema: replace runtime require with static import
- README: fix stale tool/endpoint counts (28→37, 72→93)

* fix: second round code review — mesh locks, routines DAG, MCP input validation, README counts

- mesh.ts: add withKeyedLock on action writes in receive path, add IPv6 private ranges to SSRF check
- routines.ts: validate DAG (duplicate orders, unknown deps), set dep actions to blocked, refresh stepStatus in routine-status
- checkpoints.ts: runtime type enum validation, set linked pending actions to blocked
- leases.ts: validate ttlMs is finite positive number
- export-import.ts: add skip strategy checks for all v0.5.0 import blocks
- mcp/server.ts: typeof guards on tags, config JSON.parse, actionIds, linkedActionIds, categories
- README.md: Tools 18→37, Functions 33→50, stats line updated
- test: update checkpoint test for new blocked-on-create behavior

* fix: third round review — lease renew/release safety, mesh locking, export replace cleanup, MCP input guards

- leases.ts: renew extends from max(now, existing expiry) instead of now; release verifies action ownership before mutation
- mesh.ts: withKeyedLock on memory writes in mesh-receive; applySyncData validates id/updatedAt and locks both memory and action writes
- routines.ts: stepStatus maps "blocked" to "pending" explicitly; progress includes blocked/cancelled counts; routine-freeze wrapped in withKeyedLock
- export-import.ts: remove unused RoutineRun import; replace strategy clears all orchestration namespaces; skip strategy for graphNodes/graphEdges/semantic/procedural
- mcp/server.ts: sentinel_trigger JSON.parse with typeof+try/catch; facet_query typeof guards on matchAll/matchAny; remove redundant requires cast; .filter(Boolean) on concepts/files/tags/requires CSV splits
- README.md: clarify API table is a representative subset

* fix: fourth round review — redirect SSRF, blocked-on-create, missing action handling, boolean normalization

- mesh.ts: add redirect:"error" to both outbound fetch calls to prevent SSRF via redirect
- routines.ts: create actions with status "blocked" directly when hasDeps (eliminates two-pass race); handle missing actions in routine-status as cancelled; progress.total uses run.actionIds.length
- mcp/server.ts: sentinel config accepts object values directly; normalize unreadOnly/dryRun for both JSON booleans and string values
- README.md: consistent bundle size (365KB) across both occurrences

* feat: v0.6.0 advanced retrieval — triple-stream search, stemming, real benchmarks

Search improvements:
- Porter stemmer for word normalization (authentication ↔ authenticating)
- 40+ coding-domain synonym groups (db ↔ database, k8s ↔ kubernetes)
- Binary-search prefix matching replaces O(n) full scan
- Session diversification (max 3 results per session)
- Co-occurrence graph edges between all concept pairs

New retrieval modules:
- Sliding window inference pipeline (context enrichment at ingestion)
- Adaptive query expansion (LLM-generated reformulations)
- Triple-stream search (BM25 + Vector + Graph with RRF fusion)
- Append-only temporal knowledge graph (versioned edges, point-in-time queries)
- Graph-augmented retrieval (entity search + neighborhood expansion)
- Ebbinghaus retention scoring (decay + tiered hot/warm/cold/evictable)

Real-world benchmarks (240 observations, 20 labeled queries):
- Quality eval: 64.1% recall@10 with Xenova embeddings (vs 55.8% grep)
- Scale eval: 92-100% token savings vs built-in memory at 240-50K observations
- Cross-session: 12/12 queries found vs 10/12 for 200-line MEMORY.md cap
- Token measurement uses actual search results (fixed fake constant bug)
- Removed old microbenchmarks (bench.ts, run-bench.ts, COMPARISON.md)
2026-03-18 08:47:35 +00:00

5.6 KiB

agentmemory v0.6.0 — Scale & Cross-Session Evaluation

Date: 2026-03-18T07:45:03.529Z Platform: darwin arm64, Node v20.20.0

1. Scale: agentmemory vs Built-in Memory

Every built-in agent memory (CLAUDE.md, .cursorrules, Cline's memory-bank) loads ALL memory into context every session. agentmemory searches and returns only relevant results.

Observations Sessions Index Build BM25 Search Hybrid Search Heap Context Tokens (built-in) Context Tokens (agentmemory) Savings Built-in Unreachable
240 30 177ms 0.112ms 0.63ms 9MB 10,504 1,924 82% 17%
1,000 125 155ms 0.317ms 1.709ms 6MB 43,834 1,969 96% 80%
5,000 625 810ms 1.496ms 8.58ms 25MB 220,335 1,972 99% 96%
10,000 1250 1657ms 3.195ms 17.49ms 1MB 440,973 1,974 100% 98%
50,000 6250 9182ms 22.827ms 108.722ms 316MB 2,216,173 1,981 100% 100%

What the numbers mean

Context Tokens (built-in): How many tokens Claude Code/Cursor/Cline would consume loading ALL memory into the context window. At 5,000 observations, this is ~250K tokens — exceeding most context windows entirely.

Context Tokens (agentmemory): How many tokens the top-10 search results consume. Stays constant regardless of corpus size.

Built-in Unreachable: Percentage of memories that built-in systems CANNOT access because they exceed the 200-line MEMORY.md cap or context window limits. At 1,000 observations, 80% of your project history is invisible.

Storage Costs

Observations BM25 Index Vector Index (d=384) Total Storage
240 395 KB 494 KB 0.9 MB
1,000 1,599 KB 2,060 KB 3.6 MB
5,000 8,006 KB 10,298 KB 17.9 MB
10,000 16,005 KB 20,596 KB 35.7 MB
50,000 80,126 KB 102,979 KB 178.8 MB

2. Cross-Session Retrieval

Can the system find relevant information from past sessions? This is impossible for built-in memory once observations exceed the line/context cap.

Query Target Session Gap BM25 Found BM25 Rank Hybrid Found Hybrid Rank Built-in Visible
How did we set up OAuth providers? ses_005-009 24 Yes #1 Yes #1 Yes
What was the N+1 query fix? ses_010-014 18 Yes #1 Yes #2 Yes
PostgreSQL full-text search setup ses_010-014 17 Yes #1 Yes #1 Yes
bcrypt password hashing configuration ses_005-009 20 Yes #1 Yes #1 Yes
Vitest unit testing setup ses_020-024 9 Yes #1 Yes #1 Yes
webhook retry exponential backoff ses_015-019 14 Yes #1 Yes #1 Yes
ESLint flat config migration ses_000-004 29 Yes #1 Yes #1 Yes
Kubernetes HPA autoscaling configuration ses_025-029 4 Yes #1 Yes #1 No
Prisma database seed script ses_010-014 16 Yes #1 Yes #1 Yes
API cursor-based pagination ses_015-019 14 Yes #1 Yes #1 Yes
CSRF protection double-submit cookie ses_005-009 24 Yes #1 Yes #1 Yes
blue-green deployment rollback ses_025-029 4 Yes #1 Yes #1 No

Summary: agentmemory BM25 found 12/12 cross-session queries. Hybrid found 12/12. Built-in memory (200-line cap) could only reach 10/12.

3. The Context Window Problem

Agent context window: ~200K tokens
System prompt + tools:  ~20K tokens
User conversation:      ~30K tokens
Available for memory:  ~150K tokens

At 50 tokens/observation:
  200 observations  =  10,000 tokens  (fits, but 200-line cap hits first)
  1,000 observations =  50,000 tokens  (33% of available budget)
  5,000 observations = 250,000 tokens  (EXCEEDS total context window)

agentmemory top-10 results:
  Any corpus size     =  ~1,924 tokens  (0.3% of budget)

4. What Built-in Memory Cannot Do

Capability Built-in (CLAUDE.md) agentmemory
Semantic search No (keyword grep only) BM25 + vector + graph
Scale beyond 200 lines No (hard cap) Unlimited
Cross-session recall Only if in 200-line window Full corpus search
Cross-agent sharing No (per-agent files) MCP + REST API
Multi-agent coordination No Leases, signals, actions
Temporal queries No Point-in-time graph
Memory lifecycle No (manual pruning) Ebbinghaus decay + eviction
Knowledge graph No Entity extraction + traversal
Query expansion No LLM-generated reformulations
Retention scoring No Time-frequency decay model
Real-time dashboard No (read files manually) Viewer on :3113
Concurrent access No (file lock) Keyed mutex + KV store

5. When to Use What

Use built-in memory (CLAUDE.md) when:

  • You have < 200 items to remember
  • Single agent, single project
  • Preferences and quick facts only
  • Zero setup is the priority

Use agentmemory when:

  • Project history exceeds 200 observations
  • You need to recall specific incidents from weeks ago
  • Multiple agents work on the same codebase
  • You want semantic search ("how does auth work?") not just keyword matching
  • You need to track memory quality, decay, and lifecycle
  • You want a shared memory layer across Claude Code, Cursor, Windsurf, etc.

Built-in memory is your sticky notes. agentmemory is the searchable database behind them.


Scale tests: 5 corpus sizes. Cross-session tests: 12 queries targeting specific past sessions.