Day zero

· vshulcz ·

16 min read Original article ↗

what each memory tool knows the minute you install it

I maintain deja-vu, which is one of the eight tools measured below. Read the rest with that in mind. The corpus, every driver and the scoring rule are in the repository, so a row you doubt can be re-run.

Memory tools get measured warm. You install one, it fills up for a few weeks, and then you ask it about what it captured. Day zero is the other case. The laptop already holds months of Claude Code, Codex and Cursor sessions, and the tool went on five minutes ago. Can it answer from what is already on disk, and how long before it can? And what did installing it cost, counting the processes it leaves running and the tokens it spends on every turn before you have asked anything?

Two answers surprised me. The vector index with a reranker, built by Hugging Face, spends 31 minutes embedding this history and then ranks below plain BM25 at every depth. And the cost nobody puts in a README: a wired-in MCP server spends between 477 and 8,283 tokens of your context on tool definitions every single turn, whether the agent calls it or not.

There is a second measurement further down, on one real task in a real repository rather than on questions: the same job cost 53,558 tokens with deja wired in and 126,222 with no memory at all, and 52,815 against 103,443 when it was run again on a later build.

The setup

19,195 sessions from the LongMemEval-S cleaned set go on disk first, in the layouts the tools actually read: ~/.claude/projects/<project>/<id>.jsonl and ~/.codex/sessions/<date>/rollout-<id>.jsonl. Nothing is installed until they are there.

The obvious objection first: those are not coding sessions. LongMemEval's conversations are personal-assistant chat about travel plans and purchases. I used them anyway because the benchmark guarantees something a real laptop cannot, which is that each of the 100 questions has exactly one session holding its answer. Without that there is no ground truth to score against. The file layouts are real, the volume is realistic, and the retrieval problem is the same shape. The content is not, and that is the one thing you should discount.

A tool scores at rank k when the session holding the answer comes back k-th. hit@1 counts rank 0, hit@5 counts anything under 5, found@50 counts anything in the first fifty. The control points deja at an empty directory and scores 0 of 100, which is what any memory that only records forward sees on day zero.

Expect low absolute numbers. These are LongMemEval questions asked against nineteen thousand sessions instead of the fifty or so the benchmark ships per question. On that standard setup deja's hit@1 is 88.1%, and the benchmarks page runs it. The haystack here is roughly four hundred times bigger. A third of the questions are the kinds any lexical index is bad at, like stated preferences and facts assembled from several sessions. Nobody looks good here. The interesting part is the gaps between the rows and what each row cost to produce.

Each tool runs with its defaults and its documented way of reading existing history: deja index, cass index --full, agentsview sync, agentmemory import-jsonl, mempalace mine --mode convos. Latency is the wall time of one search call from a warm process; first is the very first call after the build. claude-mem has no row: it captures at session end through hooks and needs an AI provider to write memory, so on day zero it has nothing to answer from.

deja-vufunesctxCASSagentsviewagentmemoryMemPalaceclaude-mem
Reads the history already on disk✓ 35 agents, no capture step✓ 4 harnesses✓ 39 providers✓ 25+ agents✓ 62 agents✓ Claude Code JSONL import, ≤1000 files per call✓ mine per file tree— records from install on
Install18.6 MB binary172 MB binary + 1.2 GB of models on first use43 MB binary58.5 MB binary115 MB binary689 MB npm + 28 MB engine312 MB + 168 MB embedding modelnpm + Claude Code plugin
Runs in the backgroundnothingnothingdaemonnothingdaemon, for searchworker + engine, four portsChromaworker
Needs a model or keynolocal embeddings + rerankerno by default; optional local semantic modelnono; optional embeddings endpointkeyless mode: BM25 onlylocal embeddingsyes
Tool definitions in context, every turn477 tokens, 1 tool2,289, 6 tools1,897, 6 toolsnone — CLI, no MCP server ‖8,283, 7 tools745, 7 tools8,183, 45 tools2,307, 15 tools
Build over 19,195 sessions17.6 s31 min ‡72 s §56 min3 min 56 s95 s≈3 h—
First answer after the build26 ms6.4 s1.04 s0.38 s1.44 s185 ms4.5 s—
Search latency, p5097 ms8.6 s1.03 s0.37 s113 ms145 ms2.6 s—
hit@1 / 100193 (7 with recency off)75 †7 ¶14140
hit@5 / 1003615 (17)148 †13 ¶34200
found@50 / 1006542 (42)3416 †19 ¶65460
Serves the agent by itself✓ MCP, hooks at session start and before a tool runs✓ MCP + per-turn hooks✓ MCP, integrationsrobot CLI, MCPMCP, read-only✓ MCP, hooks✓ MCP, hooks✓ hooks

‖ Tool definitions are sent with every request for as long as a server is wired in, so they are paid once a turn whether or not the agent calls them — the one cost on this page a user pays without asking a question. Counted with scripts/day0compare/toolcost.py: tools/list over the server's own stdio command, tiktoken o200k_base over the JSON of the tools array, each server given an empty home so it answers for its defaults. Measured September 24, 2026, and so on the releases current that day rather than the ones the accuracy rows were run on: deja 0.21.1 (deja mcp), funes 1.3.3, ctx 1.4.12, agentsview 0.44.0, agentmemory 0.9.29, MemPalace 3.9.0 and claude-mem 13.25.3. CASS 0.8.0 has no MCP server — its agent surface is the CLI, which stands in no context and is paid per call instead. The number moves with a release: deja's own was 828 until the schema was cut to one tool with modes.

† CASS is a keyword search: the question as typed returns nothing (a strict AND over every word, stop words included), so its row is the question with stop words removed — one of two tools, with agentsview, given anything but the question verbatim. The released 0.7.1 fails every query on this index with Quill query fuel exhausted (their #441, fixed on main after this was measured); the row is the main build. It also indexed both layouts as separate conversations, 38,394 in all.

¶ agentsview 0.43.0 (kenn-io, Go, SQLite, 62 agents in its table), synced with agentsview sync: 3 min 56 s, all 19,195 sessions — it reads both layouts, 38,390 in its database, 398,998 messages, a 915 MB store. Search runs through its daemon, so the driver starts one after the sync. Its default search returns nothing for all 100 questions as typed, and --fts returns nothing for 88 of them, so like CASS the row is --fts with stop words removed; including one-shot and automated sessions changes none of the three counts. Results are messages; the driver ranks the sessions they belong to. Its semantic and hybrid modes need an embeddings endpoint and are not in the row. Measured September 21, 2026.

§ ctx 1.4.12 (ctx.rs, Rust, Tantivy, lexical by default, 39 providers): its own ctx setup publishes no index for this directory. 19,195 sessions on disk, "Sessions 0, Data scanned 0 B" after 25 minutes, with the daemon and without it. So the row is ctx import --provider claude --path …, which took 72 s and indexed all 19,195 sessions (199,499 events, a 553 MB store). Scored with ctx search --limit 50 --format json on provider_session_id. 1.4 added a semantic mode, off by default; enabled here with its own daemon it embedded 0 of 99,046 records in 1 h 38 min and its job file reports "status": "budget_exhausted", so the row is the lexical index it actually serves. Measured September 21, 2026.

Re-run on ctx 1.6.3 on September 24, 2026, because two minor versions had landed and its site now says setup discovers and reads these sources on its own. ctx setup --wait finds the history and indexes none of it: "Roots 19 included roots, Sessions 0, Messages 0, Data 0 B processed", twice from a clean data root. The explicit import does work on 1.6.3 and indexes everything (19,195 sessions, 199,499 messages, 216.5 MiB), but it takes long enough that the command gives up on its own daemon first — it exits non-zero with admission_pending with no observable progress for 300 seconds while the refresh goes on to finish in the background. The accuracy figures in the row are still the 1.4.12 ones; what was re-measured is the setup path.

‡ funes 1.3.2 (Hugging Face, Rust, local bge-small embeddings and a bge-reranker, four harnesses), indexed with funes index <path> --yes --no-thinking on the same laptop: 52,371 chunks in 31 min, all 19,195 sessions, none skipped. It computed 311,583 chunks and put 52,371 of them in the index, 16.8%, counted from its own per-session lines. On a machine with no index the first funes recall downloads 1.2 GB of models and then stops with no index yet; the budgeted funes index run takes 933 of the 19,195 sessions, offers to finish the rest in about an hour, and answers 0 of 100 until it does. The accuracy row is with its defaults; the bracketed figures are --half-life 0, since the default 30-day recency weighting costs it hits on history this old. Measured September 21, 2026.

Versions and dates: deja main at 9906e203, CASS 0.7.1 and main (38f0412, built from source), agentmemory @agentmemory/agentmemory 0.9.29 from npm on Sep 2, MemPalace 3.9.0 from uv on Sep 2, measured September 2–3, 2026 on an Apple Silicon laptop; the deja, funes and ctx rows re-measured there, and agentsview 0.43.0 added, on September 21, 2026. The deja row was run again on 0.21.1 on September 23, 2026 and came back the same: build 20.4 s, first answer 26 ms, p50 103 ms, hit@1 19, hit@5 36, found@50 65. The deja row is go run ./scripts/day0bench -data longmemeval_s_cleaned.json -limit 100 -corpus 500 -keep DIR; the other rows are the drivers under scripts/day0compare run over DIR.

Reading the numbers

agentmemory ties deja on the measure that matters most here: found@50 is 65 for both. It is two behind at rank five and five behind at rank one. Both are plain BM25 with no model, and the honest summary is that they retrieve about as well as each other. What differs is the bill. deja got there in 17.6 seconds with nothing else running; agentmemory took 95 seconds and leaves a worker and an engine up on four ports, out of a 689 MB npm install.

funes is the result I did not expect. Hugging Face built it on local bge-small embeddings with a bge-reranker, which is the architecture everyone assumes wins. On this pile it does not rank ahead of deja at any depth. With its 30-day recency weighting off, which this old a history needs, it finds the answer in the top five for 17 questions against 36. It charges 31 minutes of embedding and 8.6 seconds a query for that. It also indexes 52,371 of the 311,583 chunks it computes, 16.8%, by its own accounting. On a fresh machine the first funes recall pulls down 1.2 GB of models and then tells you there is no index yet; the budgeted run it offers instead covers 933 of the 19,195 sessions and answers none of the hundred until it finishes.

ctx is the closest thing to deja in design: a Rust lexical index over thirty-nine providers' files, with MCP and a blame command of its own. It finds a little over a third of deja's hits at rank one and rank five, and about half at fifty, with queries ten times slower. The part worth reporting is that its automatic setup never indexed this directory at all. The semantic mode it added embedded nothing here either.

MemPalace pays for being a semantic store twice on day zero: an afternoon of mining with a model to build, and 2.6 seconds a query afterwards, for 46 found against 65.

CASS takes 56 minutes to index, longer than everything except MemPalace, and the released 0.7.1 could not answer at all on this index. The row is a build from their main branch. It is a keyword tool, so it wants the question rewritten before it returns anything.

agentsview is an archive with a web UI over sixty-two agents' files, and as a search over this pile it is quick to build: 3 minutes 56 seconds. Cut the question down to its keywords and it finds the answer at rank one for 7 questions and within fifty for 19, against 19 and 65.

claude-mem scores zero, and that is not a criticism of it. It starts recording at the first session after you install it, which is a deliberate design and a perfectly reasonable one. It just means day zero is empty.

Two of the rows were given help the others were not. CASS and agentsview both return nothing for the questions as typed, so their rows use the question with stop words removed. That is a concession in their favour, and it is the only place any tool got different input.

CASS and deja want the same thing at the start and different things at the end. Both index every agent's files with no model. Then CASS gives you a TUI and a robot interface to search with, and deja hands the session to the agent without being asked, at session start and before a file is edited. If the job is reading your own history, CASS is better at it than we are.

What finishing one piece of work costs

Retrieval scores are a proxy. The thing you actually care about is whether an agent finishes the job faster because the answer was already on the machine. So: same repository, same task, same model, and the only difference is which memory is wired in.

The stand is in scripts/taskcost. The repository is 329 files of Go, and the answer to the task sits in two places nothing points at: a build tag and an environment variable. There is a document in the repo that names the wrong variable, which is the trap. Two prior sessions on the machine did this work already and recorded the command that passed.

The task, given verbatim to each arm: "find the exact command that runs the golden test and makes it pass, without modifying any file, and show the proof". Same repository, same model (opencode 1.18.32, openai/gpt-5.6-luna), same history on disk. The only variable is which memory is wired, each one installed the way its own installer does it. Eleven runs per arm. Every run solved the task and none of them modified the tree.

tool callsmedian tokensmeancache readwall
nothing wired17.9126,222119,14087,64567 s
agentmemory 0.9.2915.9104,974124,96090,95070 s
deja-vu9.853,55859,16735,70045 s

The comparison worth stating out loud, because the table has it but does not say it: deja took the median run from 126,222 tokens to 53,558, which is 57.6% off, and 17.9 tool calls to 9.8. agentmemory brought the median down to 104,974 and pushed the mean up, to 124,960 against 119,140 with no memory wired at all, and it finished 4.5% slower. A memory tool is a thing you add to the context window, and adding it is not free. On the mean, in this stand, one of the two did not earn its keep.

The spread inside an arm is wide. deja's eleven runs went from 24k to 92k tokens, and what decides is whether the agent reads a file before it starts globbing, because reading a file is when the hook fires and the history reaches it. The ordering held anyway: deja's most expensive run of the eleven still cost less than the median run of either other arm.

This stand has been run three times and the saving is a range, not a figure. The table above is eleven runs an arm, and it puts the median at 53,558 tokens against 126,222, which is 58% off. A second run of the same fixture, after the point-of-action line for opencode landed in 0.21.2, came out at 40k against 136k with nothing wired, 71% off, on fewer than eleven runs. The third is eleven runs an arm again, on deja main at 54f6f6ab (after 0.21.2, unreleased when measured), same model and harness, on September 25, 2026, and this time the two arms alternate run by run:

tool callsmedian tokensmeancache readwall
nothing wired17.8103,443107,62478,47662 s
deja-vu7.152,81548,59032,62843 s

That is 49% off the median and 55% off the mean, every run of both arms solved, none modified the tree. Alternating matters more than I expected. The same afternoon the deja arm's own median moved between 24k and 44k from one hour to the next, with nothing changed on this side, and a sequential pair measured that morning — all deja runs first, then all runs with nothing wired — came out at 80% off. That number is not on this page as a result, because an arm run an hour later is partly measuring the model's hour. Read the three together: deja cut the bill every time, by about half when the arms share the same hours, and what moves the saving run to run is whether the agent runs the recalled command straight away (about 16k tokens) or opens the files first to check it (33k to 100k).

Two things about the agentmemory arm belong in the text and not hidden in the number. It was given the same two prior sessions through its own import-jsonl, and it found them: recall returns the matching observations by id and title. The agent then goes and reads the repository anyway, and that is where its 15.9 tool calls go. It also runs keyless here, which is its own documented no-key mode. Attach a model and it summarises instead, which is a different arm and this page does not measure it.

What is not measured

Nothing here says anything about the memory a tool writes for itself over weeks, which is the thing most of them are actually built for. A warm store is a different measurement and probably a more important one.

Semantic recall with a model attached is also absent. deja has it as an option, MemPalace and agentmemory have it on by default once you give them a key, and none of those configurations are in this table.

Paraphrased questions are the case where every lexical tool loses ground, deja included. The warm benchmark puts numbers on that one.

Every driver, the corpus builder and the scoring rule are in the repository. If a row looks wrong, re-run it and tell me what you got.