What Memory Is Worth: +55 Points, Same Model
Hold the reader constant and toggle only the memory: 13.9% closed-book becomes 68.9% with Velixar. Underneath it, retrieval at 98.3% recall@25 — and the measurement conditions nobody else prints.
The one number that's attributable
Most memory benchmarks blend two things and report one number: whether the right evidence was retrieved, and whether the model reading it answered correctly. The blend moves for reasons that have nothing to do with the memory layer — we've measured a 22-point swing from reader choice alone. So the headline we lead with is the one comparison where nothing moves except the memory:
A memory layer that can't show its lift against the same model answering unaided hasn't demonstrated it's doing anything. This is ours.
Why the lift is that large
Because the evidence is almost always on the table. Measured with no model in the loop at all — scored against LoCoMo's own evidence labels — Velixar returns the labelled evidence for 98.3% of questions at a retrieval budget of 25, and the complete evidence set for 97.4% (n=310). No reader, no judge, no prompt, nothing to drift when a model gets deprecated.
When retrieval runs at 98%, the remaining error is downstream — which is exactly what we found when we measured downstream exhaustively.
The +55 points are the memory's, not the reader's — the same model produces them only when Velixar is behind it. That's the distinction that matters at purchase. The reader is a commodity you'll swap every quarter; the institutional knowledge your system accumulates — what it has learned about your customers, your codebase, your operation — is the asset that compounds and survives every model change. You are not buying a better reader. You are buying the part that lasts.
The hidden variables, measured
In The Jury Nobody Names we ran every pairing of seven reader models and seven judge models on identical retrieved context — 49 cells. Three spreads fell out:
- 22.2 points from reader choice alone
- 6.8 points from judge choice alone
- +0.9 points mean self-preference — real, and the smallest of the three by far
That reader spread is larger than the published gap between most competing memory systems. An end-to-end score that doesn't name its reader, judge, and retrieval budget isn't a measurement — it's a configuration. Ours are printed on every figure.
The comparison we won't fake
No competitor leaderboard here. Our own grid proves that lining up two end-to-end scores produced with different readers and judges carries no information. The only honest comparison is both systems in one harness, one reader, one judge — and every figure above came through the standard client path against the shipping product, so you can run that comparison yourself.
Full methodology, recall curves, and per-cell grid data: Retrieved, Not Read · The Jury Nobody Names
Velixar is governed memory infrastructure — persistent, auditable memory for AI systems. Get in touch to run the benchmark against your own stack.