rec 2026-07-22v1.0benchmarks3 min

What Memory Is Worth: +55 Points, Same Model

Hold the reader constant and toggle only the memory: 13.9% closed-book becomes 68.9% with Velixar. Underneath it, retrieval at 98.3% recall@25 — and the measurement conditions nobody else prints.

Velixar AI

The one number that's attributable

Most memory benchmarks blend two things and report one number: whether the right evidence was retrieved, and whether the model reading it answered correctly. The blend moves for reasons that have nothing to do with the memory layer — we've measured a 22-point swing from reader choice alone. So the headline we lead with is the one comparison where nothing moves except the memory:

+55.0 pts
Accuracy lift from memory — reader held constant, paired questions, LoCoMo, k=25
The same model answers 13.9% of these questions closed-book and 68.9% with Velixar. One reader on both sides — the difference is attributable to memory, not to model choice.

A memory layer that can't show its lift against the same model answering unaided hasn't demonstrated it's doing anything. This is ours.

025507510013.9%68.9%Closed-booksame modelWith Velixarsame model+55.0
One reader on both sides. Removing only the memory drops the same model from 68.9% to 13.9% — a +55.0-point lift that belongs to the memory layer. LoCoMo, paired, k=25.

Why the lift is that large

Because the evidence is almost always on the table. Measured with no model in the loop at all — scored against LoCoMo's own evidence labels — Velixar returns the labelled evidence for 98.3% of questions at a retrieval budget of 25, and the complete evidence set for 97.4% (n=310). No reader, no judge, no prompt, nothing to drift when a model gets deprecated.

98.3%
Labelled evidence retrieved
recall@25 · LoCoMo, n=310
97.4%
Complete evidence set
every required session returned

When retrieval runs at 98%, the remaining error is downstream — which is exactly what we found when we measured downstream exhaustively.

What you're actually buying

The +55 points are the memory's, not the reader's — the same model produces them only when Velixar is behind it. That's the distinction that matters at purchase. The reader is a commodity you'll swap every quarter; the institutional knowledge your system accumulates — what it has learned about your customers, your codebase, your operation — is the asset that compounds and survives every model change. You are not buying a better reader. You are buying the part that lasts.

The hidden variables, measured

In The Jury Nobody Names we ran every pairing of seven reader models and seven judge models on identical retrieved context — 49 cells. Three spreads fell out:

  • 22.2 points from reader choice alone
  • 6.8 points from judge choice alone
  • +0.9 points mean self-preference — real, and the smallest of the three by far
0510152025Spread in score (percentage points)Reader choice22.2 ptsJudge choice6.8 ptsSelf-preference0.9 pts
The score moves most when you change the reader (22.2 pts), less with the judge (6.8 pts), and least from a model preferring its own output (0.9 pts). The memory layer is held fixed throughout.

That reader spread is larger than the published gap between most competing memory systems. An end-to-end score that doesn't name its reader, judge, and retrieval budget isn't a measurement — it's a configuration. Ours are printed on every figure.

The comparison we won't fake

No competitor leaderboard here. Our own grid proves that lining up two end-to-end scores produced with different readers and judges carries no information. The only honest comparison is both systems in one harness, one reader, one judge — and every figure above came through the standard client path against the shipping product, so you can run that comparison yourself.

Full methodology, recall curves, and per-cell grid data: Retrieved, Not Read · The Jury Nobody Names


Velixar is governed memory infrastructure — persistent, auditable memory for AI systems. Get in touch to run the benchmark against your own stack.