rec 2026-07-21v1.0benchmarks5 min

80.6% Answered. 98.3% Retrieved. Two Numbers, Two Different Jobs.

With a frontier reader, Velixar-backed answers reach 80.6% end-to-end. Underneath: 98.3% retrieval recall@25, a +55-point lift over the same reader without memory, and every condition printed.

Velixar AI

The headline

Pair Velixar's memory layer with a frontier reader and the stack answers 80.6% of long-horizon memory questions end-to-end.

80.6%
End-to-end accuracy — LongMemEval, n=120, k=25, reader gpt-5.4, averaged over 6 independent judges
End-to-end accuracy is a joint measurement of the memory layer and the reader model. We name every condition because the number moves without them — by 22.2 points, as it turns out. Averaged across judges rather than reported under a single favourable grader; the best individual cell was 81.7%.

That's the number most vendors would put in a hero section and stop. We publish it with its conditions attached, plus the numbers underneath it — because the difference between those layers is the entire reason to buy a memory system rather than a bigger context window.

What's actually ours

Velixar is a retrieval and memory layer. The model that reads what we return is the customer's choice. So the number that describes our product is the one measured with no model in the loop at all.

On LoCoMo, the standard long-conversation memory benchmark, each question is labelled with the session holding its evidence. That makes retrieval scoring deterministic — did the labelled evidence come back, or not.

98.3%
Retrieval recall@25 — LoCoMo, n=310 evidence-bearing questions
A retrieval figure, not an answer-accuracy figure, and not comparable to any published end-to-end score — those measure something else. Counting only questions where EVERY required evidence session came back: 97.4%. No reader, no judge, no prompt, nothing to drift when a model is deprecated.

These two numbers come from different benchmarks, and we are not going to blur them. The 80.6% is LongMemEval; the 98.3% is LoCoMo. They measure different halves of the same problem on different corpora, and stacking them into a single funnel — "98 in, 80 out" — would be the kind of arithmetic this post exists to argue against. On LoCoMo, where we have both halves, retrieval runs at 98.3% and end-to-end lands at 68.9% with the same frontier reader. That gap is the point, and it is examined in Retrieved, Not Read.

98.3%
Retrieved · recall@25
LoCoMo — no model in the loop
68.9%
Answered end-to-end
same corpus, same frontier reader

What memory is worth

One more number, and it is the most attributable one we have: hold the reader constant and toggle only the memory.

+55.0 pts
Memory lift — LoCoMo, reader held constant on both arms, paired questions
The same model answers 13.9% of these questions closed-book and 68.9% with Velixar behind it. Because the reader is identical on both sides, the difference is attributable to the memory layer and nothing else.
025507510013.9%68.9%Closed-booksame modelWith Velixarsame model+55.0
One reader on both sides. Removing only the memory drops the same model from 68.9% to 13.9% — a +55.0-point lift that belongs to the memory layer. LoCoMo, paired, k=25.

A memory layer that can't show its lift against the same model answering unaided hasn't demonstrated it's doing anything.

Which number comes with you

Of these two numbers, the one that's yours is the retrieval: 98.3% recall, measured with no model in the loop, unchanged when the reader is deprecated. The 80.6% is a joint figure with a reader you rent and replace. What compounds and stays is the institutional knowledge the memory accumulates — what lasts is that corpus, not the model reading it.

Why we print the conditions

Because we measured what happens when you don't. In The Jury Nobody Names we ran a complete 7×7 grid — seven readers, seven judges, 49 cells, identical retrieved context — and found:

  • 22.2 points of spread from reader choice alone
  • 6.8 points from judge choice alone
  • +0.9 points mean self-preference — real, and the smallest of the three
0510152025Spread in score (percentage points)Reader choice22.2 ptsJudge choice6.8 ptsSelf-preference0.9 pts
The score moves most when you change the reader (22.2 pts), less with the judge (6.8 pts), and least from a model preferring its own output (0.9 pts). The memory layer is held fixed throughout.

A 22-point swing from the reader is larger than the published gap between most competing memory systems. An end-to-end score that doesn't name its reader, judge, and retrieval budget isn't a measurement — it's a configuration.

That applies to our own history. Earlier figures we published for this benchmark were produced with a single model both reading and grading, on single runs. They came out higher — and we are reporting the lower number, because it is the one measured with a judge distinct from the reader, across full runs, averaged over six graders. When better method moves a number down, the number moves down.

The comparison we won't fake

You won't find a competitor leaderboard here. Lining our 80.6% up against another system's end-to-end score, run with a different reader and a different judge, would carry no information — our own grid proves it. The only honest comparison is both systems in one harness with one reader and one judge.

If you want that comparison, run us in yours. Every figure above came through the standard client path against the shipping product, and the conditions are printed on all of them.

Full methodology, recall curves, and per-cell grid data: Retrieved, Not Read and The Jury Nobody Names.