80.6% Answered. 98.3% Retrieved. Two Numbers, Two Different Jobs.
With a frontier reader, Velixar-backed answers reach 80.6% end-to-end. Underneath: 98.3% retrieval recall@25, a +55-point lift over the same reader without memory, and every condition printed.
The headline
Pair Velixar's memory layer with a frontier reader and the stack answers 80.6% of long-horizon memory questions end-to-end.
That's the number most vendors would put in a hero section and stop. We publish it with its conditions attached, plus the numbers underneath it — because the difference between those layers is the entire reason to buy a memory system rather than a bigger context window.
What's actually ours
Velixar is a retrieval and memory layer. The model that reads what we return is the customer's choice. So the number that describes our product is the one measured with no model in the loop at all.
On LoCoMo, the standard long-conversation memory benchmark, each question is labelled with the session holding its evidence. That makes retrieval scoring deterministic — did the labelled evidence come back, or not.
These two numbers come from different benchmarks, and we are not going to blur them. The 80.6% is LongMemEval; the 98.3% is LoCoMo. They measure different halves of the same problem on different corpora, and stacking them into a single funnel — "98 in, 80 out" — would be the kind of arithmetic this post exists to argue against. On LoCoMo, where we have both halves, retrieval runs at 98.3% and end-to-end lands at 68.9% with the same frontier reader. That gap is the point, and it is examined in Retrieved, Not Read.
What memory is worth
One more number, and it is the most attributable one we have: hold the reader constant and toggle only the memory.
A memory layer that can't show its lift against the same model answering unaided hasn't demonstrated it's doing anything.
Of these two numbers, the one that's yours is the retrieval: 98.3% recall, measured with no model in the loop, unchanged when the reader is deprecated. The 80.6% is a joint figure with a reader you rent and replace. What compounds and stays is the institutional knowledge the memory accumulates — what lasts is that corpus, not the model reading it.
Why we print the conditions
Because we measured what happens when you don't. In The Jury Nobody Names we ran a complete 7×7 grid — seven readers, seven judges, 49 cells, identical retrieved context — and found:
- 22.2 points of spread from reader choice alone
- 6.8 points from judge choice alone
- +0.9 points mean self-preference — real, and the smallest of the three
A 22-point swing from the reader is larger than the published gap between most competing memory systems. An end-to-end score that doesn't name its reader, judge, and retrieval budget isn't a measurement — it's a configuration.
That applies to our own history. Earlier figures we published for this benchmark were produced with a single model both reading and grading, on single runs. They came out higher — and we are reporting the lower number, because it is the one measured with a judge distinct from the reader, across full runs, averaged over six graders. When better method moves a number down, the number moves down.
The comparison we won't fake
You won't find a competitor leaderboard here. Lining our 80.6% up against another system's end-to-end score, run with a different reader and a different judge, would carry no information — our own grid proves it. The only honest comparison is both systems in one harness with one reader and one judge.
If you want that comparison, run us in yours. Every figure above came through the standard client path against the shipping product, and the conditions are printed on all of them.
Full methodology, recall curves, and per-cell grid data: Retrieved, Not Read and The Jury Nobody Names.