rec 2026-07-19v1.0benchmarks1 min

Measuring memory, in the open

Our first research report is up: why end-to-end memory benchmark scores are joint measurements — and how we evaluate Velixar honestly.

Velixar

We published our first research report: Measuring Memory.

The short version: an end-to-end memory benchmark score is not a property of the memory system alone. It's a joint measurement of retrieval, the reader model, and the judge. Freeze the retrieval and swap only the reader, and our LongMemEval score moves 16.4 points — larger than the published gap between several competing systems.

The report lays out what that means for how these numbers should be read, where Velixar lands (LongMemEval-S, 82.1% ± 1.8, with its full configuration), and the bar we hold a number to before treating any ranking as settled. Every headline figure ships with its caveat — because anyone can run Velixar and check.

Read the report →