category · 11 records
benchmarks
- rec 2026-08-21BENCHMARKS5 min
The Librarian Test
Bring the right books, check the card, truly withdraw the withdrawn. GateMem scores all three at once — and Velixar is the only librarian that passes the whole test.
- rec 2026-07-27BENCHMARKS9 min
What the Numbers Actually Say
Velixar on HaluMem — extraction leads the field, QA and Update don't, and we ran down both confounds before publishing either result.
- rec 2026-07-26BENCHMARKS13 min
The Corpus Outlives the Model
98.3% recall, a 97.4% complete evidence set, and a 22.2-point spread in what seven models do with them — the case for a component-isolated evaluation standard.
- rec 2026-07-26BENCHMARKS12 min
The Corpus Outlives the Model
98.3% recall, a 97.4% complete evidence set, and a 22.2-point spread in what seven models do with them — the case for a component-isolated evaluation standard.
- rec 2026-07-22BENCHMARKS3 min
What Memory Is Worth: +55 Points, Same Model
Hold the reader constant and toggle only the memory: 13.9% closed-book becomes 68.9% with Velixar. Underneath it, retrieval at 98.3% recall@25 — and the measurement conditions nobody else prints.
- rec 2026-07-22BENCHMARKS7 min
Measuring Memory
An end-to-end memory benchmark score is a joint measurement of the memory system and the model reading it. We measured how much — and tightened our protocol until the number held.
- rec 2026-07-21BENCHMARKS5 min
Same Memory. Same Questions. 22 Points Apart.
We ran one memory system through seven reader models and seven judges — all 49 combinations. The scores say more about the industry's benchmarks than about any memory system.
- rec 2026-07-21BENCHMARKS5 min
80.6% Answered. 98.3% Retrieved. Two Numbers, Two Different Jobs.
With a frontier reader, Velixar-backed answers reach 80.6% end-to-end. Underneath: 98.3% retrieval recall@25, a +55-point lift over the same reader without memory, and every condition printed.
- rec 2026-07-21BENCHMARKS6 min
Retrieved, Not Read
Measured in isolation, Velixar's retrieval returns the labelled evidence for 98.3% of LoCoMo questions — and the complete evidence set for 97.4%.
- rec 2026-07-21BENCHMARKS9 min
The Jury Nobody Names
Velixar's retrieval returns the labelled evidence for 98.3% of questions. Then we measured what happens after retrieval — across every pairing of 7 readers and 7 judges.
- rec 2026-07-19BENCHMARKS1 min
Measuring memory, in the open
Our first research report is up: why end-to-end memory benchmark scores are joint measurements — and how we evaluate Velixar honestly.