append-only · 11 records

The Ledger

Engineering notes, benchmark results, and release records. Corrections are appended, never overwritten.

  1. rec 2026-08-21BENCHMARKS5 min

    The Librarian Test

    Bring the right books, check the card, truly withdraw the withdrawn. GateMem scores all three at once — and Velixar is the only librarian that passes the whole test.

  2. rec 2026-08-08ENGINEERING9 min

    The Real Problem With Long Context

    Long context beat Velixar on GateMem, and we published it. The real question isn't benchmark scale — it's what happens to cost, latency, and governance as history grows.

  3. rec 2026-07-27ENGINEERING10 min

    We Thought a Bigger AI Model Would Win. The Benchmark Said Otherwise.

    We benchmarked Velixar on HaluMem expecting a bigger extraction model to win. It didn't — and our obvious fix for the weakest category made things worse. What component isolation taught us.

  4. rec 2026-07-27BENCHMARKS9 min

    What the Numbers Actually Say

    Velixar on HaluMem — extraction leads the field, QA and Update don't, and we ran down both confounds before publishing either result.

  5. rec 2026-07-26BENCHMARKS13 min

    The Corpus Outlives the Model

    98.3% recall, a 97.4% complete evidence set, and a 22.2-point spread in what seven models do with them — the case for a component-isolated evaluation standard.

  6. rec 2026-07-22BENCHMARKS3 min

    What Memory Is Worth: +55 Points, Same Model

    Hold the reader constant and toggle only the memory: 13.9% closed-book becomes 68.9% with Velixar. Underneath it, retrieval at 98.3% recall@25 — and the measurement conditions nobody else prints.

  7. rec 2026-07-21BENCHMARKS5 min

    Same Memory. Same Questions. 22 Points Apart.

    We ran one memory system through seven reader models and seven judges — all 49 combinations. The scores say more about the industry's benchmarks than about any memory system.

  8. rec 2026-07-21BENCHMARKS5 min

    80.6% Answered. 98.3% Retrieved. Two Numbers, Two Different Jobs.

    With a frontier reader, Velixar-backed answers reach 80.6% end-to-end. Underneath: 98.3% retrieval recall@25, a +55-point lift over the same reader without memory, and every condition printed.

  9. rec 2026-07-19COMPANY2 min

    Introducing the Ledger

    Velixar's public record surface. Every post is an append-only entry — corrections are appended, never silently overwritten.

  10. rec 2026-07-19BENCHMARKS1 min

    Measuring memory, in the open

    Our first research report is up: why end-to-end memory benchmark scores are joint measurements — and how we evaluate Velixar honestly.

  11. rec 2026-07-13RELEASE1 min

    One command to connect any MCP host

    velixar-mcp-server now installs itself into Claude Desktop, Cursor, Windsurf, or any MCP host with a single command — no hand-edited JSON.