category · 2 records

engineering

← all records

  1. rec 2026-08-08ENGINEERING9 min

    The Real Problem With Long Context

    Long context beat Velixar on GateMem, and we published it. The real question isn't benchmark scale — it's what happens to cost, latency, and governance as history grows.

  2. rec 2026-07-27ENGINEERING10 min

    We Thought a Bigger AI Model Would Win. The Benchmark Said Otherwise.

    We benchmarked Velixar on HaluMem expecting a bigger extraction model to win. It didn't — and our obvious fix for the weakest category made things worse. What component isolation taught us.