AI · 2h ago
AI memory benchmark LOCOMO excludes 'I don't know' questions, skewing scores
LOCOMO, a key AI memory benchmark, excludes 22.5% of its questions that require abstention, and instructs models to never say 'not specified'. Mem0's harness removes category 5 adversarial questions and forces commitment, inflating accuracy scores. An independent audit reveals additional leniencies like partial credit and no retrieval budget.
Meridian48 take
The benchmark's design choices, while defensible in isolation, collectively produce a metric that overstates real-world memory performance by ignoring the hardest test: knowing when to say nothing.
Read the full reporting
The AI-memory benchmark everyone quotes forbids saying “I don't know” →
DEV Community
ai-memory-benchmarkbenchmark-flaws