MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
3.80T1 sourcearXiv cs.MA
Source record
Published by arXiv cs.MA (T1 source). The original is at https://arxiv.org/abs/2608.02613.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryMemArena benchmark for on-device personal memory assistants simulates 50 agents over 15 days, evaluating five open-weight LLMs with BM25-RAG, Oracle, Memobase, and MemSearch backends across six recall, reasoning, and trustworthiness dimensions. Key findings: backend choice outweighs reader scaling, permission-aware access fails universally, and search adds modest 7–87 ms latency on edge hardware.
Why it mattersConcrete benchmark numbers on on-device memory backends with an accompanying simulator release, directly useful for builders of private, edge-deployed agentic assistants.
Cited by
No citations on record.
