MemFit: Efficient Long-Term Agentic Memory
Most memory systems run new messages through an LLM before storing them. MemFit stores each turn word for word with no LLM call and puts the work into retrieval, using lexical and semantic search, two ways of widening the candidate pool, and a cross-encoder. It had the best overall results on LoCoMo, LongMemEval-S, and MemGallery and built memory 3.9 to 52.6× faster than the systems we compared against.
Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations. To address this limitation, we propose MemFit, a long-term memory system for conversational agents that reduces the cost and latency of memory operations. Unlike existing systems that rely on expensive LLM calls for memory construction or discard surface-level details through compression, MemFit stores each turn verbatim in an append-only store with near-instantaneous, LLM-free insertion, indexing turns with segment summaries rather than replacing them. Additionally, MemFit uses an LLM-free, multi-path retrieval strategy that combines lexical and semantic signals with cross-encoder reranking over caption-augmented episodes in both textual and multimodal settings. Empirical results on three widely used benchmarks, LoCoMo, MemGallery, and LongMemEval-S, show that MemFit achieves state-of-the-art performance while reducing memory construction time and cost several-fold, providing a scalable and efficient solution for persistent agentic memory.
@misc{piehl2026memfit,
title = {{MemFit}: Efficient Long-Term Agentic Memory},
author = {Piehl, Mitchell and Ye, Muchao},
year = {2026},
note = {Under review}
}