Publications

Newest first. You can also find these on Google Scholar.

2026

Under review

MemFit: Efficient Long-Term Agentic Memory

Mitchell Piehl, Muchao Ye

Most memory systems run new messages through an LLM before storing them. MemFit stores each turn word for word with no LLM call and puts the work into retrieval, using lexical and semantic search, two ways of widening the candidate pool, and a cross-encoder. It had the best overall results on LoCoMo, LongMemEval-S, and MemGallery and built memory 3.9 to 52.6× faster than the systems we compared against.

Summary

Long-term memory systems for large language models (LLMs) have gained popularity for extending reasoning capabilities across applications. Current memory systems rely on LLM agents to organize and consolidate memory, resulting in costly, inefficient write operations. To address this limitation, we propose MemFit, a long-term memory system for conversational agents that reduces the cost and latency of memory operations. Unlike existing systems that rely on expensive LLM calls for memory construction or discard surface-level details through compression, MemFit stores each turn verbatim in an append-only store with near-instantaneous, LLM-free insertion, indexing turns with segment summaries rather than replacing them. Additionally, MemFit uses an LLM-free, multi-path retrieval strategy that combines lexical and semantic signals with cross-encoder reranking over caption-augmented episodes in both textual and multimodal settings. Empirical results on three widely used benchmarks, LoCoMo, MemGallery, and LongMemEval-S, show that MemFit achieves state-of-the-art performance while reducing memory construction time and cost several-fold, providing a scalable and efficient solution for persistent agentic memory.

@misc{piehl2026memfit,
  title  = {{MemFit}: Efficient Long-Term Agentic Memory},
  author = {Piehl, Mitchell and Ye, Muchao},
  year   = {2026},
  note   = {Under review}
}
Under review

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection

Mitchell Piehl, Muchao Ye

Video anomaly detection with a frozen vision-language model, run entirely at test time. LATERN keeps a checked, image-grounded memory of earlier segments and groups segment scores into events, so each incident gets one explanation. It had the best AUC on UCF-Crime and the best AP on XD-Violence among the frozen-model methods we compared.

arXiv Summary

Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning capabilities and natural-language-based explainability. In this paper, we address a key limitation of such pipelines: they perform segment-level inference independently due to token constraints and reason without structured temporal context, resulting in fragmented predictions and inconsistent explanations. We instead let VLMs interpret anomalies as deviations from evolving video dynamics. To specify, we propose a context-aware framework named LATERN, which reformulates VAD as a temporal evidence aggregation process. LATERN consists of two complementary modules: Context-Aware Anomaly Scoring (CEA) and Recursive Evidence Aggregation (REA). CEA introduces a novel image-grounded memory mechanism that selects historical content based on frame diversity and visual-textual alignment, serving as expanded context to help generate reliable anomaly scores. Building upon these scores, REA performs recursive temporal aggregation to identify coherent anomaly intervals and produce event-level decisions and explanations grounded in visual-textual evidence. Experiments on standard VAD benchmarks, UCF-Crime and XD-Violence, show that LATERN enhances detection accuracy and explanation consistency for frozen VLMs during test time, while generating temporally coherent and semantically grounded event-level explanations.

@misc{piehl2026latern,
  title         = {{LATERN}: Test-Time Context-Aware Explainable Video Anomaly Detection},
  author        = {Piehl, Mitchell and Ye, Muchao},
  year          = {2026},
  eprint        = {2605.15054},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV}
}
Under review

ER-MIA: Black-Box Adversarial Memory Injection Attacks on Long-Term Memory-Augmented Large Language Models

Mitchell Piehl, Zhaohan Xi, Zuobin Xiong, Pan He, Muchao Ye

Black-box attacks that plant false memories through normal conversation and use SVD-based embedding optimization to get them retrieved. We test them on A-MEM, Mem0, and LightMem with three victim models, and we evaluate a write-time filter as a defense.

arXiv Summary

Long-term memory systems are increasingly adopted to enable persistent reasoning in large language models (LLMs) beyond finite context windows, but they simultaneously expand the attack surface against LLMs. However, adversarial attacks against memory retrieval mechanisms remain underexplored, especially in realistic black-box settings. We present a systematic study of black-box adversarial memory injection attacks targeting similarity-based retrieval in memory-augmented LLMs. Specifically, we introduce ER-MIA, a unified framework that exposes this vulnerability and formalizes two realistic attack settings: content-based attacks and question-targeted attacks. In these settings, ER-MIA includes an arsenal of composable attack primitives and ensemble attacks that combine LLM-based heuristic generation with SVD-based embedding optimization to achieve high success rates even under strict black-box constraints. Experiments across three memory systems and three victim LLMs show that targeted injected memories reduce question answering accuracy by up to 25.1 points, that larger victim models are no more robust, and that attacks optimized against one embedding model transfer to a system retrieving with a different one.

@misc{piehl2026ermia,
  title         = {{ER-MIA}: Black-Box Adversarial Memory Injection Attacks on Long-Term Memory-Augmented Large Language Models},
  author        = {Piehl, Mitchell and Xi, Zhaohan and Xiong, Zuobin and He, Pan and Ye, Muchao},
  year          = {2026},
  eprint        = {2602.15344},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG}
}

2025

ICMLA 2025

Solving Math Word Problems Using Estimation Verification and Equation Generation

Mitchell Piehl, Dillon Wilson, Ananya Kalita, Jugal Kalita

IEEE International Conference on Machine Learning and Applications (ICMLA), 2025

The LLM writes equations for a symbolic solver, then estimates the answer the way math teachers tell students to, and uses that estimate to check the exact result. New state-of-the-art results on numeric and algebraic word problems, plus two new datasets, SVAMPClean and Trig300.

arXiv Code

Large Language Models (LLMs) excel at various tasks, including problem-solving and question-answering. However, LLMs often find Math Word Problems (MWPs) challenging because solving them requires a range of reasoning and mathematical abilities with which LLMs seem to struggle. Recent efforts have helped LLMs solve more complex MWPs with improved prompts. This study proposes a novel method that initially prompts an LLM to create equations from a decomposition of the question, followed by using an external symbolic equation solver to produce an answer. To ensure the accuracy of the obtained answer, inspired by an established recommendation of math teachers, the LLM is instructed to solve the MWP a second time, but this time with the objective of estimating the correct answer instead of solving it exactly. The estimation is then compared to the generated answer to verify. If verification fails, an iterative rectification process is employed to ensure the correct answer is eventually found. This approach achieves new state-of-the-art results on datasets used by prior published research on numeric and algebraic MWPs, improving the previous best results by nearly two percent on average. In addition, the approach obtains satisfactory results on trigonometric MWPs, a task not previously attempted to the authors' best knowledge. This study also introduces two new datasets, SVAMPClean and Trig300, to further advance the testing of LLMs' reasoning abilities.

@inproceedings{piehl2025evoss,
  title     = {Solving Math Word Problems Using Estimation Verification and Equation Generation},
  author    = {Piehl, Mitchell and Wilson, Dillon and Kalita, Ananya and Kalita, Jugal},
  booktitle = {2025 International Conference on Machine Learning and Applications (ICMLA)},
  year      = {2025},
  publisher = {IEEE}
}