arXiv NLP
techCenter
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Modelstranslating…
1 min readUnknownarXiv Press Group
arXiv:2608.26295v1 Announce Type: new
Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing evaluations measure source preference without establishing source correctness. We introduce…